<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The good, the bad and the ugly</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Markus Strohmaier</string-name>
          <email>markus.strohmaier@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GESIS &amp; University of Koblenz Unter Sachsenhausen</institution>
          <addr-line>6-8 50667 Cologne</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>General Terms Experimentation</institution>
          ,
          <addr-line>Human Factors</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>1141</volume>
      <abstract>
        <p>According to the Computational Social Science Society of the Americas (CSSSA), computational social science is “The science that investigates social phenomena through the medium of computing and related advanced information processing technologies”. Positioned between the computer and social sciences, this new and emerging interdisciplinary field is fuelled by at least the following two developments: (i) availability of data: With the web, a huge volume of social data is now available which enables the study of traces of social interactions on new scales. (ii) increasing quantification of social theories: With recent advances in the social sciences, social theories become increasingly formal and/or mathematical and thus amenable to quantification. Taken together, these two developments give rise to a whole range of new and interesting problems on the intersection between computer and social sciences. While a multitude of social data is available on the World Wide Web, microblogs are of particular interest due to their real-time nature, their rich social fabric and their presumed on/offline coupling. In this talk, I am going to talk about the potentials and the challenges of doing computational social science based on data obtained from microblogs such as Twitter. In particular, I want to present previous work by my group and others to identify research avenues where progress has already been made or where progress is on the horizon, and contrast these with what I feel are open research challenges in this emerging field. Work that demonstrates the potential of microblogs for computational social science includes for example [1], where we have operationalized a number of theoretical constructs from sociology to characterize the nature of online conversational practices of political parties on Twitter. In another work, we have studied the ways in which users' fields of expertise can be inferred from microblog data [4]. Work that demonstrates the pitfalls and challenges of doing computational social science with microblog data include for example [5] where we have studied a network of bots who are competing against each other in attacking users on Twitter. In subsequent work, we have found that such attacks have Categories and Subject Descriptors J.4 [Social and behavioral sciences]: Sociology; I.6.0 [Simulation and Modeling]: General</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Social data</kwd>
        <kwd>computational social science</kwd>
        <kwd>social behavior</kwd>
        <kwd>web science</kwd>
        <kwd>online social networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>the potential to impact the social graph of Twitter [3], i.e.
the network of who follows whom respectively who replies
to whom. In other work, [2] have shown that there is a
stark difference between the demographics of Twitter and
the general population of the US, finding that Twitter users
significantly over-represent densely populated regions and
are predominantly male. I will argue that these and other
factors need to be considered when we aim to unlock the full
potential of microblog data for computational social science
purposes.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>