<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Analysing the Role of Key Term In ections in Knowledge Discovery on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ali Hurriyetoglu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jurjen Wagemaker</string-name>
          <email>wagemaker@floodtags.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nelleke Oostdijk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antal van den Bosch</string-name>
          <email>a.vandenboschg@let.ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Language Studies, Radboud University</institution>
          ,
          <addr-line>P.O. Box 9103, NL-6500 HD, Nijmegen</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Floodtags</institution>
          ,
          <addr-line>Binckhorstlaan 36, M2.11, 2511 BE Den Haag</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Statistics Netherlands</institution>
          ,
          <addr-line>P.O. Box 4481, 6401 CZ Heerlen</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We introduce our methodology for collecting tweets and identifying event-related actionable information using key terms and in ections. The vast amount of user-generated content makes it challenging to detect relevant information. Therefore, we aim to facilitate extracting morphological, syntactic, and semantic features of a key term semiautomatically. The results of our study show that handling the in ections of key terms separately has its advantages: in the disaster scenario that we are investigating we are successful in discovering relevant and irrelevant information e ectively.</p>
      </abstract>
      <kwd-group>
        <kwd>knowledge discovery</kwd>
        <kwd>in ection analysis</kwd>
        <kwd>ood</kwd>
        <kwd>text mining</kwd>
        <kwd>machine learning</kwd>
        <kwd>social media analysis</kwd>
        <kwd>Twitter</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Twitter provides a platform that contains a rich source of user generated content.
Any user of this platform can post tweets, and language use tends to be rich and
creative. Consequently, increasingly large quantities of novel content is awaiting
proper analysis on Twitter. Especially, in a disaster context, the need of detecting
precise and complete information is in urgent need of a solution.</p>
      <p>In this paper, we present our analysis method and apply it on a tweet
collection concerning oods. The method consists of collecting data with a single key
term (a stem, and its in ections), pre-processing, extracting extensive features
for use in machine learning, and determining coherent clusters in tweet subsets.
Human input is used when labeling the clusters.</p>
      <p>A central concept in our method is the information thread, i.e, a group of
related tweets. Relatedness is determined by the expert who uses the method.
For example, the word ` ood' has multiple senses, including `to cover or ll with
water' but also `to come in great quantities'4. A water manager will probably
want to focus on only the water-related sense. At the same time, he will want
to discriminate between di erent contextualizations of this sense: past, current,
up-coming events, e ects, measures taken, etc. By incrementally clustering and
labeling the tweets, the collection is analyzed into di erent information threads.</p>
      <p>Here, we report on the e ect that the use of key term in ections has on
knowledge discovery in a tweet collection. In the following sections, we summarize
the related research, the tool we developed to support our methodology, the data
we collected, and the results obtained.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Studies</title>
      <p>
        Identifying di erent uses of a word is a word sense induction task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which is
especially challenging for tweets [
        <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
        ]. Since a tweet collection potentially contains
innumerable information threads, we propose an incremental-iterative search for
information threads in which we include the human in the loop to determine the
nal result. By means of this approach, we can manage the ambiguity of key
terms as well as the complexity of a tweet collection5.
      </p>
      <p>
        Detecting relevant information in mass emergencies has considerable
complexity and urgency [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Therefore, we direct our e orts to contribute to this line
of research without sacri cing precision, speed, and recall.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Relevancer</title>
      <p>We implemented an open source tool, Relevancer6. Relevancer supports the
analysis of the tweet collections with the following steps:
Preprocessing Retweets that are identi ed by the respective JSON eld of an
API response and the tweets where `RT @' occurs at the beginning of the
tweet text are eliminated. Moreover, user names and URLs are converted to
`usrusrusr' and `urlurlurl' respectively. After that, we exclude exact duplicate
tweets.</p>
      <p>Feature extraction Any token that occurs in a tweet text is used as a
feature. Tokens are detected based on a split on a single punctuation mark or
an arbitrary number of space characters. Features can be words, hashtags,
letters, numbers, letter and number combinations, emoticons, or 2, 3, or 4
length punctuation mark combinations.</p>
      <p>Near-duplicate detection Tweets that contain the same pentagram (here: 5
consecutive words larger than 2 letters), or have a cosine similarity higher</p>
      <sec id="sec-3-1">
        <title>4 http://www.oed.com/view/Entry/71808</title>
        <p>5 The article in the following URL provides an excellent example of the
ambiguity caused by lexical meaning and syntax: http://speld.nl/2016/05/22/
man-rijdt-met-180-kmu-over-a2-van-harkemase-boys/
6 https://bitbucket.org/hurrial/relevancer
than 0.85, based on the features extracted in the previous step, are
considered as a group of near duplicates. We keep only one tweet from each
near-duplicate group of tweets.</p>
        <p>Information thread detection Information threads related to the key term
and its in ections are detected using an unsupervised clustering method, viz.
K-Means. Clustering and cluster selection steps are repeated in iterations
until the requested number of coherent clusters is reached. The tweets in the
detected clusters are not included in the following clustering iteration.
Annotation An expert assigns an information thread label to each cluster
based on its coherency and relevancy.</p>
        <p>Remaining and new tweets can be clustered or analyzed with the knowledge
discovered in previous iterations of the clustering and annotation process.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Use Case: Flood tweet collection</title>
      <p>We collected tweets from the public Twitter API7 with the key term ` ood'
and its in ections ` oods', ` ooded', and ` ooding' between December 17 and 31
2015. We applied the analysis steps supported by Relevancer to each subset.</p>
      <p>Detailed statistics of the collected tweets are represented in Table 1. The
number of collected tweets are provided in the column #All. The columns #unique
and unique% contain the counts after eliminating the retweets and duplicate
tweets and the percentage of unique tweets in each subset of the collection.
The unique tweet ratio for the terms ` ooded' and ` ood' is highest and lowest
respectively.</p>
      <p>Each subset of unique tweets was clustered in order to identify information
threads. The annotation was performed with the labels `relevant', `irrelevant, and
`incoherent', which are the most general information threads that can be handled
with this approach. We generated 50 clusters for each subset. The number of
tweets that were placed in a cluster is presented in the column clustered.</p>
      <p>A cluster is relevant, if it is about a relevant information thread, e.g. an
actionable insight, an observation, witness reaction, event information available
from citizens or authorities that can help people avoid harm. Otherwise, the</p>
      <sec id="sec-4-1">
        <title>7 https://dev.twitter.com/rest/public</title>
        <p>label is irrelevant. Clusters that are not clearly about any information thread
are labeled as incoherent. Respective columns in Table 1 provide information
about the number of tweets in each thread. Having many incoherent clusters
points towards the ambiguity of a term.</p>
        <p>The detailed cluster analysis revealed the characteristics of the information
threads for the key term and its in ections. The key term ` ood' (stem form) is
mostly used by the authorities and automatic bots that provide updates about
disasters in the form of a ` ood alert' or ` ood warning'. Moreover, mentioning
the name of the authorities, e.g, ood advisory, enables tweets to fall in the same
cluster. The irrelevant tweets using this term are about ads or product names.</p>
        <p>Each in ection of the term ` ood' has a di erent set of uses with a
considerable number of overlapping uses. The form ` ooded' is mostly used for expressing
observations and opinions toward a disaster event and news article related tweets.
General comments and expressions of empathy toward the victims of disasters
are found in tweets that contain the form ` oods'. Finally, the form ` ooding'
mostly occurs in tweets that are about the consequences of a disaster.</p>
        <p>Another aspect that emerged from the cluster analysis is that common and
speci c multi-word expressions containing the key term or one of its in ections,
e.g. ` ooding back', ` ooding timeline', form at least a cluster around them.
Tweets that contain such expressions can be transferred from the remaining
tweets set to the respective cluster. For example, we identi ed 806 and 592
tweets that contain ` ooding back' and ` ooding timeline' respectively.</p>
        <p>Finally, named entities, such as the name of a storm, river, bridge, road,
web platform, person, institution, or place, and emoticons enable the clustering
algorithm to detect coherent clusters of tweets containing such named entities.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Directions</title>
      <p>A novel tweet collection and analysis methodology was presented in order to
discover detailed and precise information about disaster events on a microblogging
platform and in a challenging domain. Our results shows that determining and
handling separate uses of a key term and its in ections reveal di erent angles of
the knowledge we can discover in tweet collections. The majority of the tweets
that were not placed in any cluster can be handled with the knowledge we
discovered in the annotation phase, which provides insights about the semantic and
syntactic space each subset covers.</p>
      <p>This methodology can be applied to any term in any language. The labeled
tweets can be used to create supervised machine learning models in order to
handle the remaining tweets in the collection as well as new tweets.
Acknowledgments. This research was supported by the Dutch national
research programme COMMIT. We gratefully acknowledge the contribution by
Floodtags.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gella</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <source>One Sense per Tweeter... and Other Lexical Semantic Tales of Twitter. Proceedings of the 14th Conference of the European Chapter of the Association for Computational</source>
          Linguistics pp.
          <volume>215</volume>
          {
          <issue>220</issue>
          (
          <year>2014</year>
          ), http://www.aclweb.org/anthology/E14-4042
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Imran</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieweg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Processing social media messages in mass emergency: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 47(4)</source>
          ,
          <volume>67</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lau</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Computing</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Word sense induction for novel sense detection</article-title>
          .
          <source>Proceedings of the 13th Conference of the European Chapter of the Association for computational Linguistics (EACL</source>
          <year>2012</year>
          ) pp.
          <volume>591</volume>
          {
          <issue>601</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mccarthy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Apidianaki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erk</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Word Sense Clustering and Clusterability</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>42</volume>
          (
          <issue>2</issue>
          ),
          <volume>4943</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>