<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Word Embeddings for Information Extraction from Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Surajit Dasgupta</string-name>
          <email>surajit.techie@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abhash Kumar</string-name>
          <email>abhashmaddi@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dipankar Das</string-name>
          <email>ddas@cse.jdvu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sudip Kumar Naskar</string-name>
          <email>sudip.naskar@cse.jdvu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jadavpur University</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our approach on \Information Extraction from Microblogs Posted during Disasters" as an attempt in the shared task of the Microblog Track at Forum for Information Retrieval Evaluation (FIRE) 2016 [2]. Our method uses vector space word embeddings to extract information from microblogs (tweets) related to disaster scenarios, and can be replicated across various domains. The system, which shows encouraging performance, was evaluated on the Twitter dataset provided by the FIRE 2016 shared task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Social media plays a very important role in the
dissemination of real-time information such as disaster outbreaks.
E cient processing of information from social media
websites such as Twitter can help us to pursue proper disaster
mitigation strategies. Extracting relevant information from
tweets proves to be a challenging task, owing to their short
and noisy nature. Information extraction from social media
text is a well researched problem [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Approaches using bag-of-words model, n-grams based methods
and machine learning have been extensively used to extract
information from microblogs.
      </p>
    </sec>
    <sec id="sec-2">
      <title>TASK DEFINITION</title>
      <p>A set of tweets, T = ft1; t2; t3 : : : tng and a set of topics,
Q = fq1; q2; q3 : : : qmg are given. Each topic contains a title,
a brief description and a detailed narrative on what type
of tweets are considered relevant to the topic. The tweets
given in the task were posted during the Nepal earthquake1
in April 2015. Each topic contains a broad information need
during a disaster, such as { availability or requirement of
general or medical resources by the population in the
disaster a ected area, availability or requirement of resources in a
geographical region, reports of relief being carried out by an
organization and reports of damage to infrastructure. The
main objective of this task is to extract all tweets, ti 2 T
that are relevant to each topic, qj 2 Q with high precision
and high recall, and rank them in their order of relevance.
3.</p>
    </sec>
    <sec id="sec-3">
      <title>DATA AND RESOURCES</title>
      <p>This section describes the dataset and resources provided
to the shared task participants. A text le containing 50,068
tweet identi ers that were posted during the Nepal
earthquake in April 2015, was provided by the organizers. A
Python script was provided that downloaded the tweets
using the Twitter API2 into a JSON encoded tweet le which
was processed during the task. A text le of topic
descriptions in TREC3 format was provided, that contained
information necessary for the extraction of relevant tweets.
The topic le consisted of 7 topics: FMT1, FMT2, FMT3,
FMT4, FMT5, FMT6 and FMT7. Each topic consisted of
the following 4 sections:
&lt;num&gt; : Topic number.
&lt;title&gt; : Title of the topic.
&lt;desc&gt; : Description of the topic.
&lt;narr&gt; : A detailed narrative which describes what
types of tweets would be considered relevant to the
topic.
4.
4.1</p>
    </sec>
    <sec id="sec-4">
      <title>SYSTEM DESCRIPTION</title>
    </sec>
    <sec id="sec-5">
      <title>Preprocessing</title>
      <p>We parsed the JSON encoded tweets and retrieved the
following attributes { tweet identi er, tweet, geolocation. From
the tweets, we removed the Twitter handles starting with @,
URLs and all punctuation marks except the instances of a
single \ . " (period) and a single \ , " (comma), using
regular expressions. We removed the ASCII characters from
the tweets and converted the remaining tweet to lower case
characters.</p>
      <p>We also preprocessed the &lt;narr&gt; sections of the topic
le. We removed the punctuation marks, stop words and
converted the text to lower case characters. The
preprocessed &lt;narr&gt; sections for each topic was used for building
the word bags.
*Indicates equal contribution.
1http://en.wikipedia.org/wiki/April 2015 Nepal earthquake
2https://dev.twitter.com/overview/api/tweets
3http://trec.nist.gov
4.2</p>
      <p>To build the topic-speci c word bags, the preprocessed
&lt;narr&gt; section was manually checked to retain the relevant
words for each topic. The topic words were expanded using
the synonyms obtained from NLTK WordNet4. The past,
past participle and present continuous forms of verbs were
obtained using the NodeBox5 library for Python. Vowels,
except the initial character, were removed to create
unnormalized version of the words which are generally used in
Twitter owing to the 140 character limit. The resultant set
of words were used to create the word bags for each topic.
4.3</p>
    </sec>
    <sec id="sec-6">
      <title>Entity Detection</title>
      <p>For the topics FMT5 and FMT6, location and
organization information was required to be detected from the tweet.
To extract the location information, we used the geo-location
attribute from the tweets and the Stanford NER tagger6 to
extract location names from the tweet text. Similarly, we
used the Stanford NER tagger to detect organizations in
the tweet text.</p>
      <p>We split the tweet le into 10 les containing 5,000 tweets
each. The Stanford NER tagger was used in parallel on the
10 splitted les to identify the location and organization, if
any. This reduced the computation time by 85%.
4.4</p>
    </sec>
    <sec id="sec-7">
      <title>Word Vectors</title>
      <p>
        We used the pre-trained 200 dimensional GloVe [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] word
vectors on Twitter data7 (2 billion tweets) to create the
vectors of the preprocessed tweets and the word bags.
      </p>
      <p>The tweet vectors were created by taking the normalized
summation of the vectors of the words in the tweets, which
were present in the vocabulary of the pre-trained GloVe
model. In cases where the word was not a part of the model
vocabulary, it was assigned to the null vector.</p>
      <p>#
ti =</p>
      <p>1</p>
      <p>Nv (ti) j=1
# #
and, uij = 0 ;</p>
      <p>Nv(ti) #</p>
      <p>X uij
if uij 2= vocabulary
where,
#
ti = Tweet vector of ith tweet, ti.</p>
      <p>Nv (ti) = Number of words in ti present in vocabulary.
#
uij = Vector of ith word in jth tweet.</p>
      <p>Similarly, the word bag vectors were created by taking
the normalized summation of the vectors of the words in
the word bags, which were present in the vocabulary of the
pre-trained GloVe model. Out of vocabulary words were
assigned to the null vector.</p>
      <p>#
qi =</p>
      <p>1</p>
      <p>Nv (qi) j=1
# #
and, wij = 0 ;</p>
      <p>Nv(qi) #</p>
      <p>X wij
if wij 2= vocabulary
4http://www.nltk.org/howto/wordnet.html
5https://www.nodebox.net/code/index.php/Linguistics
6http://nlp.stanford.edu/software/CRF-NER.shtml
7http://nlp.stanford.edu/projects/glove/
#
qi = Topic vector of ith word bag, qi.</p>
      <p>Nv (qi) = Number of words in qi present in vocabulary.
#
wij = Vector of ith word in jth word bag.</p>
      <p># #
The tweet vector, ti and the word bag vector, qi are used
to calculate the similarity.</p>
      <p>
        The Word2Vec [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] library for Gensim8 was used to
create the tweet vectors and the topic vectors using the
pretrained GloVe model. The GloVe vectors were converted to
Word2Vec vectors using code from the GitHub repository,
manasRK/glove-gensim9.
4.5
      </p>
    </sec>
    <sec id="sec-8">
      <title>Similarity Metric</title>
      <p>We used cosine similarity measure to calculate the cosine
similarity, S between the tweet vector and the topic vector.
# #
S = cosine-sim(ti; qj )</p>
      <p># #
= #ti q#j</p>
      <p>jjtijj jjqj jj
A high v#alue of S denotes higher similarity between the tweet
#
vector, ti and the topic vector, qj and vice versa.</p>
      <p>For topics such as FMT5 and FMT6, where entity
information such as location (LOC) or organization (ORG) was
required, the consolidated score, S0 was calculated as
follows:
where,</p>
      <p>S0 =
I =</p>
      <p>S + I</p>
      <p>2
8
&lt;1, if LOC or ORG is present.
:0,</p>
      <p>otherwise.</p>
      <p>The consolidated value, S0 shifts the cosine similarity
towards 1 if the location or organization information is present
(high relevance) and towards 0, otherwise (low relevance).
5.</p>
    </sec>
    <sec id="sec-9">
      <title>RESULTS AND ERROR ANALYSIS</title>
      <p>JU NLP 1 0.4357
JU NLP 2 0.3714
JU NLP 3 0.3714
0.3420 0.0869
0.3004 0.0647
0.3004 0.0647
0.1125
0.0881
0.0881</p>
      <p>The secondary performance obtained in Run 2, 3 is a
result of the averaging which approximated the actual cosine
similarity value between the tweet and topic vectors. Runs
2 and 3, which are identical in nature, used cosine distance
as their similarity metric.</p>
      <p>8https://radimrehurek.com/gensim/models/word2vec.html
9https://github.com/manasRK/glove-gensim</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>In this paper, we presented a brief overview of our
system to address the information extraction from microblog
data. We have observed that, building word bags which
contained all the topic words relevant to the topic showed
better results than splitting the word bags. Therefore, Run
1 exhibited better results than the rest. Considering
hashtags as a feature should also improve the performance of the
system.</p>
      <p>As a future work, we work like to explore more
sophisticated techniques to build the vectors of the tweets, given the
vectors of its constituent words, by considering the sequence
of the words into account. We also plan to incorporate more
topic speci c features to improve the performance of our
system.
7.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Naskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bandyopadhyay</surname>
          </string-name>
          .
          <article-title>Entity extraction from social media using machine learning approaches</article-title>
          .
          <source>In Working Notes in Forum for Information Retrieval Evaluation (FIRE</source>
          <year>2015</year>
          ), pages
          <fpage>106</fpage>
          {
          <fpage>109</fpage>
          ,
          <string-name>
            <surname>Gandhinagar</surname>
          </string-name>
          , India,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          .
          <article-title>Overview of the FIRE 2016 Microblog track: Information Extraction from Microblogs Posted during Disasters</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Elbassuoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Meier</surname>
          </string-name>
          .
          <article-title>Practical extraction of disaster-relevant information from social media</article-title>
          .
          <source>In Proceedings of the 22Nd International Conference on World Wide Web, WWW '13 Companion</source>
          , pages
          <volume>1021</volume>
          {
          <fpage>1024</fpage>
          , New York, NY, USA,
          <year>2013</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Elbassuoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Meier</surname>
          </string-name>
          .
          <article-title>Extracting information nuggets from disaster-related messages in social media</article-title>
          .
          <source>Proc. of ISCRAM</source>
          , Baden-Baden, Germany,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In EMNLP</source>
          , volume
          <volume>14</volume>
          , pages
          <fpage>1532</fpage>
          {
          <fpage>43</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sriram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fuhry</surname>
          </string-name>
          , E. Demir,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ferhatosmanoglu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Demirbas</surname>
          </string-name>
          .
          <article-title>Short text classi cation in twitter to improve information ltering</article-title>
          .
          <source>In Proceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '10</source>
          , pages
          <fpage>841</fpage>
          {
          <fpage>842</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vieweg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Hughes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Starbird</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Palen</surname>
          </string-name>
          .
          <article-title>Microblogging during two natural hazards events: what twitter may contribute to situational awareness</article-title>
          .
          <source>In Proceedings of the SIGCHI conference on human factors in computing systems</source>
          , pages
          <volume>1079</volume>
          {
          <fpage>1088</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Ounis.</surname>
          </string-name>
          <article-title>Using word embeddings in twitter election classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1606.07006</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>