<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semi-automatic keyword based approach for FIRE 2016 Microblog Track</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Supervised classification; Information extraction; Terrier; Twitter.</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ganchimeg Lkhagvasuren Évora University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>José Saias Évora University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Teresa Gonçalves Évora University</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Computing methodologies Support Vector • Information systems Information Extraction</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>This paper describes our semi-automatic keyword based approach for the four topics of Information Extraction from Microblogs Posted during Disasters task at Forum for Information Retrieval Evaluation (FIRE) 2016. The approach consists three phases;</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Machine</p>
    </sec>
    <sec id="sec-2">
      <title>1 INTRODUCTION</title>
      <p>
        It is undeniable that microblogging sites have become key
resources of significant information during disaster event [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. One
of these microblogging site, Twitter, is a social networking
website which enables users to generate 140-character messages
named “tweets” everyday. A giant number of tweets is posted
including informative and non-informative messages, which
makes opportunities for information extraction [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        However, dealing with tweets and identifying specific keywords
are challenging work due to the nature of Twitter. The small,
noisy and fragmented tweets mean they have very simple
discourse and pragmatic structure, issues which still challenge
state-of-the-art NLP systems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Task description: The aim is to retrieve a number tweets relevant
to each topic provided with high precision as well as high recall.
The titles of topics are provided in TREC format as the following:</p>
      <sec id="sec-2-1">
        <title>What resources were available</title>
      </sec>
      <sec id="sec-2-2">
        <title>What resource were required</title>
      </sec>
      <sec id="sec-2-3">
        <title>What medical resource were available</title>
      </sec>
      <sec id="sec-2-4">
        <title>What medical resource were required What were the requirements or availability of resources at specific location</title>
      </sec>
      <sec id="sec-2-5">
        <title>What were the activities of various government organizations</title>
      </sec>
      <sec id="sec-2-6">
        <title>NGOs or</title>
        <p>What infrastructure damage or restoration were reported
Dataset: Approximately 50,000 tweets that posted during Nepal
earthquake disaster were given in JSON format. A main feature
of the task is that a gold standard dataset was not provided.
In terms of our approach, we propose to achieve the first four
topics using keywords extraction with manual work and
classification methods.</p>
        <p>
          This paper is organized as follows. First, the components of
approach are described separately. Then, the result analysis and
conclusion are presented. Our work is submitted in FIRE 2016
Microblog track [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2 KEYWORD BASED APPROACH</title>
      <p>Our approach for the Microblog track comprises three phases,
Keyword extraction, Retrieval, and Classification (see Figure 1).
In the first phase, we extracted all relief resources (keywords) that
were available or required. Using those keywords and Terrier1
search engine, we retrieved a number of tweets that each tweet
includes a keyword at least in the middle phase. In the last phase,
the retrieved tweets are classified into the first and second topics
using Support Vector Machine (SVM).</p>
    </sec>
    <sec id="sec-4">
      <title>2.1 Extracting keywords</title>
      <p>
        In order to extract keywords, we used separately the following
two methods with manual work. The keywords that we first
extracted is provided in the topic descriptions, such as food,
water, volunteer, money, medicine and transportation. The
quantitative results explored in this phase is presented in Table 1.
Since tweets are usually written in an informal style, the most of
NLP tools show poor performance on Twitter datasets. So we
tried to exploit specific NLP tools which are [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Word embedding: Based on these relief resources we mentioned
before, we attempted to obtain more keywords from the given
dataset. In order to do that, First of all, we tagged all tweets by
GATE twitter Part-of-speech tagger [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. After distinguished all
nouns, each noun is represented by a Word2Vec model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that
was trained particularly on Twitter datasets to deal with noisy
tweets. Then 50 nearest neighbor nouns of each of the keywords
extracted from the descriptions are found as candidates. From
these candidates which are more likely be relief resources during
the earthquake, we labeled 86 nouns as keywords manually.
However, it was clear that there are more keywords we could not
extract, such as Nepali words.
      </p>
      <p>
        Chunking and Wordnet: One of the basic technique for
information extraction, chunking, is used to identify keywords in
our approach. We defined some chunk grammar, for example,
“CHUNK: {&lt;NNP.?&gt;*&lt;VB.?&gt;+&lt;DT&gt;?&lt;JJ&gt;*&lt;NN|NNS&gt;+}”
based on tagging by POS in the previous step. Next, the nouns
were filtered by Wordnet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and specific verbs such as distribute,
give, provide, support and hand. Then we enriched the keyword
list from filtered nouns manually.
      </p>
    </sec>
    <sec id="sec-5">
      <title>2.2 Retrieval</title>
      <p>Once we had a bunch of keywords extracted in the previous
phase, we retrieved all tweets (around 8620 tweets) that include at
least one keyword using those keywords on the Terrier. There are
few open search engines however, we chose Terrier taking some
of its advantages into consideration. In term of scoring model, we
employed BM25 which is based on probabilistic retrieval
framework. The rank and scores are used to compute the
relevance of a tweet to a topic in further.</p>
    </sec>
    <sec id="sec-6">
      <title>2.3 Classifying into topics</title>
      <p>The most of tweets that retrieved in previous phase can be
significantly related to the first two topics while some of them
cannot. For instance, Even though the following two tweets both
includes water (a keyword), the former one is related to the first
topic what resources were available, whereas the latter tweet is
not related to any topic.</p>
      <p>Anyone in need of drinking water contact me. Have some can
donate #earthquake #Nepal #bhaktapur
#ShameOnYou #nepalgov Rs 20 water cost Rs 40
#earthquakenepal #earthquake #Nepal #fuckoff don't #donate
#unknown #website
Therefore we classified the tweets into three classes – available,
required and other. In order to do that, first, we annotated 1000
tweets manually. In preprocessing of classification, all URL,
usertag and some symbols were removed. Then we employed
three classifiers with basic features such as unigram, bag-of-words
and some twitter specific features on WEKA2 open source
machine learning software. The best result was executed by SVM
(see Table 2).</p>
      <p>In term of the third and fourth topic, “What medical resources
were available” and “What medical resources were required”, we
retrieved the relevant tweets from the tweets of the first topic and
second topic using medical relief resources respectively.
It is impossible to compare our results to other participants results
because we submitted the attempts of only three topics to the
organizers. However, the results estimated by the organizers was
reasonable, which brought us encourage to complete our work.
The result is presented in Table 3.
0,8500
0,4988
0,2204</p>
      <p>Overall MAP
0,2420
4</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>In this paper, we have presented our keyword based approach for
the four topics of FIRE2016 Microblog Track. Our system is
semi-automatic, which includes manual work in the keyword
extraction phase. Moreover, the phases are not integrated with
each other.</p>
      <p>Next, we plan to improve our system to become automatic and to
use advanced methods.
5</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Elbassuoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Meier</surname>
          </string-name>
          .
          <article-title>Extracting Information Nuggets from Disaster Related Messages in Social Media</article-title>
          .
          <source>In: Proceeding of the 10th International ISCRAM Conference</source>
          ,
          <year>2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ritter</surname>
          </string-name>
          , Mausem and
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <article-title>Open domain event extraction from Twitter</article-title>
          .
          <source>KDD'</source>
          12
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Piskorski</surname>
          </string-name>
          , R.Yangerber Information Extraction:
          <article-title>Past, Present and Future</article-title>
          . In: Multi source,
          <source>Multilingual Information Extraction and Summarisation</source>
          . 2013
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Clarke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>"Twitter Part-of-Speech Tagging for All: Overcoming Sparse and Noisy Data"</article-title>
          .
          <source>In Proceedings of the International Conference on Recent Advances in Natural Language Processing</source>
          , ACL.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Godin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandersmissen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Neve</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , &amp; Van de Walle,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Multimedia Lab @ ACL W-NUT NER shared task: Named entity recognition for Twitter microposts</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>George</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <string-name>
            <surname>Miller</surname>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>WordNet: A Lexical Database for English</article-title>
          .
          <source>Communications of the ACM</source>
          Vol.
          <volume>38</volume>
          , No.
          <volume>11</volume>
          :
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          .
          <article-title>Overview of the FIRE 2016 Microblog track: Information Extraction from Microblogs Posted during Disasters</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>