<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IIT BHU at FIRE 2018 IRMiDis Track - Obtaining Factual Tweets During Natural Disasters</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Harshit Mehrotra</string-name>
          <email>harshit.mehrotra.cse15@iitbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sukomal Pal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Indian Institute of Technology (BHU) Varanasi - 221005</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Introduction - Tasks and Data</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents details of the work done by the team of IIT (BHU) Varanasi for the IRMiDis track in FIRE 2018. The task involved classifying tweets posted during a disaster into those which are fact-checkable or factual and which are not, and also match these tweets to relevant news articles. Methodologies had to be developed in context of the 2015 Nepal Earthquake.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>microblogs</kwd>
        <kwd>disaster</kwd>
        <kwd>word embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Sub-Task 1</title>
      <sec id="sec-1-1">
        <title>Methodology</title>
        <p>The methodology for the rst sub-task i.e. identi cation of fact-checkable tweets
is fully automatic in both, query generation and searching. The key steps are
indicated as follows:
1. Pre-processing of all tweets by lower-casing, removal of stopwords,
hashtags and addressing and nally stemming using porter stemmer. The term
tweet hereafter refers to the pre-processed tweet.
2. Creating a TF-IDF based ranked list of terms in the reference set of 84
tweets. Only those terms are considered that occur in the reference set more
than once. We call this set of terms R with the TF-IDF score function being
T .
3. A word2vec word embedding model is trained on the entire set of 50,000
tweets.
4. Each test tweet is now attributed to its corresponding feature vector
that is formed by an arithmetic mean of the sum of the individual terms
embeddings.
5. To form the reference feature (V ) vector against which they will be matched,
we use the following weighted mean:</p>
        <p>V =</p>
        <p>PjiR=j1 T (Ri)E(Ri)</p>
        <p>
          PjiR=j1 T (Ri)
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Sub-Task 2</title>
      <p>The methodology for the second sub-task is manual in query generation and
automatic in searching, scoring, using the Java based text search library Lucene.
Details of the constituent steps are as follows:
1. The news articles are pre-processed in the same way as tweets are in
subtask 1.
2. The headline and rst 3 sentences of each news articles are combined.</p>
      <p>This creates one testing document for each news article.
3. Now each pre-processed tweet is used as a query to match with the testing
documents of the news articles. This done using Lucene and the score of
the best matching document is seen for each tweet.
4. If this score is more than 0.30, the corresponding news article is said to be
matching the tweet, otherwise no relevant news article is said to be found
for the tweet.
5. To nd the matching sentence, the tweet as a query is matched with each
sentence of the relevant news article. The sentence with the highest score is
returned as the answer.
3</p>
      <sec id="sec-2-1">
        <title>Results</title>
        <p>The results on the two sub-tasks, based on di erent metrics are indicated in
Table 1 and 2.
4</p>
      </sec>
      <sec id="sec-2-2">
        <title>Possible Improvements</title>
        <p>Depending on the kind of data, a sentiment analysis module can be augmented
in the classi cation pipeline. However since such system should be ready to use
for a disaster when it happens, the weight of such an additional module can
be found as a hyperparameter by studying data from such incidents that have
already occurred.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Basu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Overview of the FIRE 2018 track: Information Retrieval from Microblogs during Disasters (IRMiDis)</article-title>
          .
          <source>In: Proceedings of FIRE 2018 - Forum for Information Retrieval Evaluation (December</source>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>