<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information Extraction from Microblogs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Partha Pakray Computer Science and Engineering National Institute of Technology Mizoram</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Prashant Bhardwaj Computer Science and Engineering National Institute of Technology Agartala</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The micro blogging sites contain the emotion and expression of the public in raw format. The data can be used to extract much meaningful information that could be used to develop technologies for future use. There are numerous micro blogging sites available these days that are used in different contexts. Some are used basically for conversation, some for image and video sharing, and some for formal and official purposes. Twitter is one of the most outspoken platform for sharing the emotions and comments on almost every topic starting from sport to entertainment, religion to politics and many more. The paper attempts to extract information from the database of the tweets collected from Twitter. The task is to develop methodologies for extracting tweets that are relevant to each topic with high precision. This paper presents the nita_nitmz team participation in FIRE 2016 Microblog track.</p>
      </abstract>
      <kwd-group>
        <kwd>Information Retrieval</kwd>
        <kwd>Micro blogging</kwd>
        <kwd>Twitter</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Computing methodologies ~ Natural</title>
    </sec>
    <sec id="sec-2">
      <title>Processing</title>
    </sec>
    <sec id="sec-3">
      <title>Information systems ~ Information extraction</title>
      <sec id="sec-3-1">
        <title>1 INTRODUCTION</title>
        <p>
          The faster growth of Internet in current period provides new
sources of information. Now a day’s people prefer to express
themselves more often on social sites than any print media. The
idea of Information Extraction from Microblogs Posted during
Disasters was introduced by Sarah Vieweg et.al., in 2010
Proceedings of the SIGCHI Conference on Human Factors in
Computing Systems [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Henceforth it has become one of the
most researched topics considering the possibilities it contained
for the proper accessing of any incidence. The importance of the
said topic can be attributed to the fact that people rarely provide
any false information in the social sites and pour their emotions
according to their knowledge and wisdom. This paper presents the
experiments carried out at National Institute of Technology
Agartala as part of the participation in the Forum for Information
Retrieval Evaluation (FIRE) 2016 in Information Extraction from
Microblogs Posted during Disasters [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The experiments carried
out by us for FIRE 2016 are based on stemming, zonal indexing,
theme identification, TF-IDF based ranking model and positional
information. The data contained 48845 tweets out of 50,000
tweets mentioned in the workshop website. Query was provided
by the organizing committee and each query was specified using
title, narration and description format.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2 RELATED WORKS</title>
        <p>
          The problem of Information Extraction from Microblogs Posted
during Disasters is researched for a couple of years starting from
2010 by Sarah Vieweg et.al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and Leysia Palen et.al.[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. But
there has been tremendous work since then and a new field of
information retrieval has come into existence. Sudha Verma et.al.
wrote on Situational Awareness through tweets [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The research
on location of disaster hit area, the response and the information
extraction has been going on since then [4][
          <xref ref-type="bibr" rid="ref5">5</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ][
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. One of the
important part of the information retrieval part is the part of
speech tagging in the code mixed microblog
data[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ][
          <xref ref-type="bibr" rid="ref9">9</xref>
          ][
          <xref ref-type="bibr" rid="ref10">10</xref>
          ][
          <xref ref-type="bibr" rid="ref11">11</xref>
          ][
          <xref ref-type="bibr" rid="ref12">12</xref>
          ][
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Several researchers even work on
information extraction from mixed script analysis in social media
websites and forums. English used to dominate the micro
blogging sites previously such as Twitter and Facebook.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3 TASK DESCRIPTION</title>
        <p>A large set of microblogs (tweets) posted during a recent disaster
event was be made available, along with a set of topics (in TREC
format). Each ‘topic’ identified a broad information need during a
disaster, such as – what resources are needed by the population in
the disaster affected area, what resources are available, what
resources are required / available in which geographical region,
and so on. Specifically, each topic contained a title, a brief
description, and a more detailed narrative on what type of tweets
will be considered relevant to the topic. The participants are
required to develop methodologies for extracting tweets that are
relevant to each topic with high precision as well as high recall.</p>
        <sec id="sec-3-3-1">
          <title>The data contained:</title>
          <p>•
•</p>
          <p>Around 50,000 microblogs (tweets) from Twitter, those
were posted during the Nepal earthquake in April 2015.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Only the tweetids of the tweets was provided, along</title>
          <p>with a script that was be used to download the tweets
using the Twitter API. Out of 50,000 tweets only 48845
could be downloaded on my experimental setup.
A set of 5 – 8 topics in TREC format, each containing a
title, a brief description, and a more detailed narrative.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>4 METHODOLOGY</title>
        <p>For the given task we created the required searching configuration
on Apache Nutch 0.9 which is a highly extensible and scalable
open source web crawler software project. The implementation of
the said task is done in two steps. First is creating the searching
environment, Secondly apply appropriate test queries to search the
results from the previously configured Nutch using Tomcat
server.</p>
      </sec>
      <sec id="sec-3-5">
        <title>4.1 Preparation of the data</title>
        <p>The python script provided by the organizers helped to get the
tweets. But file generated was of json type. So another script had
to be developed to extract the tweets from the json file. The json
file contained the tweets, tweet_id and many more metadata. First
problem arises when after extracting tweets from the json file, we
found out that only 48845 tweets were downloaded via the given
script. As per the norms of the Apache Nutch, the tweets had to
separated into different files. Since the task was to extract relevant
tweets , the files containing the tweets had to be named according
to their tweet_ids. We developed two files one containing the
tweets only and other containing the tweet_ids. After that we
wrote a code to take the tweet_id from one file , create a text file
with that name , take a tweet from another file and store it in the
newly created file. But due to some unavoidable errors, only
48815 file could be created, each having the name as tweet_id and
containing the corresponding tweet inside.</p>
      </sec>
      <sec id="sec-3-6">
        <title>4.2 Crawling using Nutch</title>
        <p>We used another code to store the addresses of the different file in
the urls.txt file. We started the crawling part using the nutch. It
works in following steps:
Injecter: All the URLs are taken by the injector from the
seed.txt (here urls.txt) file, compares urls with
regexurlfiler regex and update crawldb with supported urls..
The crawldb maintains information on all known URLs
(fetch schedule, fetch status, metadata, …).</p>
        <p>Generator: Based on the data of crawldb, the generator selects
best-scoring urls due for fetch and the segments
directory is created.</p>
        <p>Fetcher and CrawlDb Update: Next, the fetcher fetches
the remote pages of the URLs on the fetch list and
updates it to the segment directory. This step takes a lot
of time.</p>
        <p>Parser: The contents of each web page are parsed. If the crawl
produces another extension to an already existing one,
the updater adds the new data to the crawldb.</p>
        <p>Inverting: The links need to be inverted before indexing. The
fact that the number of incoming links is more valuable
than the outgoing links is taken care off, similar to how
Google PageRank works. The inverted links are saved
in the linkdb.</p>
        <p>Indexing, Deduplicating and Merging: Using data from
crawldb, linkdb and segments, the indexer creates an
index and saves it. The Lucene library is used for</p>
        <p>Indexing.</p>
      </sec>
      <sec id="sec-3-7">
        <title>4.3 Searching for test Queries</title>
        <p>Now, we can search for tweets regarding the crawled database.
The searching of the test query takes place in the following steps:
Stop Word Removal: From the given query the stop words
have to be removed, because they do not contribute
much to the searching procedure.</p>
        <p>Query Segmentation: for a given set of words in a query, the
search engine may not give proper results, so it searched
for every combination of the words contained in the
search query.</p>
        <p>Merging: As discussed in the Query segmentation, the results
for a given query may not be available; so the results
obtained from the different combination of the words of
the query text is merged to get one of the probable
results of the search.
5</p>
      </sec>
      <sec id="sec-3-8">
        <title>RESULT AND CONCLUSION</title>
        <p>We submitted 37 results of the four query subtexts. The run
submission was accepted in the category of semi automatic run.
The results from the organizers after judging the submitted run is
provided in Table 1.
The results are not encouraging, but considering the fact that we
started from the scratch, we have much to learn. The different
participating teams have employed different algorithms to extract
the results. We would try to enhance our methodology for future
research.
6</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Vieweg</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>L. Hughes A.</given-names>
            ,
            <surname>Starbird</surname>
          </string-name>
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Plaen</surname>
          </string-name>
          <string-name>
            <surname>L.</surname>
          </string-name>
          ,
          <year>2010</year>
          .
          <article-title>Microblogging during two natural hazards events: what twitter may contribute to situational awareness</article-title>
          .
          <source>In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(Atlanta</source>
          ,
          <string-name>
            <surname>GA</surname>
          </string-name>
          , USA, April
          <volume>10</volume>
          -
          <issue>15</issue>
          ,
          <year>2010</year>
          ).
          <source>CHI '10. ACM</source>
          , Atlanta, GA,
          <fpage>1079</fpage>
          -
          <lpage>1088</lpage>
          . DOI=
          <volume>10</volume>
          .1145/1753326.1753486.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Palen</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M. Anderson</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mark</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sicker</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmer</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grunwald</surname>
            <given-names>D.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>A vision for technologymediated support for public participation &amp; assistance in mass emergencies &amp; disasters</article-title>
          .
          <source>In Proceedings of the 2010 ACM-BCS Visions of Computer Science Conference (Swindon, UK</source>
          ,
          <fpage>13</fpage>
          -
          <lpage>16</lpage>
          April
          <year>2010</year>
          ). ACM-BCS '
          <fpage>10</fpage>
          ,
          <string-name>
            <surname>Swindon</surname>
          </string-name>
          , UK.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Verma</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieweg</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>J. Corvey W.</surname>
          </string-name>
          , Plaen L.,
          <string-name>
            <given-names>H. Martin J.</given-names>
            ,
            <surname>Palmer</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Schram</surname>
          </string-name>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>M. Anderson</surname>
          </string-name>
          <string-name>
            <surname>K.</surname>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Natural Language Processing to the Rescue?: Extracting “Situational Awareness” Tweets During Mass Emergency</article-title>
          .
          <source>Association for the Advancement of Artificial Intelligence.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Watanabe</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ochi</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okabe</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Onai</surname>
            <given-names>R.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Jasmine: a real-time local-event detection system based on geolocation information propagated to microblogs</article-title>
          .
          <source>In Proceedings of the 20th ACM international conference on Information and knowledge management(Glasow</source>
          ,UK,
          <fpage>24</fpage>
          -
          <lpage>28</lpage>
          October
          <year>2011</year>
          ). CIKM '
          <volume>11</volume>
          ,
          <string-name>
            <surname>Glasow</surname>
          </string-name>
          ,UK,
          <fpage>2541</fpage>
          -
          <lpage>2544</lpage>
          , DOI=10.1145/2063576.2064014.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Lingad</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karimi</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            <given-names>J.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Location extraction from disaster-related microblogs</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on World Wide Web(Rio De Janerio, Brazil</source>
          ,
          <fpage>13</fpage>
          -17 May).
          <source>WWW '13 Companion</source>
          , Rio De Janerio, Brazil,
          <fpage>1017</fpage>
          -
          <lpage>1020</lpage>
          . DOI=
          <volume>10</volume>
          .1145/2487788.2488108
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Imran</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbassuoni</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diaz</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meier</surname>
            <given-names>P.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Practical extraction of disaster-relevant information from social media</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on World Wide Web(Rio De Janerio, Brazil</source>
          ,
          <fpage>13</fpage>
          - 17 May).
          <source>WWW '13 Companion</source>
          , Rio De Janerio, Brazil,
          <fpage>1021</fpage>
          -
          <lpage>1024</lpage>
          . DOI=
          <volume>10</volume>
          .1145/2487788.2488109
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Imran</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucas</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meier</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieweg</surname>
            <given-names>S.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>AIDR: artificial intelligence for disaster response</article-title>
          .
          <source>In Proceedings of the 23nd International Conference on World Wide Web(Seoul</source>
          , South Korea,
          <fpage>7</fpage>
          -11 April).
          <source>WWW '14 Companion</source>
          , Seoul, South Korea,
          <fpage>159</fpage>
          -
          <lpage>162</lpage>
          . DOI=
          <volume>10</volume>
          .1145/2567948.2577034
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Anupam</given-names>
            <surname>Jamatia</surname>
          </string-name>
          ,
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          .
          <article-title>Part-of-Speech Tagging System for Indian Social Media Text on Twitter</article-title>
          .
          <source>In Proceedings Workshop on Language Technologies For Indian Social Media(SOCIAL-INDIA)</source>
          , Pages
          <fpage>21</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Pinaki</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Partha Pakray</surname>
          </string-name>
          , Sivaji Bandyopadhyay,
          <article-title>Theme Based English and Bengali Ad-hoc Monolingual Information Retrieval in FIRE 2010</article-title>
          ,
          <string-name>
            <surname>In</surname>
            <given-names>FIRE</given-names>
          </string-name>
          2010, Working Notes. [2010]
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Barman</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wagner</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2014</year>
          <article-title>CodeMixing: A Challenge for Language Identification in the Language of Social Media</article-title>
          .
          <source>In The 1st Workshop on Computational Approaches</source>
          to Code Switching,
          <string-name>
            <surname>EMNLP</surname>
          </string-name>
          <year>2014</year>
          , October,
          <year>2014</year>
          , Doha, Qatar.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Björn</given-names>
            <surname>Gambäck</surname>
          </string-name>
          , and
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Comparing the Level of CodeSwitching in Corpora</article-title>
          .
          <source>In the 10th edition of the Language Resources and Evaluation Conference (LREC)</source>
          ,
          <source>2328 May</source>
          <year>2016</year>
          ,
          <string-name>
            <surname>Portorož</surname>
          </string-name>
          (Slovenia).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Anupam</surname>
            <given-names>Jamatia</given-names>
          </string-name>
          , Björn Gambäck, and
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Collecting and Annotating Indian Social Media CodeMixed Corpora</article-title>
          .
          <source>In the 17th International Conference on Intelligent Text Processing and Computational Linguistics (CICLING), April 3-9</source>
          , Konya, Turkey.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Kunal</given-names>
            <surname>Chakma</surname>
          </string-name>
          , and
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          .
          <article-title>CMIR:A Corpus for Evaluation of Code Mixed Information Retrieval of HindiEnglish Tweets</article-title>
          .
          <year>2016</year>
          .
          <source>In the 17th International Conference on Intelligent Text Processing and Computational Linguistics (CICLING), April 3-9</source>
          , Konya, Turkey.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          .
          <article-title>Overview of the FIRE 2016 Microblog track: Information Extraction from Microblogs Posted during Disasters</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>