<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the FIRE 2016 Microblog track: Information Extraction from Microblogs Posted during Disasters</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Saptarshi Ghosh</string-name>
          <email>sghosh@cs.iiests.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kripabandhu Ghosh</string-name>
          <email>kripa.ghosh@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of CST</institution>
          ,
          <addr-line>IIEST Shibpur</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Indian Statistical Institute</institution>
          ,
          <addr-line>Kolkata</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p>The FIRE 2016 Microblog track focused on retrieval of microblogs (tweets posted on Twitter) during disaster events. A collection of about 50,000 microblogs posted during a recent disaster event was made available to the participants, along with a set of seven practical information needs during a disaster situation. The task was to retrieve microblogs relevant to these needs. 10 teams participated in the task, submitting a total of 15 runs. The task resulted in comparison among performances of various microblog retrieval strategies over a benchmark collection, and brought out the challenges in microblog retrieval.</p>
      </abstract>
      <kwd-group>
        <kwd>Information systems ! Query reformulation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Microblogging sites such as Twitter (https://twitter.com)
have become important sources of situational information
during disaster events, such as earthquakes, oods, and
hurricanes [
        <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
        ]. On such sites, a lot of content is posted
during disaster events (in the order of thousands to
millions of tweets), and the important situational information
is usually immersed in large amounts of general
conversational content, e.g., sympathy for the victims of the disaster.
Hence, automated IR techniques are needed to retrieve
speci c types of situational information from the large amount
of text.
      </p>
      <p>
        There have been few prior attempts to develop IR
techniques over microblogs posted during disasters, but there has
been little effort till now to develop a benchmark dataset /
test collection using which various microblog retrieval
methodologies can be compared and evaluated. The objectives of
the FIRE 2016 Microblog track are two-fold { (i) to develop
a test collection of microblogs posted during a disaster
situation, which can serve as a benchmark dataset for evaluation
of microblog retrieval methodologies, and (ii) to evaluate and
compare the performance of various IR methodologies over
the test collection. The track is inspired by the TREC
Microblog Track [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which aims to evaluate microblog retrieval
strategies in general. In contrast, the FIRE 2016 Microblog
Track focuses on microblog retrieval in a disaster situation.
      </p>
      <p>In this track, a collection of about 50,000 microblogs posted
during a recent disaster event was made available to the
participants, along with a set of seven practical information
needs that are faced in a disaster situation by the agencies
responding to the disaster. Details of the collection are
discussed in Section 2. The task was to retrieve microblogs
relevant to the information needs (see Section 3. 10 teams
participated in the track, submitting a total of 15 runs that
are described in Section 4). The runs were evaluated against
a gold standard developed by human assessors, using
standard measures like Precision, Recall, and MAP.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>THE TEST COLLECTION</title>
      <p>
        In this section, we describe how the test collection for the
Microblog track was developed. Following the Cran eld
style [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], we describe the creation of topics (information
needs), document set (here, microblogs or tweets)
collection and relevance assessment to prepare the gold standard
necessary for evaluation of IR methodologies.
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Topics for retrieval</title>
      <p>In this track, our objective was to develop a test
collection to evaluate IR methodologies for extracting
information (from microblogs) that can potentially help responding
agencies to respond to a disaster situation such as an
earthquake or a ood. To this end, we consulted members of
some NGOs who regularly work in disaster-affected regions
{ such as, Doctors For You (http://doctorsforyou.org/) and
SPADE (http://www.spadeindia.org/) { to know what are
the typical information requirements during a disaster
relief operation. They identi ed certain information needs
such as what resources are required / available (especially
medical resources), what infrastructure damages are being
reported, the situation at speci c geographical locations, the
ongoing activities of various NGOs and government
agencies (so that the operations of various responding agencies
can be coordinated), and so on. Based on their feedback,
we identi ed seven topics on which information needs to be
retrieved during a disaster.</p>
      <p>Table 1 states the seven topics which we have developed
as a part of the test collection. These topics are written
in the format conventionally used for TREC topics.1 Each
topic contains an identifying number (num), a textual
representation of the information need (title), a brief description
(desc) of the same and a more detailed narrative (narr)
explaining what type of documents (tweets) will be considered
relevant to the topic, and what type of tweets would not be
considered relevant.</p>
      <sec id="sec-3-1">
        <title>1trec.nist.gov/pubs/trec6/papers/overview.ps.gz</title>
        <p>
          &lt;num&gt; Number: FMT1
&lt;title&gt; What resources were available
&lt;desc&gt; Identify the messages which describe the availability of some resources.
&lt;narr&gt; A relevant message must mention the availability of some resource like food, drinking water, shelter, clothes, blankets,
human resources like volunteers, resources to build or support infrastructure, like tents, water lter, power supply and so on.
Messages informing the availability of transport vehicles for assisting the resource distribution process would also be relevant.
However, generalized statements without reference to any resource or messages asking for donation of money would not be relevant.
&lt;num&gt; Number: FMT2
&lt;title&gt; What resources were required
&lt;desc&gt; Identify the messages which describe the requirement or need of some resources.
&lt;narr&gt; A relevant message must mention the requirement / need of some resource like food, water, shelter, clothes, blankets,
human resources like volunteers, resources to build or support infrastructure like tents, water lter, power supply, and so on. A
message informing the requirement of transport vehicles assisting resource distribution process would also be relevant. However,
generalized statements without reference to any particular resource, or messages asking for donation of money would not be relevant.
&lt;num&gt; Number: FMT3
&lt;title&gt; What medical resources were available
&lt;desc&gt; Identify the messages which give some information about availability of medicines and other medical resources.
&lt;narr&gt; A relevant message must mention the availability of some medical resource like medicines, medical equipments, blood,
supplementary food items (e.g., milk for infants), human resources like doctors/staff and resources to build or support medical
infrastructure like tents, water lter, power supply, ambulance, etc. Generalized statements without reference to medical resources
would not be relevant.
&lt;num&gt; Number: FMT4
&lt;title&gt; What medical resources were required
&lt;desc&gt; Identify the messages which describe the requirement of some medicine or other medical resources.
&lt;narr&gt; A relevant message must mention the requirement of some medical resource like medicines, medical equipments,
supplementary food items, blood, human resources like doctors/staff and resources to build or support medical infrastructure like tents,
water lter, power supply, ambulance, etc. Generalized statements without reference to medical resources would not be relevant.
&lt;num&gt; Number: FMT5
&lt;title&gt; What were the requirements / availability of resources at speci c locations
&lt;desc&gt; Identify the messages which describe the requirement or availability of resources at some particular geographical location.
&lt;narr&gt; A relevant message must mention both the requirement or availability of some resource, (e.g., human resources like
volunteers/medical staff, food, water, shelter, medical resources, tents, power supply) as well as a particular geographical location.
Messages containing only the requirement / availability of some resource, without mentioning a geographical location would not
be relevant.
&lt;num&gt; Number: FMT6
&lt;title&gt; What were the activities of various NGOs / Government organizations
&lt;desc&gt; Identify the messages which describe on-ground activities of different NGOs and Government organizations.
&lt;narr&gt; A relevant message must contain information about relief-related activities of different NGOs and Government organizations
in rescue and relief operation. Messages that contain information about the volunteers visiting different geographical locations would
also be relevant. However, messages that do not contain the name of any NGO / Government organization would not be relevant.
&lt;num&gt; Number: FMT7
&lt;title&gt; What infrastructure damage and restoration were being reported
&lt;desc&gt; Identify the messages which contain information related to infrastructure damage or restoration.
&lt;narr&gt; A relevant message must mention the damage or restoration of some speci c infrastructure resources, such as structures
(e.g., dams, houses, mobile tower), communication infrastructure (e.g., roads, runways, railway), electricity, mobile or Internet
connectivity, etc. Generalized statements without reference to infrastructure resources would not be relevant.
We collected a large set of tweets related to the devastating
earthquake that occurred in Nepal and parts of India on
25th April 2015.2 We collected tweets using the Twitter
Search API [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], using the keyword `nepal', that were posted
during the two weeks following the earthquake. We collected
only tweets in English (based on language identi cation by
Twitter itself), and collected about 100K tweets in total.
        </p>
        <p>
          Tweets often contain duplicates and near-duplicates since
the same information is frequently retweeted / re-posted by
multiple users [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. However, duplicates are not desirable
in a test collection for IR, since the presence of duplicates
can result in over-estimation of the performance of an IR
methodology. Additionally, the presence of duplicate
documents also creates information overload for human
annotators while developing the gold standard [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Hence, we
removed duplicate and near-duplicate tweets using a
simplied version of the methodologies discussed in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], as follows.
        </p>
        <p>Each tweet was considered as a bag of words (excluding
2https://en.wikipedia.org/wiki/April 2015 Nepal
earthquake
standard English stopwords and URLs), and the similarity
between two tweets was measured as the Jaccard
similarity between the two corresponding bags (sets) of words. If
the Jaccard similarity between two tweets was found to be
higher than a threshold value (0:7), the two tweets were
considered near-duplicates, and only the longer tweet
(potentially more informative) was retained in the collection.</p>
        <p>After removing duplicates and near-duplicates, we obtained
a set of 50,068 tweets, which was used as the test collection
for the track.
2.3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Developing gold standard for retrieval</title>
      <p>
        Evaluation of any IR methodology requires a gold standard
containing the documents that are actually relevant to the
topics. As is the standard procedure, we used human
annotators to develop this gold standard. A set of three human
annotators were used, each of whom is pro cient in English
and is a regular user of Twitter, and has prior experience of
working with social media content posted during disasters.
The development of gold standard involved three phases.
Phase 1: Each annotator was given the set of 50,068 tweets,
and the seven topics (in TREC format, as stated in Table 1).
Each annotator was asked to identify all tweets relevant to
each topic, independently, i.e., without consulting the other
annotators. To help the annotators, the tweets were indexed
using the Indri IR system [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which helped the annotators to
search for tweets containing speci c terms. For each topic,
the annotators were asked to think of appropriate
searchterms, retrieve tweets containing those search terms (using
Indri), and to judge the relevance of the retrieved tweets.
      </p>
      <p>After the rst phase, we observed that the set of tweets
identi ed to be relevant to the same topic by different
annotators, was considerably different. This difference was
because different annotators used different search-terms to
retrieve tweets.3 Hence, we conducted a second phase.
Phase 2: In this phase, for a particular topic, all tweets
that were judged relevant by at least one annotator (in the
rst phase) were considered. The decision whether a tweet
is relevant to a topic was nalised through discussion among
all the annotators and mutual agreement.</p>
      <p>
        Phase 3: The third phase used standard pooling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] (as
commonly done in TREC tracks) { the top 30 results of all
the submitted runs were pooled (separately for each topic),
and judged by the annotators. In this phase, all annotators
were judging a common set of tweets, hence inter-annotator
agreement could be measured. There was agreement among
all annotators for over 90% of the tweets; for the rest, the
relevance was decided through discussion among all the
annotators and mutual agreement.
      </p>
      <p>The nal gold standard contains the following number of
tweets judged relevant to the seven topics { FMT1: 589,
FMT2: 301, FMT3: 334, FMT4: 112, FMT5: 189, FMT6:
378, FMT7: 254.
2.4</p>
    </sec>
    <sec id="sec-5">
      <title>Insights from the gold standard development process</title>
      <p>Through the process described above, we understood that
for any of the topics, there are several tweets which are
definitely relevant to the topic, but which were difficult to
retrieve even for human annotators. This is evident from the
fact that, many of the relevant tweets could initially be
retrieved by only one out of the three annotators (in the rst
phase), but when the tweets were shown to the other
annotators (in the second phase), they unanimously agreed that
the tweet was relevant. These observations highlight the
challenges in microblog retrieval.</p>
      <p>
        Note that our approach for developing the gold standard
is different from that used in TREC tracks, where the gold
standard is usually developed by pooling few top-ranked
documents retrieved by different submitted systems, and
then annotating these top-ranked documents [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In other
words, only the third phase (as described above) is applied
in TREC tracks.
      </p>
      <p>Given that it is challenging to identify many of the tweets
relevant to a topic (as discussed above), annotating only a
relatively small pool of documents retrieved by IR
methodologies has the potential risk of missing many of the relevant
documents which are more difficult to retrieve. We believe
3Since the different annotators retrieved and judged very
different sets of tweets, it is not meaningful to report
interannotator agreement in this case.
that our approach, where the annotators viewed the entire
dataset instead of a relatively small pool, is likely to be more
robust, and is expected to have resulted in development of
a more complete gold standard which is irrespective of the
performance of any IR methodology.
3.</p>
    </sec>
    <sec id="sec-6">
      <title>DESCRIPTION OF THE TASK</title>
      <p>The participants were given the tweet collection and the
seven topics described earlier. It can be noted that the
Twitter terms and conditions prohibit direct public sharing of
tweets. Hence, only the tweet-ids4 of the tweets were
distributed among the participants, along with a Python script
using which the tweets can be downloaded via the Twitter
API.</p>
      <p>The participants were invited to develop IR methodologies
for retrieving tweets relevant to the seven topics. The
participants were asked to submit a ranked list of tweets that they
judge relevant to each topic. The ranked list was evaluated
based on the gold standard (developed as described earlier)
using the following measures: (i) Precision at 20 (Prec@20),
i.e., what fraction of the top-ranked 20 results are actually
relevant according to the gold standard, (ii) Recall at 1000
(Recall@1000), i.e., what fraction of all tweets relevant to a
topic (as identi ed in the gold standard) is present among
the top-ranked 1000 results, (iii) Mean Average Precision at
1000 (MAP@1000), and (iv) Overall MAP considering the
full retrieved ranked list. Out of these, we only report the
Prec@20 and MAP measures (in the next section).</p>
      <p>The track invited three types of methodologies { (i)
Automatic, where both query formulation and retrieval are
automated, and (ii) Semi-automatic, where manual intervention
is involved in the query formulation stage (but not in the
retrieval stage), and (iii) Manual, where manual intervention
is involved in both query formulation and retrieval stages.</p>
      <p>15 runs were submitted by the participants, out of which,
one run was fully automatic, while the others were
semiautomatic. The methodologies are summarized and
compared in the next section.
4.</p>
    </sec>
    <sec id="sec-7">
      <title>METHODOLOGIES</title>
      <p>Ten teams participated in the FIRE 2016 Microblog track.
A summary of the methodologies used by each team is given
in the next sub-section. Table 2 shows the evaluation
performance of each submitted run, along with a brief summary.
For each type, the runs are arranged in the decreasing order
of the primary measure, i.e., Precision@20. In case of a tie,
the arrangement is done in the decreasing order of MAP.
4.1</p>
    </sec>
    <sec id="sec-8">
      <title>Method summary</title>
      <p>We now summarize the methodologies adopted in the
submitted runs.</p>
      <p>dcu fmt16: This team participated from ADAPT
Centre, School of Computing, Dublin City University,
Ireland. It used WordNet5 to perform synonym-based
query expansion and submitted the following two runs:
1. dcu fmt16 1: This is an Automatic run (i.e. no
manual step involved). First, the words in
&lt;title&gt; and &lt;narr&gt; were considered, from which the
4Twitter assigns a unique numeric id to each tweet, called
the tweet-id.
5https://wordnet.princeton.edu/
iiest saptarashmi bandyopadhyay 1</p>
      <p>Run Id</p>
      <sec id="sec-8-1">
        <title>WordNet, Query Expansion</title>
      </sec>
      <sec id="sec-8-2">
        <title>Correlation, NER, Word2Vec WordNet, Query Expansion, NER, GloVe</title>
        <p>WordNet, Query Expansion,</p>
        <p>Relevance Feedback
WordNet, Query Expansion,
NER, GloVe, word bags split
WordNet, Query Expansion,
NER, GloVe, word bags split</p>
        <p>Lucene default model</p>
        <p>Relevancer system, Clustering
Manual labelling, Naive Bayes classi cation
Word2vec, Query Expansion,</p>
        <p>equal term weight</p>
        <p>Word2vec, Query Expansion,
unequal term weights, WordNet</p>
        <p>Word-overlap, POS tagging</p>
        <p>WordNet, wup score, POS tagging
Apache Nutch 0.9, query segmentation,
result merging</p>
      </sec>
      <sec id="sec-8-3">
        <title>Entity and action verbs relationships,</title>
        <p>Temporal Importance
Keyword extraction, Part-of-speech tagger,</p>
        <p>Word2Vec, WordNet, Terrier, Retrieval,</p>
        <p>
          Classi cation, SVM
stopwords were removed. Thus the initial query
was formed. Then, for each word in the query,
the synonyms were added using WordNet,
resulting in the expanded query. Retrieval was done
from this expanded query using the BM25 model
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
2. dcu fmt16 2: This is a Semi-automatic run (i.e.
manual step was involved). First an initial ranked
list was generated using the original topic. From
the top 30 tweets, 1-2 relevant tweets were
manually identi ed and query expansion was done from
these relevant tweets. The expanded query was
further expanded using WordNet just as done for
dcu fmt16 1. This nal expanded query was used
for retrieval.
iiest saptarashmi bandyopadhyay: This team
participated from Indian Institute of Engineering Science
and Technology, Shibpur, India. It submitted one
Semiautomatic run described below:
{ iiest saptarashmi bandyopadhyay 1: Correlation
between the topic words and the tweet was
calculated and this value determined the relevance
score for a given topic-tweet pair. The Stanford
NER tagger6 was used to identify the
LOCATION, ORGANIZATION and PERSON names in
the tweets. For each topic, some keywords were
manually selected on which a number of tools
(e.g., PyDictionary, NodeBox toolkit etc.) were
used to nd the corresponding synonyms, in
ectional variants etc. The bag of words for each
topic was further converted into a vector using
Word2Vec package.7 Finally, the relevance score
was calculated from the correlation between the
vector representations of the topic word bags and
the tweet text.
        </p>
        <p>
          JU NLP: This team participated from Jadavpur
University, India. It submitted three Semi-automatic runs
described as below:
1. JU NLP 1: This run was generated by using word
embeddings. For each topic, relevant words were
manually chosen and expanded using the synonyms
obtained from NLTK WordNet toolkit. In
addition, past, past participle and present
continuous forms of verbs were obtained using the
NodeBox library for Python. For the topics FMT5
and FMT6, location and organization
information was extracted using Stanford NER tagger.
GloVe[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] model was trained on the twitter
collection. A tweet vector, as well as, a query vector
was formed by taking the normalized summation
of the vector (obtained from GloVe) of the
constituent words. Then for each query-tweet pair,
6nlp.stanford.edu/software/Stanford-ner-2015-04-20.zip
7https://deeplearning4j.org/word2vec
the similarity score was calculated by the cosine
similarity of the corresponding vectors.
2. JU NLP 2: This run is similar to JU NLP 1
except that here word bags were split categorically
and average similarity between the tweet vector
and the split topic vectors was calculated.
        </p>
        <p>3. JU NLP 3: This is identical to JU NLP 2.
iitbhu fmt16: This team participated from
Department of Computer Science and Engineering, Indian
Institute of Technology (BHU) Varanasi, India. It
submitted one Semi-automatic run { iitbhu fmt16 1
described as follows:
{ iitbhu fmt16 1: The Lucene8 default similarity model,
which combines Vector Space Model (VSM) and
probabilistic models (e.g., BM25), was used to
generate the run. StandardAnalyzer, which
handled names and email address and lowercased each
token, and removed stopwords and punctuations,
was used. The query formulation stage involved
manual intervention.
daiict irlab: This team participated from DAIICT,
Gandhinagar, India and LDRP, Gandhinagar, India.
It submitted two Semi-automatic runs described as
follows:
1. daiict irlab 1: This run was generated using query
expansion, where the 5 similar words and
hashtags from the Word2vec model, trained on the
tweet corpus, were added to the original query.</p>
        <p>Equal weight was assigned to each term.
2. daiict irlab 2: This run was generated in the same
way as daiict irlab 1 except that different weights
were assigned to the expanded terms than the
original terms. More weights were assigned to
the words like required and available. These terms
were also expanded using WordNet.
trish iiest: This team participated from Indian
Institute of Engineering Science and Technology, Shibpur,
India. It submitted two Semi-automatic runs described
below:
1. trish iiest ss: The similarity score between a query
and a tweet is the word-overlap between them,
normalized by the query length. In each topic, the
nouns, identi ed by the Stanford Part-Of-Speech
Tagger, were selected to form the query. In
addition, more weight is assigned on words like
availability or requirement.
2. trish iiest ws: For this run, wup9 score is
calculated on the synsets of each term obtained from
WordNet.
nita nitmz: This team participated from National
Institute of Technology, Agartala, India and National
Institute of Technology, Mizoram. It submitted one
Semi-supervised run described as below:</p>
        <sec id="sec-8-3-1">
          <title>8https://lucene.apache.org/(2016,August20)</title>
          <p>9http://search.cpan.org/dist/WordNet-Similarity/lib/
WordNet/Similarity/wup.pm
{ nita nitmz 1: This run was generated on Apache
Nutch 0.9. Search was done using the different
combination of words present in the query. The
results obtained from different combinations of
query were merged.</p>
          <p>Helpingtech: This team participated from Indian
Institute of Technology, Patna, Bihar, India and
submitted the following Semi-automatic run (on 5 topics
only):
{ Helpingtech 1: For each query, relationships
entities and action verbs were de ned through manual
inspection. The ranking score was calculated on
the basis of the presence of these pre-de ned
relationships in the tweet for a given query. More
importance was given to a tweet which indicated
immediate action than a one which indicated a
proposed action for future.</p>
          <p>GANJI: This team participated from Evora
University, Portugal. It submitted three retrieval results (GANJI 1,
GANJI 2, GANJI 3) for the rst three topics only
using Semi-automatic methodology, described below:
{ GANJI 1, GANJI 2, GANJI 3 (combined): First,
keyword extraction was done using Part-of-speech
tagger, Word2Vec (to obtain the nouns) and
WordNet (to obtain the verbs). Then, retrieval was
performed on Terrier10 using the BM25 model.
Finally, SVM classi er was used to classify the
retrieved tweets into available, required and other
classes.
relevancer ru nl: This team participated from
Radboud University, the Netherlands and submitted the
following Semi-automatic run:
{ relevancer ru nl: This run was produced by a tool
Relevancer. After a pre-processing step, the tweet
collection was clustered to identify coherent
clusters. Each such cluster was manually labelled by
some experts as relevant or non-relevant. This
training data was used for Naive Bayes based
classi cation. For each topic, the test tweets
predicted as relevant by the classi er were
submitted.
5.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION AND FUTURE DIRECTIONS</title>
      <p>The FIRE 2016 Microblog track successfully created a
benchmark collection of microblogs posted during disaster events,
and compared the performance of various IR methodologies
over the collection.</p>
      <p>In subsequent years, we hope to conduct extended versions
of the Microblog track, where the following extensions can
be considered:</p>
      <p>Instead of just considering binary relevance (where a
tweet is either relevant to a topic or not), graded
relevance can be considered, e.g., based on factors like how
important or actionable the information contained in
the tweet is, how useful the tweet is likely to be to the
agencies responding to the disaster, and so on.
10http://terrier.org
The challenge in this year's track considered a static
set of microblog. But in reality, microblogs are
obtained in a continuous stream. The challenge can be
extended to retrieve relevant microblogs dynamically,
e.g., as and when they are posted.</p>
      <p>It can be noted that even the best performing method
submitted in the track achieved a relatively low MAP score
of 0:24 (considering only three topics), which highlights the
difficulty and challenges in microblog retrieval during a
disaster situation. We hope that the test collection developed
in this track will help development of better models for
microblog retrieval in future.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgements</title>
      <p>The track organizers thank all the participants for their
interest in this track. We also acknowledge our assessors,
notably Moumita Basu and Somenath Das, for their help in
developing the gold standard for the test collection. We also
thank the FIRE 2016 organizers for their support in
organizing the track.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cleverdon</surname>
          </string-name>
          .
          <article-title>The cran eld tests on index language devices</article-title>
          . In K. Sparck Jones and P. Willett, editors,
          <source>Readings in Information Retrieval</source>
          , pages
          <volume>47</volume>
          {
          <fpage>59</fpage>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Vieweg</surname>
          </string-name>
          .
          <source>Processing Social Media Messages in Mass Emergency: A Survey. ACM Computing Surveys</source>
          ,
          <volume>47</volume>
          (
          <issue>4</issue>
          ):
          <volume>67</volume>
          :1{
          <fpage>67</fpage>
          :
          <fpage>38</fpage>
          ,
          <year>June 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Efron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sherman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          .
          <article-title>Overview of the TREC-2015 Microblog Track</article-title>
          . Available at: https://cs.uwaterloo.ca/ ~jimmylin/publications/Lin etal TREC2015.pdf,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>I.</given-names>
            <surname>Ounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Soboroff.</surname>
          </string-name>
          <article-title>Overview of the TREC-2011 Microblog Track</article-title>
          . Available at: http://trec.nist.gov/pubs/trec20/ papers/MICROBLOG.OVERVIEW.pdf,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>1532</fpage>
          {
          <fpage>1543</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Robertson</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Zaragoza</surname>
          </string-name>
          .
          <article-title>The probabilistic relevance framework: Bm25 and beyond</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>3</volume>
          (
          <issue>4</issue>
          ):
          <volume>333</volume>
          {
          <fpage>389</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K. Sparck</given-names>
            <surname>Jones</surname>
          </string-name>
          and
          <string-name>
            <surname>C. van Rijsbergen.</surname>
          </string-name>
          <article-title>Report on the need for and provision of an ideal information retrieval test collection</article-title>
          .
          <source>Tech. Rep</source>
          .
          <volume>5266</volume>
          , Computer Laboratory, University of Cambridge, UK,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Strohman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Metzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Turtle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Indri: A language model-based search engine for complex queries</article-title>
          .
          <source>In Proc. ICIA</source>
          . Available at: http://www.lemurproject.org/indri/,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Abel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hauff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.-J.</given-names>
            <surname>Houben</surname>
          </string-name>
          , and
          <string-name>
            <given-names>U.</given-names>
            <surname>Gadiraju</surname>
          </string-name>
          . Groundhog Day:
          <article-title>Near-duplicate Detection on Twitter</article-title>
          .
          <source>In Proc. World Wide Web (WWW)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Twitter Search API</article-title>
          . https://dev.twitter.com/rest/public/search.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vieweg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Hughes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Starbird</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Palen</surname>
          </string-name>
          .
          <article-title>Microblogging During Two Natural Hazards Events: What Twitter May Contribute to Situational Awareness</article-title>
          .
          <source>In Proc. ACM SIGCHI</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>