<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Microposts</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ugo Scaiella</string-name>
          <email>{surname}@spaziodati.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gaetano Prestia</string-name>
          <email>prestia@netseven.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emilio Del Tessandoro</string-name>
          <email>deltessa@di.unipi.it</email>
          <email>deltessa@di.unipi.it veri@di.unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Mario Verì, Dipartimento di Informatica, University of Pisa</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Michele Barbera</institution>
          ,
          <addr-line>Stefano Parmesan, Spaziodati Srl, Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Net7 Srl</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>4</volume>
      <abstract>
        <p>In this paper we describe the approach taken for the “Making Sense of Microposts challenge 2014” (#Microposts2014), where participants were asked to cross reference micro-posts extracted from Twitter with DBpedia URIs belonging to a given taxonomy. For this task we deployed dataTXT1 which is the evolution of Tagme[3], the state-of-the-art topic annotator for short texts and which has proven to be very effective and efficient in several challenging scenarios[2].</p>
      </abstract>
      <kwd-group>
        <kwd>topic annotator</kwd>
        <kwd>entity extraction</kwd>
        <kwd>datatxt</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The #Microposts2014 challenge[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] focuses on the task of
annotating micro-posts with DBpedia entities belonging to
a given taxonomy. With respect to traditional Information
Retrieval tasks, such data poses new challenges in terms of
the effectiveness and efficiency of the algorithms and
applications because data is so short and noisy that it is difficult
to mine significant statistics that are rather available when
texts are long and well written. Additionally, participants
have to deal with the issue of associating extracted entities
with the provided taxonomy, that makes this challenge even
harder.
      </p>
      <p>
        For this challenge, we deployed dataTXT, an entity
extraction system that is the evolution of Tagme[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. An
instance of dataTXT has been specifically trained using the
official training set provided in this challenge.
      </p>
    </sec>
    <sec id="sec-2">
      <title>ANATOMY OF DATATXT</title>
      <p>dataTXT is able to identify meaningful sequences of one
or more terms in unstructured texts on-the-fly and with
high accuracy, and link them to a pertinent Wikipedia page.
1http://dandelion.eu/datatxt
PCeorpmyirsisgihont tco m20ak1e4 dhieglidtalbyorahuatrhdocro(sp)i/eoswonfearl(ls)o;r cpoaprtyionfgthpiserwmoitrktefdor
poenrlsyonfoarl oprricvlaatsesraonodmaucsaediesmgircanptuedrpwositehso.ut fee provided that copies are
nPoutbmliashdeedoradsipstarirbtuotefdthfoer#prMofiticorrocpoomstms2e0rc1i4al Wadovraknsthaogpe apnrdoctheeadticnogpsi,es
baveaariltahbislenotice anadstCheEfUulRlcVitoatli-o1n14o1n (thhet tfirpst:/p/acgeeu.rT-owsco.oprygo/tVhoelr-w1i1se4,1)to
online
republish, to post on servers or to redistribute to lists, requires prior specific
p#eMrmicisrsoiponosatnsd2/0o1r4a, fAeep.ril 7th, 2014, Seoul, Korea.</p>
      <p>Copyright 20XX ACM X-XXXXX-XX-X/XX/XX ...$10.00.
dataTXT maintains the core algorithm of its predecessor,
Tagme, but adds functionality and several improvements in
terms of cleaning the input text and identifying mentions.</p>
      <p>
        The algorithm is based on the anchor texts drawn from
Wikipedia for identifying mentions in input text. When an
input text is received, it judiciously cross-references each
anchor a found in the input text T with one pertinent page pa
of Wikipedia. dataTXT first identifies for each anchor a all
possible pages pa linked by a in Wikipedia. Then, from these
pages, it selects the best association a 7→ pa by computing
a score based on a “collective agreement” between the page
pa and the pages pb that can be associated with all other
anchors b1...bn detected in T . We deploy a voting-schema,
where pages pb vote for each candidate pa according to a
function that estimates the relatedness between two
Wikipedia pages by exploiting the underlying graph. Further
details of this voting-schema and the relatedness function
can be found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Not all mentions extracted in this way
are worth annotating, so a confidence score is assigned to
all mentions. This score is based on (a) a–priori statistics
based on Wikipedia and (b) other figures representing the
coherence of the candidate entity with respect to the whole
text. It is thus possible to discard those whose confidence
score is below a given threshold.
      </p>
      <p>
        dataTXT does not rely on any linguistic feature, but
only on statistics and data extracted from Wikipedia. We
argue that this approach, derived from Tagme, yields
better results when dealing with user generated content such
as micro-posts, where well-known NLP tools, such as part–
of–speech taggers, are less effective because texts are short,
fragmented and often contain slang and/or misspelled words.
An in-depth evaluation of Tagme’s effectiveness and
comparison with others annotators was recently published in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
showing the validity of this approach.
3.
      </p>
      <p>
        TRAINING
dataTXT was designed for short texts, but it is effective
for long texts as well[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However there are some
parameters that can be amended in order to better fit the
context of this challenge. One of them is , which is used to
tune the disambiguation algorithm[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and defines whether
dataTXT should rely more on the context or favor more
common topics in order to discover entities. Using a higher
value favors more common topics, which may lead to better
results when processing fragmented inputs where the
context is not always reliable. Two other parameters have been
taken into account: (a) the minimum link probability, say
δ, that is used to discard a mention that is rarely used as
anchor texts in Wikipedia; (b) the minimum commonness,
say γ, that is used to discard a possible association a 7→ pa,
thus reducing the “ambiguity” of a mention. Refer to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], for
further details on these two thresholds. dataTXT assigns
a confidence score to each annotation so that those that are
below a given threshold, say φ, can be discarded. This
parameters can be used to balance precision vs. recall and the
best value may vary based on the application context. For
each configuration we tested, we evaluated the results using
20 values of this threshold, ranging from 0 to 1.
      </p>
      <p>Another important issue we faced, is that the
annotation task of this challenge has been restricted to entities
belonging to a limited taxonomy. dataTXT is a generic
topic annotator and it extracts all topics contained in the
input text. If we considered the overall output produced
by dataTXT, the results, and in particular the precision,
would be significantly penalized because dataTXT also
includes topics that are not part of the taxonomy. As an
example consider this tweet, which was part of the
training set: “Bank of America posts $8.8 billion loss in second
quarter due to mortgage security settlement”. The human
annotators extracted Bank of America as the only mention
of this micropost, whereas dataTXT extracts also
mortgage and security linking them to Mortgage_loan and
Security_(finance) respectively. These are not errors of the
system, but #MSM2014 focused on a limited taxonomy and
mortgage and security are not part of it. Unfortunately,
given a DBpedia URI it was not possible to automatically
check whether or not the entity belongs to that taxonomy.
To address the issue, we initially tested a naive approach
using a white–list of entities derived from the training set.
This is useful but, of course, is not generic. Thus, we
designed another approach that provides the probability that a
generic entity belongs to the taxonomy based on the
Wikipedia categories and DBpedia types associated with the entity.
We thus gathered all Wikipedia categories and all DBpedia
types associated with each entity extracted by dataTXT
from the training set. We then counted the occurrences of
categories and types for all entities that were part of the
ground-truth and the occurrences of categories and types
for those that were not. For each category/type we then
computed a probability. Given an entity e extracted by
dataTXT, we thus computed the probability that e
belongs to the taxonomy by computing a weighted sum of
probabilities of all categories and types of e. Finally, we
discarded from the results all entities whose probability is
below a given threshold. The value of this threshold, called
β, was experimentally evaluated using the training set,
together with other parameters mentioned above. We also
tested a third approach by deploying a C4.5 classifier, which
was trained by exploiting types and categories derived as
mentioned above. Categories and types were thus deployed
as features to train the classifier.</p>
      <p>For parameters tuning, we simply used a grid search in
an N –dimensional space, using a 5-fold cross evaluation, in
order to avoid over–fitting. Note that dataTXT is very
efficient, and the evaluation of a single parameter
combination (ie the annotation of more than 2K tweets) takes about
800 ms, thus this search is feasible even with a long list of
combinations. Tuned values of parameters do not change
significantly across the different folds, showing a good
stability and generality of the approach (see Table 1).
0.15
0.2
0.2
0.2
0.2</p>
      <sec id="sec-2-1">
        <title>Approach 1. White–list only 2. White–list + types prob. 3. White–list + C4.5 classifier</title>
      </sec>
      <sec id="sec-2-2">
        <title>Recall 41.3 50.7 55.5</title>
        <p>F1
50.2
57.2
64.1
4.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>RESULTS</title>
      <p>
        During the training phase, we noticed several differences
between the annotation generated by dataTXT and the
annotation provided in the ground-truth, therefore we
implemented a few post–annotation steps to improve the
performance for this challenge: (a) dataTXT does not annotate
dates or numbers, so a step that identifies these types of
mentions using simple regular expressions was added; (b)
dataTXT annotates only the first occurrence of a mention,
so a post–processing step to handle repeated mentions was
added. These steps do not affect the core algorithm and thus
were not considered during the training phase, however they
improve the performance of dataTXT for this challenge.
Table 2 shows the overall results of the cross evaluation of
our approaches using the training set. These figures are not
directly comparable those presented in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], as #MSM2014
focused on a limited set of entities, i.e. the ones specified by
the taxonomy.
5.
      </p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSIONS</title>
      <p>We have described the approach taken by our group for
the #MSM2014 challenge, where we deployed dataTXT,
the evolution of the state-of-the-art topic annotator Tagme.
Given that its algorithm does not depend on linguistic
features, dataTXT is very accurate even in this scenario. We
have also outlined a basic approach to verticalize the general–
purpose extraction algorithm to improve the performance in
the domain defined within this challenge. We believe that
this approach to verticalization could be further refined by
applying more sophisticated machine–learning techniques,
such as SVM or CRF.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. E. Cano</given-names>
            <surname>Basave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stankovic</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making Sense of Microposts (#Microposts2014) Named Entity Extraction &amp; Linking Challenge</article-title>
          .
          <source>In Proc., #Microposts2014</source>
          , pages
          <fpage>54</fpage>
          -
          <lpage>60</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornolti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciaramita</surname>
          </string-name>
          .
          <article-title>A framework for benchmarking entity-annotation systems</article-title>
          .
          <source>In WWW</source>
          ,
          <fpage>249</fpage>
          -
          <lpage>260</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          and
          <string-name>
            <given-names>U.</given-names>
            <surname>Scaiella</surname>
          </string-name>
          .
          <article-title>Fast and Accurate Annotation of Short Texts with Wikipedia Pages</article-title>
          .
          <source>IEEE Software</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <fpage>70</fpage>
          -
          <lpage>75</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>