<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Experimental Study of a Hybrid Entity Recognition and Linking System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julien Plu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Rizzo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphae¨l Troncy</string-name>
          <email>raphael.troncyg@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>EURECOM</institution>
          ,
          <addr-line>Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an experimental study of the performance of a hybrid semantic and linguistic system for recognizing and linking entities from formal and informal texts. In the current literature, systems are generally tailored to one or a few types of textual documents (e.g. narrative texts, newswire articles, informal text such as microposts). In contrast, we assess the performance of a hybrid approach that adapts the entity extraction, recognition and linking process to the type of document being analyzed. The hybrid system relies on POS taggers, gazetteers and Twitter user account dereferencing modules to extract entities, NER modules to recognize and type entities and entity popularity, string distance measures and scoring functions to disambiguate entities by ranking potential link candidates. The evaluation results show the robustness of our proposed approach in terms of text-independence compared to the current state-of-the art.</p>
      </abstract>
      <kwd-group>
        <kwd>Entity Extraction</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Entity Recognition</kwd>
        <kwd>Entity Filtering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The approach we analyze in this paper links entities from both formal and informal
textual documents to DBpedia 2014 resources. The approach can be broken down into
three tasks: entity extraction, entity recognition and entity linking. Entity extraction
refers to the task of spotting mentions that can be entities in the text. Entity recognition
refers to the task of giving a type to the extracted entity. Entity linking refers to the task
of linking the mention to a targeted knowledge base, and it is often composed of two
sub-tasks: generating candidates and ranking them according to scoring functions. We
experiment with three standard benchmark corpora: two composed of tweets and one
composed of formal texts.</p>
      <p>
        Numerous approaches have been proposed to address the task of extracting, typing
and linking entities in formal (such as newswire content) and informal (such as
microposts) text. Among the recent and best performing systems, both WAT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and TagME [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
turn the task in a multi-stage pipeline composed of extraction, typing, linking and
pruning. The two systems have been tested on both types of text. AIDA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Babelfy [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and DBpedia Spotlight [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are other systems that well-perform when processing formal
text. Specific methods have been developed for analyzing tweets, such as the E2E
system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which has introduced a new methodology for the entity linking: they recast the
extraction and linking steps as a single process. Two other similar approaches are AIDA
for tweets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and DataTXT [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] (extension of TagME [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), which have adapted their
original well-performing approach on formal text to tweets.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Architecture</title>
      <p>Our goal is to link all the entities occurring in a text to their counterparts in DBpedia
2014. Entities that do not have an entry in the knowledge base will be linked to NIL.
Tweets are notoriously problematic to process compared to formal texts because of i)
hashtags (such as #barackobama referring to Barack Obama); ii) user mentions (e.g.
@ryeong9 referring to db:Ryeo Wook1); iii) acronyms (e.g. Met for Metropolitan
Police Service); iv) short length of only 140 characters and v) syntax which is often
grammar free and words are misspelled. In the following, we briefly describe the different
components of our hybrid approach.</p>
      <p>Text Normalization. Only for tweets, this step consists of normalizing the text.
We remove emoticons, extra white spaces and punctuation symbols belonging to two
unicode categories2: other and symbol.</p>
      <p>Entity Extraction and Recognition. This step is about detecting mentions from
text that are likely to be selected as entities. The objective is achieved by using the
following components: 1) POS tagger; 2) NER; 3) gazetteer and 4) time expressions
spotter. For tweets, we add a dereferencing Twitter account system to retrieve the user
name. Our POS tagging system is the Stanford NLP POS Tagger with a model trained
specifically for tagging tweets3 in order to be case insensitive and to get independent
tags for mentions and hashtags. For formal text we use the english-bidirectional-distsim
that provides a better precision but for a higher computing time. With the POS tagger,
we spot all the noun phrases and numbers. As NER system, we use Stanford NER
properly trained with the training set of a benchmark. The gazetteers reinforce this stage
bringing a robust spotting for well-known nouns. Tweets contain Twitter accounts, so
we dereference Twitter accounts and extracts their public names. Each of those systems
are launched in parallel. The type of an entity is given by Stanford NER. If a mention
is not detected by Stanford NER then no type is assigned.</p>
      <p>Entity Resolution. Given two overlapping mentions, e.g. States of America from
Stanford NER and United States from Stanford POS tagger, we only take the union of
the two phrases. We obtain the mention United States of America and the type provided
by Stanford NER is selected.</p>
      <p>
        Entity Linking. This step is composed of three sub-tasks: i) entity generation,
where we lookup up the entity in an index built on top of both DBpedia20144 and a
dump of the Wikipedia articles5 dated from October 2014 to get possible candidates; ii)
candidates filtering based on direct inbound and outbound links between the extracted
entities in Wikipedia; iii) entity ranking based on an in-house ranking function using
Levenshtein distance between the extracted mention and the title, the set of redirect
pages and the set of disambiguation pages of each candidate, weighted by their
PageRank. If an entity does not have an entry in the knowledge base, we normally link it to
1 db stands for http://dbpedia.org/resource/
2 http://www.fileformat.info/info/unicode/category/index.htm
3 https://gate.ac.uk/wiki/twitter-postagger.html
4 http://wiki.dbpedia.org/services-resources/datasets/
datasets2014
5 https://dumps.wikimedia.org/enwiki/
NIL. The detailed method followed for the linking step and how to build the index is
explained in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experimental Study</title>
      <p>
        Our hybrid approach has been tested against the test dataset of the #Micropost2014 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
and #Micropost2015 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] NEEL challenges and the OKE2015 challenge6. The
breakdown results for each of these datasets are available7. Table 1 shows the performance of
our approach in comparison with state-of-the art systems given the F1-measure at the
final linking stage.
(DTaatgaTMXET) Babelfy WAT AIDA E2E UTwente ousia acubelab
Results for #Micropost2014 and #Micropost2015 NEEL challenges come from the official
published results. We report on the best performing systems: E2E, DataTXT, AIDA, UTwente for
#Micropost2014 and ousia, acubelab for #Micropost2015. For the OKE2015 challenge, we have
changed the test dataset in order to fix annotation issues, those changes has been approved by the
organizers and committed to the official repository. We used the neleval8 scorer instead of the
official one used in the challenge. This explains why the results are different than the ones reported
by the organizers of the challenge. Those datasets have differences listed in Table 2. We have
chosen those different datasets to show the full potential of our approach which can be adapted
depending on the dataset to be processed in terms of features and kind of text.
      </p>
      <p>Datasets</p>
      <p>OKE2015
#Microposts2014
#Microposts2015
co-references typing NIL entities dates numbers
3 3 3 7 7
7 7 7 3 3
7 3 3 7 7</p>
      <p>Table 2. Features for each datasets
The results that are reported for DBpedia Spotlight, TagME, AIDA and Babelfy for OKE2015
have been obtained using their respective APIs with the best tested settings. For DBpedia
Spotlight, those settings are: confidence=0.3 and support=20. For TagME, those settings are:
include all spots=yes and epsilon=0.5. For AIDA those settings are: technique=GRAPH,
algorithm=COCKTAIL PARTY SIZE CONSTRAINED, alpha=0.6 and coherence=0.9. For Babelfy,
those settings are: lang=en, annType=NAMED ENTITIES, annRes=WIKI, match=EXACT MATCHING,
dens=true and th=0.4. Since WAT is not publicly accessible, we did not test it with the OKE
challenge dataset. For formal text (OKE benchmark), our system outperforms the other approaches
6 https://github.com/anuzzolese/oke-challenge
7 http://multimediasemantics.github.io/adel/
8 https://github.com/wikilinks/neleval
being tested. For informal text (NEEL corpora), the hybrid approach shows the robustness in
extracting and typing entities because we jointly use linguistic and semantic methods. The results
slightly drop at linking level.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we have presented the experimental study of a text independent hybrid approach
that extracts, recognizes and links entities to DBpedia 2014. We show that a successful approach
exploits both linguistic features and semantic features extracted from DBpedia. As future work,
we plan to focus on improving the linking task by making more use of graph-based algorithms,
and to improve our ranking function.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments References</title>
      <p>This work was partially supported by the innovation activity 3cixty (14523) of EIT Digital
(https://www.eitdigital.eu).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Y. M.</given-names>
            <surname>Amir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Johannes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Yusra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Artem</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Gerhard</surname>
          </string-name>
          .
          <article-title>Adapting aida for tweets</article-title>
          .
          <source>Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Cano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          , Stankovic Milan,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making Sense of Microposts (#Microposts2014) Named Entity Extraction &amp; Linking Challenge</article-title>
          .
          <source>In 4th International Workshop on Making Sense of Microposts</source>
          , Seoul, South Korea,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yosef</surname>
          </string-name>
          , I. Bordino, H. Fu¨ rstenau, M. Pinkal,
          <string-name>
            <given-names>M.</given-names>
            <surname>Spaniol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Taneva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thater</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Robust Disambiguation of Named Entities in Text</article-title>
          .
          <source>In 8th Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>782</fpage>
          -
          <lpage>792</lpage>
          , Stroudsburg, PA, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakob</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Garc´ıa-</article-title>
          <string-name>
            <surname>Silva</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Bizer</surname>
          </string-name>
          . DBpedia Spotlight:
          <article-title>Shedding Light on the Web of Documents</article-title>
          .
          <source>In 7th International Conference on Semantic Systems (I-Semantics)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>C.</given-names>
            <surname>Ming-Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bo-June</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ricky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Kuansan</surname>
          </string-name>
          .
          <article-title>E2e: An end-to-end entity linking system for short and noisy text</article-title>
          .
          <source>Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Moro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raganato</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          .
          <article-title>Entity Linking meets Word Sense Disambiguation: a Unified Approach</article-title>
          . TACL,
          <volume>2</volume>
          :
          <fpage>231</fpage>
          -
          <lpage>244</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>F.</given-names>
            <surname>Paolo</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ugo</surname>
          </string-name>
          . Tagme:
          <article-title>On-the-fly annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In Proceedings of the 19th ACM International Conference on Information and Knowledge Management</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>F.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          .
          <article-title>From TagME to WAT: a new entity annotator</article-title>
          .
          <source>In 1st ACM International Workshop on Entity Recognition &amp; Disambiguation (ERD)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>J.</given-names>
            <surname>Plu</surname>
          </string-name>
          , G. Rizzo, and
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          .
          <article-title>A Hybrid Approach for Entity Recognition and Linking</article-title>
          .
          <source>In 12th European Semantic Web Conference, Open Knowledge Extraction Challenge</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. G. Rizzo,
          <string-name>
            <surname>Cano Amparo</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          .
          <article-title>Making sense of Microposts (#Microposts2015) named entity recognition &amp; linking challenge</article-title>
          .
          <source>In WWW</source>
          <year>2015</year>
          , 5th International Workshop on Making Sense of Microposts (#
          <source>Microposts'15)</source>
          , Florence,
          <string-name>
            <surname>ITALIE</surname>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>S.</given-names>
            <surname>Ugo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Michele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Stefano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gaetano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. T.</given-names>
            <surname>Emilio</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Mario</surname>
          </string-name>
          . Datatxt at# microposts2014 challenge.
          <source>Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>