<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Revealing Entities from Textual Documents Using a Hybrid Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julien Plu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Rizzo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphae¨l Troncy</string-name>
          <email>raphael.troncyg@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>EURECOM</institution>
          ,
          <addr-line>Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The tasks of entity extraction, recognition, and linking are largely affected by the nature of the textual documents being analyzed. In fact, a lot of research efforts have focused on improving each task for both formal text (such as newswire documents) and for informal text (such as tweets). In this work, we propose a so-called hybrid approach that aims to be agnostic of the document type. Two datasets, namely the #Micropost2014 NEEL corpus and the OKE2015 test dataset, are used to benchmark the performance of our approach. The experimental results show that the approach presented in this paper outperforms the state-of-the-art systems on OKE2015 dataset and provides good results for the #Micropost2014 dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>Entity Recognition</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Entity Filtering</kwd>
        <kwd>Entity Resolution</kwd>
        <kwd>Micropost</kwd>
        <kwd>Formal Text</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The Web of documents has significantly grown in the last decades, resulting in an ever
increasing amount of unstructured data being published, such as newswire documents,
encyclopedic content or scientific papers. On the other hand, with the advent of social
media services, such as Twitter and Facebook, new unstructured data has contributed
to make the Web larger and wider, taping a larger human audience and embracing a
different style of writing. The content being shared on social media is generally short in
length, dynamic in terms of topics and events being covered, and informal in the writing
style (less curated syntax and grammar). In this paper, we refer to respectively formal
and informal text when talking about those two categories of textual documents.</p>
      <p>For years, the task of entity recognition, along with word sense disambiguation
and entity linking, has been tested on well-formed text (e.g newswire). The advent of
microposts has introduced a breakthrough: approaches being used in the past did not
fit this new type of textual data, and tailored approaches have been proposed to cope
with text characterized by: i) less than 140 characters (Twitter), ii) strong ambiguity, iii)
spelling errors, iv) non standard and lexical items, v) non standard syntactic patterns.</p>
      <p>
        The aim of this work is to propose a hybrid approach for performing the tasks of
entity extraction, recognition and entity linking1 that would perform fairly robustly on
1 We consider an entity, a DBpedia resource that corresponds to a Wikipedia article. We consider
a mention, one of the possible way to name an entity.
both formal and informal textual documents. The approach is hybrid since it makes use
of both linguistic and semantic algorithms working together and enclosed in a pipeline
of stages that can be turn on and off. In the experimental settings, we use DBpedia2014
as knowledge base to link the entities being extracted from text, and two standard
benchmark datasets: the #Microposts2014 NEEL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] corpus which is composed of tweets and
the OKE20152 corpus which is composed of paragraphs taken from Wikipedia articles.
The proposed approach is divided in three tasks: entity extraction, entity recognition,
and entity linking. The entity extraction task refers to spotting mentions from text. The
entity recognition task refers to the task of giving a type to the extracted mention. Entity
linking refers to the task of disambiguating the mention in a targeted knowledge base,
and it is often composed of two sub-tasks: generating candidates and ranking them
according to scoring functions.
      </p>
      <p>The paper is organized as follows: Section 2 reviews the strengths and weaknesses
of state-of-the-art approaches in processing formal and informal text. Section 3
describes our approach and its technical components. Section 4 reports the results of our
evaluation. Section 5 highlights some frequent errors while Section 6 proposes some
future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>State of the Art</title>
      <p>Numerous approaches have been proposed to tackle the task of extracting, typing and
linking entities in formal texts and microposts. We have divided the state-of-the art in
two parts: approaches that deal with formal texts and approaches dealing with
microposts, generally tweets.
2.1</p>
      <sec id="sec-2-1">
        <title>Formal Text</title>
        <p>
          Amongst the recent and best performing systems, WAT [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] builds on top of TagME
algorithms and follows the four steps approach we are advocating: extraction, typing,
linking and pruning. For the extraction, a dictionary that contains titles, surface forms
and redirect pages with a list of all their possible links ranked according to a
probability score is used. The extraction performance can also be tuned with an optional
binary classifier (SVM with linear or RBF kernel) using statistical (features) for each
entity referenced in the dictionary. For typing the entities, WAT relies on OpenNLP
NER and supports the three types PERSON, LOCATION and ORGANIZATION. For
linking entities, WAT uses two methods, namely voting-based and graph-based
algorithms. The voting-based approach assigns one score to each entity. The entity having
the highest score is then selected. The graph-based approach builds a graph where the
nodes correspond to mentions or candidates (entities) and the edges correspond to
either mention-entity or entity-entity relationships, each of these two kinds of edges being
weighted with three possible scores: i) identity, ii) commonness that is the prior
probability P r(ejm) and iii) context similarity that is the BM25 similarity score used by
Lucene3. The goal is to find the subgraph that interlinks as many mentions as possible.
2 https://github.com/anuzzolese/oke-challenge
3 https://lucene.apache.org/core/
        </p>
        <p>
          DBpedia Spotlight [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] uses a gazetteer containing a set of labels from the DBpedia
lexicalization dataset for the extraction. More precisely, the LingPipe Exact
DictionaryBased Chunker with the Aho-Corasick string distance measure is being used. Extracted
mentions that only contain verbs, adjectives, adverbs and prepositions can be detected
using the LingPipe part-of-speech tagger (POS Tagger) and then discarded. For the
typing step, DBpedia Spotlight re-uses the type of the link provided by DBpedia. For
the linking, DBpedia Spotlight relies on the so-called TF*ICF (Term Frequency-Inverse
Candidate Frequency) score computed for each entity. The goal of this score is to show
that the discriminative strength of a mention is inversely proportional to the number of
candidates it is associated with. This means that a mention that commonly co-occurs
with many candidates is less discriminative.
        </p>
        <p>
          AIDA [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] uses a different approach: instead of having one method for the extraction
and one method for the typing, the system relies on Stanford NER to combine both
tasks. For the linking, AIDA uses a similar approach than the graph-based method of
WAT. The graph is built in the same way but only one score for each kind of edge
(mention-entity or entity-entity) is proposed. The score used to weight the
mentionentity edges is a combination of similarity measure and popularity while the score used
to weight the entity-entity edges is based on a combination of Wikipedia-link overlap
and type distance.
        </p>
        <p>
          Babelfy [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] uses a part-of-speech tagger in order to identify the segments in the
text which contain at least one noun and that are substring of the entities referenced in
BabelNet (extraction step). For the typing step, the Babelnet categories are used. For the
linking, the system uses another graph-based approach where two main algorithms have
been developed: random walk and a heuristic for finding the subgraph that contains most
of the relations between the recognized mentions and candidates. The nodes are pairs
(mention, entity) and the edges correspond to existing relationships in BabelNet which
are scored. The semantic graph is built using word sense disambiguation (WSD) that
extracts lexicographic concepts and entity linking for matching strings with resources
described in a knowledge base.
        </p>
        <p>Those four methods (and particularly WAT and Babelfy) perform well on formal
text, but rather poorly on informal text such as tweets. WAT is limited as they match
any entry in the dictionary, including terms that are common words such as verbs or
prepositions. For typing, WAT is limited to the kind of mentions that can be typed by
OpenNLP NER. Similarly, AIDA depends on Stanford NER and the specific model
used by the CRF algorithm. Those four systems are also limited to the fixed knowledge
base being used, thus only extracting and linking entities that are referenced which
prevents to recognize emerging entities.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Informal text</title>
        <p>
          E2E [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] is an end-to-end entity linking system specialized for short and noisy text.
The architecture of this tool is composed of four steps: i) text normalization where the
retweet symbol, special characters (such as emojis) and additional white spaces are
removed. For example, the tweet RT: I like Paris :) is transformed into I like Paris; ii)
candidate generation where a dictionary based on Wikipedia and Freebase is used to
generate the candidate surface forms; iii) joint recognition and linking where the entity
recognition and linking are seen as a single task. The system uses a supervised learning
method where for a given message and a candidate mention, all the possible entities
are ranked and one is selected. The joint recognition and disambiguation task is crucial
to properly link surface forms to their corresponding entities; iv) overlap recognition,
which aims to resolve the conflicting cases of overlapping entities, where the system
uses dynamic programming by choosing the best-scoring set of non-overlapping
mention entity mappings. This resolution enables to improve the performance of the model
consistently.
        </p>
        <p>
          DataTXT [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is the TagME version adapted for tweets. The first step is to tokenize
and parse the tweet content to find and identify entities using a dictionary of entities
related to Wikipedia pages. The best association entity - Wikipedia page is selected by
computing a score based on an algorithm called collective agreement between each
page associated to the first entity and all the other pages associated to the other entities
found in the text. This score actually estimates the relatedness between two Wikipedia
pages.
        </p>
        <p>
          AIDA [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] has also released a system tailored for processing tweets. Before the
extraction step, the tweet content is normalized thanks to some pre-processing rules. For
example, #edSnowden is transformed into Edward Snowden and @TheRealHowardW
is transformed into Howard Wolowitz. The normalized (short) text is then processed as
it was a formal text.
        </p>
        <p>These three methods are tailored to process tweets. Like for the four methods
described in the Section 2.1, they do not recognize emerging entities as they are mostly
based on dictionaries or trained model for formal text (e.g. AIDA).
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Summary</title>
        <p>
          We summarize this state-of-the art in the Tables 1 and 2, inspired from [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], where we
show the similarities and differences of each system at extraction and linking level. In
Table 2, the abbreviation LEE means Link Emerging Entities.
        </p>
        <p>System
E2E
Babelfy
Spotlight
AIDA StanfordNER
TagME and
DataTXT
WAT</p>
        <p>NEE (Named Entity Extraction)
External Tools Main Features Method Knowledge Base
- N-Grams, stop words re- rule-based (candidate Wikipedia,
moval, punctuation as token filter), dictionary Freebase</p>
        <p>- NER dictionary
N-Grams, overlap resolu- dictionary, link proba- Wikipedia
tion, Wikipedia statistics bility
OpenNLP N-Grams, Wikipedia statis- dictionary, SVM Wikipedia,
tics NER dictionary
- N-Grams, POS, superstring dictionary Babelnet</p>
        <p>matching
LingPipePOS string matching, POS</p>
        <p>For the extraction, we observe that systems mainly use dictionaries based on a
particular knowledge base (semantic-based approach). When POS tagging is being used,
it is essentially a secondary feature which aims to enforce or to discard what has been
extracted with the dictionary. Contrarily to the others, AIDA uses a pure NLP approach
based on Stanford NER. TagME claims to make an overlap resolution between the
extracted mentions at the end of this process. Our system tackles the problem using both
a linguistic-based and a semantic-based approach of equal importance which
demonstrates a higher performance at extraction level as detailed in the Section 4.</p>
        <p>
          For the linking, we can see two main approaches: graph-based and arithmetic
combination. Contrarily to the others, E2E [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] uses a pure machine learning approach using
different features. At the end of this process, TagME and WAT do a pruning. None
of these systems claims to be able to handle emerging entities, that is, disambiguation
such entities to NIL. This is mainly due to their extraction approach. Our system tackles
the problem using an arithmetic combination inspired from TagME, to detect and link
emerging entities to NIL, while using a pruning process at the end for removing the
false positives.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Architecture and Implementation</title>
      <p>The architecture of our system is made for being modular. It is composed of different
modules that can be switched on or off according to the nature of the text to be
processed. A pipeline is a sequence of activated modules that will process an input text.
Modules can have different goals: extraction, typing, recognition, linking or pruning.
Those modules, according to their goal, have to respect some standards. Each module
has to provide an output corresponding to a specific interface in order to be understood
by the others modules. For example, in our approach, we have three extraction modules
(POS, NER and gazetteer) that should provide a similar output for being understood by
the candidate generator module. This system eases the overlap resolution process as it is
easier to compare outputs that follow the same presentation. Furthermore, each module
is run independently from the others.</p>
      <p>The candidate generation process is based on an index. If an extracted mention does
not have an entry in the index, we normally link it to NIL following the TAC KBP
convention4. The index is created on top of the DBpedia 2014 Knowledge Base5 and a
dump of the Wikipedia articles6 dated from October 2014. Each record of the index has
a key which corresponds to a DBpedia resource, while the features are listed in Table 3.
To create this index, we build an index using Lucene v5.2.1. The process requires 44
hours to be completed on a 64GB RAM, 20 core CPU at 2.5Ghz machine. The feature
number 13 in Table 3 is made with an in-house library to parse the Wikipedia dump.
We first tried several libraries that parse Wikipedia such as Sweble7, GWTWiki8 and
wikipedia-parser9. However, these libraries are either too complex to use for the simple
extraction we need or too greedy in terms of memory. We have therefore developed our
own library in order to extract the pairs (Wikipedia article title, number of times the
title appears in the article). For example, in the Wikipedia article of Paris we found 3
4 http://nlp.cs.rpi.edu/kbp/2014/
5 http://wiki.dbpedia.org/services-resources/datasets/
datasets2014
6 https://dumps.wikimedia.org/enwiki/
7 http://sweble.org/
8 https://code.google.com/p/gwtwiki/
9 https://github.com/Stratio/wikipedia-parser
links to the Wikipedia article Eiffel Tower, resulting in the tuple (Eiffel Tower, 3), and
this tuple will be associated to the corresponding entry of the DBpedia resource Paris
in the index.</p>
      <p>Our system uses three different modules to extract mentions: POS, NER and gazetteer.
Our POS tagging system is the Stanford NLP POS Tagger. We use two different models
depending if we process microposts (e.g. tweets) or formal texts (e.g. newswire). The
model used for microposts is https://gate.ac.uk/wiki/twitter-postagger.
html that is trained specifically in order to be case insensitive and to get independent
tags for mentions and hashtags. The one used for formal text is
english-bidirectionaldistsim that provides a better precision but for a higher computing time. We use the
POS tagger to spot all the noun phrases and numbers. Our second module is the NER
which relies on Stanford NER, trained with either the OKE2015 training set for formal
text or trained with the #Micropost2014 training set for tweets. Moreover, we use NER
to extract dates and other time related mentions. Our last module is the gazetteer that
aims to reinforce the extraction stage bringing a robust spotting for well-known nouns
such as abbreviations. Those three components are actually divided in five different
modules: one POS module to extract only proper-nouns, one POS module to extract
only numbers, one NER module without date spotting, one NER module with only date
spotting and one gazetteer module. Finally, a sixth module has been developed to
dereference mentions in microposts, for example, the mention @TheRealHowardW will be
transformed into Howard Wolowitz.</p>
      <p>Once all extractors have finished to extract mentions, we process their output with
an overlap resolution module. This module takes two inputs and provides one output. It
is run as many times as there are activated extractor modules minus one. For example,
if a pipeline uses the extractor modules POS for proper nouns (POSNNP), the NER
and gazetteer (GAZ) modules, then, the overlap resolution will first take the two inputs
(POSNNP and NER) to provide one output (POSNNP-NER) and it will take once again
two other inputs (POSNNP-NER and GAZ) to provide one final output
(POSNNPNER-GAZ). Therefore, when three different extractor modules are used, the overlap
resolution module is run twice. We have developed such a module since, sometimes, at
least two extractors provide overlapping mentions. For example, given the two extracted
mentions States of America from the NER module and United States from the POS for
proper noun module, we detect that there is an overlap between both mentions according
to their offset in the processed sentence. We take the union of both boundaries to create
a new mention and we remove the two others. We obtain the mention United States
of America and the type provided by the NER module is selected. The mechanism is
the same if one mention is included in another one. For example, if United States and
United States of America are extracted, the later will be kept. Finally, there are cases of
ambiguity. For example, with the text Yesterday I went to Los Angeles, California, Los
Angeles, California and Los Angeles, California are three different valid extractions.
We have decided to systematically keep the longest mention, in this case Los Angeles,
California.</p>
      <p>
        Once we have independent mentions, the candidate generation module searches the
index to retrieve as many candidates as possible. This means that for each mention, we
have potentially numerous candidates, while many of them have to be filtered out
because they are most likely not related to the context. We use a filter module that creates
a graph with all the candidates of each mention and find the densest graph between all
of these candidates, similarly to [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Our approach is, however, slightly different: we
use the feature number 13 (quotes) described in the Table 3 and not BabelNet in order
to build the graph. The edges of the graph are weighted according to the number of
occurrence of the link between each candidates. For example, given the Wikipedia article
describing the Eiffel Tower, if there is one outbound link to Paris in Texas and three
to Paris in France, both candidates (Paris in Texas and Paris in France) will be kept.
However, the weight of Paris in France will be higher than the one of Paris in Texas. In
case all candidates of a mention do not have any relation with any other candidate of
the other mentions, all its candidates are kept.
      </p>
      <p>Once we have reduced the number of candidates, we use a ranking module in order
to score each of those candidates and to pick up the best one. This module is using a
rank function r(l) to compute, for each candidate, a score based on a string similarity
measure between the extracted mention and the title of the candidate, the set of redirect
and the set of disambiguation pages associated to this candidate and its PageRank.
r(l) = (a L(m; title) + b max(L(m; R)) + c max(L(m; D))) P R(l)
(1)
where the weights a, b and c have to follow the following hypothesis: a + b + c = 1 and
a &gt; b &gt; c. We take the assumption that the string distance measure between a mention
and a title is more important than the distance measure with a redirect page and is itself
more important than the distance measure with a disambiguation page.</p>
      <p>The last module is the pruning one, which is used to detect and remove the false
positive annotations10 in order to improve the precision of the system. We use a machine
learning approach, with the algorithm k-NN. We have tried four different algorithms
(Random Forest, Naive Bayes, SVM and k-NN), and have empirically assessed that
k-NN generally provides the best results. To train this algorithm and getting a model,
we use ten features, most of them are listed in Table 3: the extracted mention and the
title, type, PageRank, HITS, inLinks, outLinks, length, redirectsNumber and r(l) of the
entity. The training method uses a four steps approach: 1) run our system on a training
set; 2) classify entities as true or false according to the entities in the Gold Standard of
the training set and the ones provided by the results of our system; 3) create a file with
the features of each of these entities and their true / false classification; 4) train k-NN
with this file that contains the features to get a model. Once the model has been created,
we let k-NN classify each annotation provided by our system to true or false.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>Our hybrid approach has been benchmarked against the test dataset of the
#Micropost2014 NEEL challenge and the test dataset of the OKE2015 challenge. An example
of #Micropost2014 and OKE2015 datasets are provided in Table 4.
10 We consider an annotation as the tuple (mention, entity)
4.1</p>
      <sec id="sec-4-1">
        <title>Experimental Settings</title>
        <p>We configured two different pipelines, one for each dataset. For OKE2015, we used
the following modules: three extractors (POS for proper nouns, NER without date and
gazetteer), the index lookup, the filtering and the scoring. For #Micropost2014, we used
six extractors (POS for proper nouns, POS for numbers, NER with only dates, NER
without dates, gazetteer and Twitter account dereferencing), the index lookup, the
filtering and the scoring. We have run these two pipelines with and without the pruning
module in order to assess its performance. This shows the advantages of our hybrid
approach of being able to adapt an entity linking process by adding or removing modules
(methods) and combining them.</p>
        <p>dataset text links
#Microposts2014 Murdoch unable to answer who the db:Rupert Murdoch
top legal officer at News International db:News UK
was. (via: http://mm4a.org/q3rx06) #p2 db:News of the World
#notw #hackgate db:News International phone hacking scandal
OKE2015
Breakdown figures for the #Micropost2014 NEEL challenge with and without the
pruning reported in Table 5 have been computed with the official scorer of the challenge11.
In the breakdown figures, the results at the recognition level is not presented since
typing entities was not required by the challenge. Breakdown figures for the OKE2015
challenge with and without the pruning reported in Table 6 have been computed with
the neleval scorer12.</p>
        <p>
          The Table 7 shows the performance of our approach in comparison to other systems,
using the F1-measure at the final linking stage. Results for #Micropost2014 for TagME
(more precisely DataTXT), AIDA, E2E and UTwente are coming from the official
results of the challenge, while Babelfy and DBpedia Spotlight have not been tested as they
are not made for processing tweets. We refer the reader to [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] for the complete results of
11 https://github.com/giusepperizzo/neeleval
12 https://github.com/wikilinks/neleval
        </p>
        <p>Without Pruning With Pruning</p>
        <p>Task P R F1 P R F1
Extraction 78.2 65.4 71.2 83.8 9.3 16.8
Recognition 65.8 54.8 59.8 75.7 8.4 15.1</p>
        <p>Linking 49.4 46.6 48 57.9 6.2 11.1
Table 6. Breakdown figures on the OKE
challenge training set with and without the
pruning.
all systems having participated in the 2014 NEEL challenge while we only reproduce
in Table 7 the best performing systems. For the OKE2015 challenge, the results have
been computed with the neleval scorer.</p>
        <p>The results that are reported for DBpedia Spotlight, TagME, AIDA and Babelfy
for OKE2015 have been obtained using their respective APIs with the best tested
settings. For DBpedia Spotlight, those settings are: confidence=0.3 and support=20. For
TagME, those settings are: include all spots=yes and epsilon=0.5. For AIDA, those
settings are: technique=GRAPH, algorithm=COCKTAIL PARTY SIZE CONSTRAINED,
alpha=0.6 and coherence=0.9. For Babelfy, those settings are: lang=en,
annType=NAMED ENTITIES, annRes=WIKI, match=EXACT MATCHING, dens=true
and th=0.4. WAT is not publicly available and could not be tested with the OKE
challenge dataset.</p>
        <p>Our approach outperforms those systems when analyzing formal texts. E2E and
UTwente still perform significantly better than our approach when processing
microposts while we achieve similar results than TagME (DataTXT).
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Error Analysis</title>
      <p>For #Micropost2014, an error analysis shows that our approach misses 490 entities and
finds 578 additional entities on a total of 1460 entities contained in the test dataset.
Amongst the most common errors, our approach has problems with entities which do
not contain tokens tagged as proper nouns but as noun (e.g. phone hacking). There is
also a problem with the possessive endings (e.g Beyonce’s for Beyonce). Syntactically
speaking, a noun phrase may contain a noun phrase, which provides duplicates. For
example, in the sentence She gets into the action, overlapping Lloyd, the noun phrase
the action, overlapping Llyod can be split in two noun phrases, which are the action
and overlapping Lloyd. The extraction may lead to these three cases.</p>
      <p>For OKE2015, an error analysis shows that our approach misses 207 entities and
find 167 additional entities on a total of 594 entities contained in the test dataset. The
most common error for this challenge is that our approach does not resolve the
coreferences while the test dataset contained a lot of them. The training set was not big
enough to properly train Stanford NER. This is why the correct mention is often
extracted but is given a wrong type. The test dataset contained errors in the gold standard,
some of them have been fixed by us, but some others need a closer work with the
organizers to be resolved.</p>
      <p>
        We observe that there is still a large margin of progress to reduce the performance
drop between the results at the recognition or extraction stage and the final results at
the linking stage. This drop can be explained by the score function and the candidates
filter used at the linking stage. For the pruning, we can see that it increases a bit the
precision but at the cost of significantly decreasing the recall. It overall performs poorly
as it removes too many (correct) mentions to get good results at the linking stage. This
idea has been inspired by the WAT system [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. However, some features we choose differ
from the ones used in WAT resulting in this serious performance drop. We stay positive
on the fact that a pruning step can typically help increasing the precision when a real
high recall at the recognition or extraction level is obtained.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>The results are encouraging and show that this hybrid approach which exploits both
linguistic and semantic features can achieved the expected behavior. As future work, we
plan to improve the linking stage by making more use of graph-based algorithms with
more accurate ranking functions and filtering. Another future work is to further develop
our pruning strategy by reviewing the list of feature used to show the full potential of our
hybrid approach. We finally aim to optimize the creation of the index by parallelizing
the process and to multiply the number of knowledge bases on which entities can be
disambiguated against.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments References</title>
      <p>This work was partially supported by the innovation activity 3cixty (14523) of EIT
Digital (https://www.eitdigital.eu).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Cano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          , Stankovic Milan,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making Sense of Microposts (#Microposts2014) Named Entity Extraction &amp; Linking Challenge</article-title>
          .
          <source>In 4th International Workshop on Making Sense of Microposts</source>
          , Seoul, South Korea,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>L.</given-names>
            <surname>Chenliang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Jianshu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yuxia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Anwitaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aixin</surname>
          </string-name>
          , and L.
          <string-name>
            <surname>Bu-Sung</surname>
          </string-name>
          .
          <article-title>TwiNER: Named Entity Recognition in Targeted Twitter Stream</article-title>
          .
          <source>In 35th International Conference on Research and Development in Information Retrieval (SIGIR)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yosef</surname>
          </string-name>
          , I. Bordino, H. Fu¨rstenau, M. Pinkal,
          <string-name>
            <given-names>M.</given-names>
            <surname>Spaniol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Taneva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thater</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Robust Disambiguation of Named Entities in Text</article-title>
          .
          <source>In 8th Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>782</fpage>
          -
          <lpage>792</lpage>
          , Stroudsburg, PA, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakob</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Garc´ıa-</article-title>
          <string-name>
            <surname>Silva</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Bizer</surname>
          </string-name>
          . DBpedia Spotlight:
          <article-title>Shedding Light on the Web of Documents</article-title>
          .
          <source>In 7th International Conference on Semantic Systems (I-Semantics)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>C.</given-names>
            <surname>Ming-Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bo-June</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ricky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Kuansan</surname>
          </string-name>
          .
          <article-title>E2E: An End-to-End Entity Linking System for Short and Noisy Text</article-title>
          .
          <source>In Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Moro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raganato</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          .
          <article-title>Entity Linking meets Word Sense Disambiguation: a Unified Approach</article-title>
          . TACL,
          <volume>2</volume>
          :
          <fpage>231</fpage>
          -
          <lpage>244</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>F.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          .
          <article-title>From TagME to WAT: a new entity annotator</article-title>
          .
          <source>In 1st ACM International Workshop on Entity Recognition &amp; Disambiguation (ERD)</source>
          , pages
          <fpage>55</fpage>
          -
          <lpage>62</lpage>
          , Gold Coast, Australia,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>U.</given-names>
            <surname>Scaiella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Barbera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Parmesan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Prestia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Del Tessandoro</surname>
          </string-name>
          , and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Ver`ı</article-title>
          . DataTXT at# Microposts2014 Challenge.
          <source>In Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yosef</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ibrahim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Boldyrev</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Adapting AIDA for Tweets</article-title>
          .
          <source>In Making Sense of Microposts (# Microposts2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>