<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohamed Amir Yosef</string-name>
          <email>mamir@mpi-inf.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes Hoffart</string-name>
          <email>jhoffart@mpi-inf.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yusra Ibrahim</string-name>
          <email>yibrahim@mpi-inf.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Artem Boldyrev</string-name>
          <email>boldyrev@mpi-inf.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerhard Weikum</string-name>
          <email>weikum@mpi-inf.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Max Planck Institute for Informatics Saarbrücken</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>1141</volume>
      <abstract>
        <p>This paper presents our system for the “Making Sense of Microposts 2014 (#Microposts2014)” challenge. Our system is based on AIDA, an existing system that links entity mentions in natural language text to their corresponding canonical entities in a knowledge base (KB). AIDA collectively exploits the prominence of entities, contextual similarities, and coherence to effectively disambiguate entity mentions. The system was originally developed for clean and well-structured text (e. g. news articles). We adapt it for microposts, specifically tweets, with special focus on the named entity recognition and the entity candidate lookup.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Entity Recognition</kwd>
        <kwd>Entity Disambiguation</kwd>
        <kwd>Social Media</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Named Entity Recognition. AIDA originally uses
Stanford NER, with a model trained on newswire snippets, a
perfect fit for news texts. However, it is not optimized
for handling user generated content with typos and
abbreviations. Hence, we employ two different components for
mention detection: The first is Stanford NER with
models trained for caseless mention detection; the second is our
in-house dictionary-based NER tool. The dictionary-based
NER is performed in two stages:
1. Detection of named entity candidates using
dictionaries of all names of all entities in our knowledge base,
using partial prefix-matches for lookups to allow for
shortening of names or little differences in the later
part of a name. For example, we would recognize
the ticker symbol “GOOG” even though our dictionary
only contains “Google”. The character-wise matching
of all names of entities in our KB is efficiently
implemented using a prefix-tree data structure.
2. The large number of false positives are filtered using a
collection of heuristics, e. g. the phrase has to contain
a NNP tag or it has to end with a suffix signifying a
name such as “Ave” in “Fifth Ave”.</p>
      <p>Mention Normalization. The original AIDA did not
distinguish between the textual representation of the mention,
and its normalized form that should be used to query the
dictionary. For example, the hashtag "#BarackObama" should
be normalized to “Barack Obama” before matching it against
the dictionary. Furthermore, many mentions of named
entities are referred to in the tweet by their Twitter user ID,
such as "@EmWatson" the Twitter account of the British
actress “Emma Watson”. Because the Twitter user IDs are not
always informative we access the account metadata, which
contains the full user name most of the time. In fact, we
attach to each mention string a set of normalized mentions
and use all of them to query the dictionary. For example
"@EmWatson" will have the following normalized mentions
{“EmWatson”, “Em Watson”, “Emma Watson”}, and each
of them will be matched against the dictionary to retrieve
the set of candidate entities. As the prior probability is on
a per-mention basis, we compute the aggregate prior
probability of an entity ei given a mention mi:
prior(mi, ei) =</p>
      <p>max
m0∈N(mi)
(1)
where N (mi) is the set of normalized mentions of mi. The
maximum is taken in order not to penalize an entity if one
of the normalized mentions is rarely used to refer to it.
Approximate Matching. This step is employed iff the
previous normalization step did not produce candidate entities
for a given mention. For example, it is not trivial to
automatically split a hashtag like "#londonriots", and hence its
normalized mention set, {”londonriots”}, does not have any
candidate entities. We address this by representing both the
mention strings and dictionary keys as vectors of
charactertrigrams between which the cosine similarity is computed.
We only consider the candidate entity if cosine similarity
between the mention and candidate entity keys is above a
certain threshold (experimentally determined as 0.6).
Parameter Settings. In our graph representation, the weight
of a mention-entity edge is computed by a linear combination
of different similarity measures. To estimate the constants of
the linear combination, we split the provided tweets training
dataset into TRAIN and DEVELOP chunks, using TRAIN for the
estimation. We estimated further hyper-parameters for our
algorithm (like the importance of mention-entity vs.
entityentity edges) on DEVELOP.</p>
      <p>
        Unlinkable Mentions. Some mentions should not be
disambiguated to a entity, even though there are candidates
for it. This is especially frequent in the case of social
media, where a large number of user names are ambiguous but
do not refer to any existing KB entity – imagine how many
Will Smiths exists besides the famous actor. We address this
problem by thresholding on the disambiguation confidence
as defined in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where a mention is considered unlinkable
and thus removed if the confidence is below a certain
threshold, estimated as 0.4 on DEVELOP.
      </p>
    </sec>
    <sec id="sec-2">
      <title>4. EXPERIMENTS</title>
      <p>
        We conducted our experiments on the dataset provided in
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We carried out experiments with three different setups.
First we used Stanford NER trained for entity detection,
along with mention prior probability and key-phrases
matching for entity disambiguation. In the second experiment we
added coherence graph disambiguation to the previous
setting. The third setting is similar to the first one, but we use
our dictionary-based NER instead of Stanford’s for entity
detection. Note that we automatically annotate all
digitonly tokens as mentions using a regular expression, as all
numbers were annotated in the training data. The results
of running the three experiments on the testing dataset are
correspondingly provided with the following ids: AIDA 1,
AIDA 2 and AIDA 3.
      </p>
      <p>During our experiments, our runs achieved around 51% F1
on the DEVELOP part of the training data, where a mention
is counted as true positive only if both the mention span
matches the ground truth perfectly and the entity label is
correct.</p>
    </sec>
    <sec id="sec-3">
      <title>5. CONCLUSION</title>
      <p>AIDA is a robust framework that can be adapted to any
type of natural language text, here we use it to
disambiguate names to entities in tweets. We found that using
a dictionary-based NER worked well for the sometimes
illformatted inputs. An approximate candidate lookup
crucially improves recall, which in combination with discarding
low-confidence mentions improves the results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. E. Cano</given-names>
            <surname>Basave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stankovic</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making Sense of Microposts (#Microposts2014) Named Entity Extraction &amp; Linking Challenge</article-title>
          . In #Microposts2014.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Finkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grenager</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Incorporating Non-local Information into Information Extraction systems by gibbs sampling</article-title>
          .
          <source>ACL 2005</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Altun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum. Discovering Emerging</surname>
          </string-name>
          <article-title>Entities with Ambiguous Names</article-title>
          .
          <source>WWW 2014</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yosef</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Bordino</surname>
          </string-name>
          et al.
          <source>Robust Disambiguation of Named Entities in Text. EMNLP 2011</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Holt</surname>
          </string-name>
          . Twitter in numbers,
          <year>March 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Witten. An</given-names>
            <surname>Effective</surname>
          </string-name>
          ,
          <article-title>Low-Cost Measure of Semantic Relatedness Obtained from Wikipedia Links</article-title>
          . WIKIAI workshop at AAAI 2008
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yosef</surname>
          </string-name>
          et al.
          <article-title>AIDA: An online tool for accurate disambiguation of named entities in text and tables</article-title>
          .
          <source>VLDB 2011</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>