<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amparo E. Cano Giuseppe Rizzo</string-name>
          <email>amparo.cano@open.ac.uk</email>
          <email>amparo.cano@open.ac.uk giuseppe.rizzo@di.unito.it</email>
          <email>giuseppe.rizzo@di.unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthew Rowe</string-name>
          <email>m.rowe@lancaster.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milan Stankovic</string-name>
          <email>milstan@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Varga</string-name>
          <email>a.varga@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aba-Sah Dadzie</string-name>
          <email>a.dadzie@cs.bham.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Knowledge Media Institute Università di Torino</institution>
          ,
          <addr-line>Italy</addr-line>
          ,
          <institution>The Open University</institution>
          ,
          <addr-line>UK EURECOM</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science, University of Birmingham</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computing and</institution>
          ,
          <addr-line>Communications</addr-line>
          ,
          <institution>Lancaster University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Sépage, France, Université Paris-Sorbonne</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>The OAK Group, The University of Sheffield</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>1141</volume>
      <fpage>54</fpage>
      <lpage>60</lpage>
      <abstract>
        <p>Microposts are small fragments of social media content and a popular medium for sharing facts, opinions and emotions. They comprise a wealth of data which is increasing exponentially, and which therefore presents new challenges for the information extraction community, among others. This paper describes the 'Making Sense of Microposts' (#Microposts2014) Workshop's Named Entity Extraction and Linking (NEEL) Challenge, held as part of the 2014 World Wide Web conference (WWW'14). The task of this challenge consists of the automatic extraction and linkage of entities appearing within English Microposts on Twitter. Participants were set the task of engineering a named entity extraction and DBpedia linkage system targeting a predefined taxonomy, to be run on the challenge data set, comprising a manually annotated training and a test corpus of Microposts. 43 research groups expressed intent to participate in the challenge, of which 24 signed the agreement required to be given a copy of the training and test datasets. 8 groups fulfilled all submission requirements, out of which 4 were accepted for the presentation at the workshop and a further 2 as posters. The submissions covered sequential and joint methods for approaching the named entity extraction and entity linking tasks. We describe the evaluation process and discuss the performance of the different approaches to the #Microposts2014 NEEL Challenge.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Microposts</kwd>
        <kwd>Named Entity</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Extraction</kwd>
        <kwd>Linking</kwd>
        <kwd>Disambiguation</kwd>
        <kwd>Challenge</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The purpose of this challenge was to set up an open and
competitive environment that would encourage participants to deliver
novel or improved approaches to extract entities from Microposts
and link them to their DBpedia counterpart resources (if defined).
This report describes the #Microposts2014 NEEL Challenge, our
collaborative annotation of a corpus of Microposts and our
evaluation of the performance of each submission. We also describe
the approaches taken in the participants’ systems – which use both
established and novel, alternative approaches to entity extraction
and linking. We describe how well they performed and how
system performance differed across approaches. The resulting body
of work has implications for researchers interested in the task of
information extraction from social media.</p>
    </sec>
    <sec id="sec-2">
      <title>2. THE CHALLENGE</title>
      <p>In this section we describe the goal of the challenge, the task set,
and the process we followed to generate the corpus of Microposts.
We conclude the section with the list of the accepted submissions.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 The Task and Goal</title>
      <p>The NEEL Challenge task required participants to build semi-automated
systems in two stages:
(i) generally known as Named Entity Extraction (NEE) – in which
participants were to extract entity mentions from a tweet; and
(ii) known as Named Entity Linking (NEL), in which each entity
extracted is linked to an English DBpedia v3.9 resource.</p>
      <p>
        For this task we considered the definition of anentity in the general
sense of being, in which an object or a set of objects do not
necessarily need to have a material existence, but which however must
be characterized as an instance of a taxonomy class. To facilitate
the creation of the gold standard (GS) we limited the entity types
evaluated in this challenge by specifying the taxonomy to be used:
the NERD ontology v0.51 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. To this we added a few concepts
from the DBpedia taxonomy. The taxonomy was not considered as
normative in the evaluation of the submissions, nor for the ranking.
This is a deliberate choice, to increase the complexity of the task
and to let participants perform taxonomy matching starting from
the distribution of the entities in the GS. The list of classes in the
taxonomy used is distributed with the released GS2.
      </p>
      <p>Beside the typical word-tokens found in a Micropost, new to this
year’s challenge we considered special social media markers as
entity mentions as well. These Twitter markers are tokens introduced
with a special symbol. We considered two such markers: hashtags,
prefixed by #, denoting the topic of a Micropost (e.g.
#londonriots, #surreyriots, #osloexpl), and mentions prefixed by @, referring
to Twitter user names, which include entities such as organizations
(e.g. @bbcworldservice) and celebrities (e.g. @ChadMMurray,
@AmyWinehouse).</p>
      <p>Participants were required to recognize these different entity types
within a given Micropost, and to extract the corresponding entity
link tuples. Consider the following example, taken from our
annotated corpus:
The 2nd token (the mention @bbcworldservice) in this Micropost
refers to the international broadcaster, the BBC World Service; the
7th token refers to the location Oslo; while the 8th token (the
hashtag #oslexp) refers to the 2011 Norway terrorist attack. An entry
to the challenge would be required to spot these tokens and display
the result as a set of annotations, where each line corresponds to a
tab-separated entity mention3 and entity link4:
1http://nerd.eurecom.fr/ontology/nerd-v0.5.
n3
2The NEEL Challenge GS available for download from: http:
//ceur-ws.org/Vol-1141/microposts2014-neel_
challenge_gs.zip
3Note that the annotated result returns tokens without the social
media markers (# and @) in the original Micropost.
4In this example “dbpedia:” refers to the namespace prefix of a
DBpedia resource (see http://dbpedia.org/resource)
Correctly formatted result:
bbcworldservice dbpedia:BBC_World_Service
Oslo dbpedia:Oslo
oslexp dbpedia:2011_Norway_attacks
We also consider the case where an entity is referenced in a tweet
either as a noun or a noun phrase, if it:
a) belongs to one of the categories specified in the taxonomy;
b) is disambiguated by a DBpedia URI within the context of the
tweet. Hence any (single word or phrase) entity without a
disambiguation URI is disregarded;
c) subsumes other entities. The longest entity phrase within a
Micropost, composed of multiple sequential entities and that can
be disambiguated by a DBpedia URI, takes precedence over its
component entities.</p>
      <p>Consider the following examples:</p>
      <sec id="sec-3-1">
        <title>1. [Natural History Museum at Tring];</title>
        <p>2. [News International chairman James
evidence to MPs on phone hacking;
3. [Sony]’s [Android Honeycomb] Tablet</p>
      </sec>
      <sec id="sec-3-2">
        <title>Murdoch]’s</title>
        <p>For the 3nd case, even though they may appear to be a coherent
phrases, since there are no DBpedia URIs for [Sony’s Android
Honeycomb] or [Sony’s Android Honeycomb Tablet],
the entity phrase is split into what are the (valid) component
entities highlighted above.</p>
        <p>To encourage competition we solicited sponsorship for the winning
submission. This was provided by the European project LinkedTV5,
who offered a prize of an iPad This generous sponsorship is
testament to the growing interest in issues related to automatic
approaches for gleaning information from (the very large amounts of)
social media data.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2.2 Data Collection and Annotation</title>
      <p>The challenge data set comprises 3,505 tweets extracted from a
collection of over 18 million tweets. This collection, provided by the
Redites project6, covers event-annotated tweets collected for the
period 15th July 2011 to 15th August 2011 (31 days). It extends
over multiple notable events, including the death of Amy
Winehouse, the London Riots and the Oslo bombing. Since the NEEL
Challenge task is to automatically extract and link entities, we built
our data set considering both event and non-event tweets. Event
tweets are more likely to contain entities; non-event tweets
therefore enable us to evaluate the performance of the system in avoiding
false positives in the entity extraction phase.</p>
      <p>Statistics describing the training and test sets are provided in
Table 1. The dataset was split into training (70%) and test (30%)
sets. The training set contains 2,340 tweets, with 41,037 tokens
and 3,819 named entities; the test set contains 1,165 tweets, with
20,224 tokens and 1,458 named entities. The tweets are relatively
5http://www.linkedtv.eu
6http://demeter.inf.ed.ac.uk/redites
long in both data sets; the average number of tokens per tweet is
17.54±5.70 in the training, and 17.36±5.59 in the test set. The
average number of entities per tweet is also relatively high, at
3.26±3.37 for the training and 2.50±2.94 for the test dataset. The
percentage of tweets without any valid entities is 32% (775 tweets)
in the training, and 40% (469 tweets) in the test set. There is a fair
bit of overlap of entities between the training and test data: 13.27%
(316) of the named entities in the training data also occurs in the
test dataset. With regard to the tokens in the original tweets with
hashtag and mention social media markers, a total of 406 hashtags
represented valid entities in the training, with 184 in the test set.
The total number of valid entity mentions was 133 in the training,
and 73 in the test data set.</p>
      <p>
        The annotation of each Micropost in the training set gave all
participants a common base from which to learn extraction patterns. In
order to assess the performance of the submissions we used an
underlying gold standard (GS), generated by 14 annotators, who had
different backgrounds, including computer scientists, social
scientists, social web experts, semantic web experts and linguists.
The annotation process comprised the following phases7
Phase 1. Unsupervised annotation of the corpus was performed, to
extract candidate links that were used as input to the next
stage. The candidates were extracted using the NERD
framework [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Phase 2. The data set was divided into batches, with three different
annotators to each batch. In this phase annotations were
performed using CrowdFlower8. The annotators were
asked to analyze the NERD links generated in phase 1 by
adding or removing entity-annotations as required. The
annotators were also asked to mark any ambiguous cases
encountered.</p>
      <p>Phase 3. In the final stage, consistency checking, three experts
double-checked the annotations and generated the GS (for
both the training and test sets). Three main tasks were
carried out here: (1) cross-consistency check of entity
types; (2) cross-consistency check of URIs; (3) resolution
of ambiguous cases raised by the 14 annotators.</p>
      <p>The complete data set, including a list of changes and the gold
stan7We aim to provide a more detailed explanation of the annotation
process and the rest of the NEEL Challenge evaluation process in a
separate publication.
8http://crowdflower.com
dard, is available for download9 with the #Microposts2014
Workshop proceedings, accessible under the Creative Commons
AttributionNonCommercial-ShareAlike 3.0 Unported License10.</p>
    </sec>
    <sec id="sec-5">
      <title>2.3 Challenge Submissions</title>
      <p>The challenge attracted a lot of interest from research groups spread
across the world. Initially, 43 groups expressed their intent to
participate in the challenge; however only 8 completed submission.
Each submission consisted of a short paper explaining the system
approach, and up to three different test set annotations generated
by running the system with different settings. After peer review,
4 submissions were accepted, and a further 2 as posters. The
submission run with the best overall performance for each system was
used in the rankings (see Table 4). The submissions accepted are
listed in Table 2.</p>
    </sec>
    <sec id="sec-6">
      <title>2.4 System Descriptions</title>
      <p>We present next an analysis of the participants’ systems for the
Named Entity Extraction and Linking (NEEL) tasks. Except for
submission 18, who treated the NEEL task as a joint task of Named
Entity Extraction (NEE) and Named Entity Linking (NEL); all
participants approached the NEEL task as two sequential sub-tasks
(i.e. NEE first, followed by NEL). A summary of these approaches
includes:
i) use of external systems;
ii) main features used;
iii) type of strategy used;
9http://ceur-ws.org/Vol-1141/
microposts2014-neel_challenge_gs.zip
10Following the Twitter ToS we only provide tweet IDs and
annotations for the training set; and tweet IDs for the test set.
2
3
1
3
2
1
iv) use of external sources.</p>
      <p>
        The NEE task on Microposts is on its own challenging. One of the
main strategies was to use off-the-shelf named entity recognition
(NER) tools, improved through the use of extended gazetteers.
System 18 approached the NEE task from scratch using a rule-based
approach; all others made use of external toolkits. Some of these
were Twitter-tuned and were applied for:
i) feature extraction, including the use of the TwitterNLP (2013)
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and TwitterNLP (2011) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] toolkits for POS tagging
(systems 16, 20);
ii) entity extraction with TwiNER [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Ritter’s NER [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and
      </p>
      <p>
        TAGME [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (systems 13, 16, 19).
      </p>
      <p>
        Other external toolkits which address NEE in longer newswire texts
were also applied, including Stanford NER [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and DBpedia
Spotlight [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] (systems 15, 20).
      </p>
      <p>
        Another common trend across these systems was the use of
gazetteerbased, rule-matching approaches to improve the coverage of the
off-the-shelf tools. System 13 applied simple regular expression
rules to detect additional named entities not found by the NE
extractor (such as numbers, and dates); systems 15 and 18 applied
rules to find candidate entity mentions using a knowledge base
(among others, Freebase [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). Some systems also applied name
normalization for feature extraction (systems 15, 18). This strategy
was particularly useful for catering for entities originally appearing
as hashtags or username mentions. For example, hashtags such as
#BarackObama were normalized into a composite entity mention
“Barack Obama"; and “@EmWatson" into “Emma Watson".
The NEL task involved in some cases the use of off-the-self tools,
for finding candidate links for each entity mention and/or for
deriving mention features (systems 13, 19, 20). A common trend across
systems was the use of external knowledge sources including:
i) NER dictionaries (e.g. Google CrossWiki [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]);
ii) Knowledge Base Gazetteers (e.g. Yago [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], DBpedia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]);
iii) Weighted lexicons (using e.g. Freebase [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Wikipedia);
iv) other sources (e.g. Microsoft Web N-gram [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]).
      </p>
      <p>A wide range of different features was investigated for the linking
strategies. Some systems characterized an entity using
Micropostderived features with Knowledge base (KB)-derived features
(systems 13, 15, 16, 19). Micropost-derived features include the use of
lexical (e.g., N-grams, capitalization) and syntactical (e.g., POS)
features, while KB-derived features included the use of URIs,
anchor text and link-based probabilities (see Table 3). Additionally,
features were extended by capturing jointly the local (within a
Micropost) and global (within a knowledge base) contextual
information of an entity, via graph-based features (such as entity
semantic cohesiveness) (system 18). Further novel features included the
use of Twitter account metadata for characterizing mentions and
popularity-based statistical features for characterizing entities
(systems 16, 18).</p>
      <p>The classificationstrategies used for entity linking included
supervised approaches (systems 13, 15, 16, 18, 19) existing off-the-shelf
approaches enhanced with simple heuristics (e.g. the search+rules)
(system 20).</p>
    </sec>
    <sec id="sec-7">
      <title>3. EVALUATION OF CHALLENGE SUBMIS</title>
    </sec>
    <sec id="sec-8">
      <title>SIONS</title>
      <p>We describe next the evaluation measures used to assess the
goodness of the submissions and conclude with the final challenge
rankings, with submissions ordered according to the F1 measure.</p>
    </sec>
    <sec id="sec-9">
      <title>3.1 Evaluation Measures</title>
      <p>We evaluate the goodness of a system S in terms of the performance
of the system to both recognize and link an entity from a test set
T S. Per each instance in T S, a system provides a set of pairs P
of the form: entity mention (e), and link (l). A link is any valid
DBpedia URI11 that points to an existing resource (e.g. http://
dbpedia.org/resource/Barack_Obama). The evaluation
consists of comparing submission entry pairs against those in the
gold standard GS. The measures used to evaluate each pair are
precision P , recall R, and f-measure F1. The evaluation is based
on micro-averages.</p>
      <p>First, a cleansing stage is performed over each submission,
resolving where needed, the redirects. Then, to assess the correctness of
the pairs provided by a system S, we perform an exact-match
evaluation, in which a pair is correct only if both the entity mention and
the link match the corresponding set in the GS. Pair order is also
relevant. We define(e, l)S ∈ S as the set of pairs extracted by the
system S, (e, l)GS ∈ GS denotes the set of pairs in the gold
standard. We define the set of true positives T P , false positives F P ,
and false negatives F N for a given system as:</p>
      <p>T P = {(e, l)S |(e, l)GS ∈ (S ∩ GS)}
F P = {(e, l)S |(e, l)GS ∈ S ∧ (e, l) ∈/ GS)}</p>
      <p>F N = {(e, l)S |(e, l)GS ∈ GS ∧ (e, l) ∈/ S}
Thus T P defines the set of relevant pairs inT S, in other words the
set of pairs in T S that match corresponding ones in GS. F P is the
set of irrelevant pairs in T S, in other words the pairs in T S that do
not match the pairs in GS. F N is the set of false negatives denoting
the pairs that are not recognised by T S, yet appear in GS. Since
our evaluation is based on a micro-average analysis, we sum the
individual true positives, false positives, and false negatives of each
system across all Microposts. As we require an exact-match for
pairs (e, l) we are looking for strict entity recognition and linking
matches; each system has to link each entity e recognised to the
correct resource l.</p>
      <p>From this set of definitions, we define precision, recall, and
fmeasure as follows:</p>
      <p>P =</p>
      <p>|T P |
|T P ∪ F P |
11We consider all DBpedia v3.9 resources valid.
(1)
(2)
(3)
(4)
w
T
3
1
]
8
[
)
1
1
0
2
(
P
L
N
r
e
t
t
i
m
e ]
t
s 0
y 1
S [</p>
      <p>N R
la R rd E
rn E fo rN
tex iNw tan ttie</p>
      <p>w
] T
6 ,
[ ]</p>
      <p>4
R 1</p>
      <p>[
E
d
a
b
k a
c r t
ten lan eyb fso
ew xP H rco
T a i</p>
      <p>T</p>
      <p>I
U M I M
oog rany
,rG tio
az ik
P
x
a</p>
      <p>T</p>
      <p>I
M I
t
f
o
s
o
r
c
i
P
P
N S</p>
      <p>U</p>
      <p>M
]
1 i’ re
[</p>
      <p>k p
’ i o</p>
      <p>l
B We
D ‘ v
‘ o e</p>
      <p>t d
o /
t d /:
d te s
te ia tp
a t
i v h
ev re ,
r b h
bb ab rca
a a e</p>
      <p>i
a S
i d
d ep leg
e i
p</p>
      <p>k o
B i o
D WG
a b c
The evaluation framework used in the challenge is available at https:
//github.com/giusepperizzo/neeleval.</p>
    </sec>
    <sec id="sec-10">
      <title>3.2 Evaluation Results</title>
      <p>Table 4 reports the performance of participants’ systems, using the
best run for each. The ranking is based on the F1.
System 18 clearly outperformed other systems, with F1 more than
15% higher than the next best system. System 18 differed from all
other systems, by using a joint approach to the NEEL task. The
others each divided the task into a sequential entity extraction and
linking task. The approach in System 18 made use of features which
capture jointly an entity’s local and global contextual information,
resulting in the best approach submitted to the #Microposts2014
NEEL Challenge.</p>
    </sec>
    <sec id="sec-11">
      <title>4. CONCLUSIONS</title>
      <p>The aim of the #Microposts2014 Named Entity Extraction &amp;
Linking Challenge was to foster an open initiative that would encourage
participants to develop novel approaches for extracting and linking
entity mentions appearing in Microposts. The NEEL task involved
the extraction of entity mentions in Microposts and the linking of
these entity mentions to DBpedia resources (where such exist).
Our motivation for hosting this challenge is the increased
availability of third-party entity extraction and entity linking tools. Such
tools have proven to be a good starting point for entity linking,
even for Microposts. However, the evaluation results show that the
NEEL task remains challenging when applied to social media
content with its peculiarities, when compared to standard length text
employing regular language.</p>
      <p>As a result of this challenge, and the collaboration of annotators
and participants, we also generated a manually annotated data set,
which may be used in conjunction with the NEEL evaluation
framework (neeleval). To the best of our knowledge this is the largest
publicly available data set providing entity/resource annotations for
Microposts. We hope that both the data set and the neeleval
framework will facilitate the development of future approaches in
this and other such tasks.</p>
      <p>The results of this challenge highlighted the relevance of
normalization and time-dependent features (such as popularity) for dealing
with this type of progressively changing content. It also indicated
that learning entity extraction and linking as a joint task may be
beneficial for boosting performance in entity linking in Microposts.
We aim to continue to host additional challenges targeting more
complex tasks, within the context of data mining of Microposts.</p>
    </sec>
    <sec id="sec-12">
      <title>5. ACKNOWLEDGMENTS</title>
      <p>The authors thank Nikolaos Aletras for helping to set-up the
crowdflower experiments for annotating the gold standard data set. A
special thank to Andrés García-Silva, Daniel Preo¸tiuc-Pietro, Ebrahim
Bagheri, José M. Morales del Castillo, Irina Temnikova, Georgios
Paltoglou, Pierpaolo Basile and Leon Derczynski, who took part in
the annotation tasks. We also thank the participants who helped us
improve the gold standard. We finally thank the LinkedTV project
for supporting the challenge by sponsoring the prize for the
winning submission.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          , G. Kobilarov,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          , and
          <string-name>
            <surname>Z. Ives.</surname>
          </string-name>
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>In 6th International Semantic Web Conference (ISWC'07)</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bollacker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Paritosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sturge</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          . Freebase:
          <article-title>A collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In ACM SIGMOD International Conference on Management of Data (SIGMOD'08)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. E. Cano</given-names>
            <surname>Basave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stankovic</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making Sense of Microposts (#MSM2013) Concept Extraction Challenge</article-title>
          .
          <source>In Making Sense of Microposts (#MSM2013) Concept Extraction Challenge</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          and W.-T. Yih.
          <article-title>Dual coordinate descent algorithms for efficient large margin structured prediction</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          and
          <string-name>
            <given-names>U.</given-names>
            <surname>Scaiella</surname>
          </string-name>
          .
          <article-title>Fast and accurate annotation of short texts with Wikipedia pages</article-title>
          .
          <source>IEEE Software</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <fpage>70</fpage>
          -
          <lpage>75</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Finkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grenager</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Incorporating non-local information into information extraction systems by Gibbs sampling</article-title>
          .
          <source>In 43rd Annual Meeting on Association for Computational Linguistics (ACL'05)</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <article-title>Greedy function approximation: A gradient boosting machine</article-title>
          .
          <source>Annals of Statistics</source>
          ,
          <volume>29</volume>
          :
          <fpage>1189</fpage>
          -
          <lpage>1232</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Gimpel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Mills</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Eisenstein</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Heilman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Yogatama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Flanigan</surname>
            , and
            <given-names>N. A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Part-of-speech tagging for Twitter: Annotation, features, and experiments</article-title>
          .
          <source>In 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoffart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Berberich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Lewis-Kelham</surname>
          </string-name>
          , G. de Melo, and
          <string-name>
            <surname>G. Weikum.</surname>
          </string-name>
          <article-title>YAGO2: Exploring and querying world knowledge in time, space, context, and many languages</article-title>
          .
          <source>In 20th International Conference Companion on World Wide Web (WWW'11)</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Datta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>TwiNER: Named entity recognition in targeted Twitter stream</article-title>
          .
          <source>In 35th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '12)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García-Silva</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>DBpedia spotlight: shedding light on the web of documents</article-title>
          .
          <source>In 7th International Conference on Semantic Systems (I-Semantics'11)</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Wordnet: A lexical database for English</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O.</given-names>
            <surname>Owoputi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gimpel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schneider</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Improved part-of-speech tagging for online conversational text with word clusters</article-title>
          .
          <source>In NAACL</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Clark</surname>
          </string-name>
          , Mausam, and
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          .
          <article-title>Named entity recognition in tweets: An experimental study</article-title>
          .
          <source>In EMNLP</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          .
          <article-title>NERD: A framework for unifying named entity recognition and disambiguation extraction tools</article-title>
          .
          <source>In 13th Conference of the European Chapter of the Association for computational Linguistics (EACL'12)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hellmann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Bruemmer</surname>
          </string-name>
          .
          <article-title>NERD meets NIF: Lifting NLP extraction results to the Linked Data Cloud</article-title>
          .
          <source>In Proceedings of the 5th International Workshop on Linked Data on the Web (LDOW'12)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V. I.</given-names>
            <surname>Spitkovsky</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. X.</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <article-title>A cross-lingual dictionary for English Wikipedia concepts</article-title>
          .
          <source>In LREC</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Turian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ratinov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Word representations: A simple and general method for semi-supervised learning</article-title>
          .
          <source>In 48th Annual Meeting of the Association for Computational Linguistics (ACL'10)</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thrasher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Viegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.-j. P.</given-names>
            <surname>Hsu</surname>
          </string-name>
          .
          <article-title>An overview of Microsoft Web N-gram corpus and applications</article-title>
          .
          <source>In NAACL HLT 2010 Demonstration Session (HLT-DEMO'10)</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Svore</surname>
            , and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gao</surname>
          </string-name>
          .
          <article-title>Adapting boosting for information retrieval measures</article-title>
          .
          <source>Information Retrieval</source>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          ):
          <fpage>254</fpage>
          -
          <lpage>270</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>