<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Datasets and GATE Evaluation Framework for Benchmarking Wikipedia-Based NER Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Milan Dojchinovski</string-name>
          <email>milan.dojchinovski@fit.cvut.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomas Kliegr</string-name>
          <email>tomas.kliegr@vse.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information and Knowledge Engineering Faculty of Informatics and Statistics University of Economics</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Web Engineering Group Faculty of Information Technology Czech Technical University in Prague</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present a wiki er evaluation framework consisting of software support and two datasets (News and Tweets), which were derived from datasets previously published at WEKEX 2011 and MSM Challenge 2013. Entities recognized in the original datasets were enriched with new annotations { a link to Wikipedia and the most speci c type from the DBpedia Ontology. The annotations were created by two annotators and a judge. The datasets are supplemented by plugins for their import to the GATE NLP framework and a DBpedia Ontology-aware plugin for aligning annotations created by a wiki er with the ground truth.</p>
      </abstract>
      <kwd-group>
        <kwd>Named Entity Recognition and Classi cation</kwd>
        <kwd>Benchmark</kwd>
        <kwd>Wikipedia</kwd>
        <kwd>DBpedia</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Wiki ers are systems that recognize entities in text and assign them URLs of
Wikipedia articles that describe these entities. Some wiki ers also assign types
from a taxonomy, making their output comparable with that of Named Entity
Recognition (NER) systems. We identify two elementary tasks that the wiki er
performs: i) disambiguation (linking of entities to Wikipedia articles, ii)
assignment of ne-grained types.</p>
      <p>In this paper, we present a framework for evaluation of these two tasks
consisting of two datasets, News and Tweets, and three plugins for the GATE text
engineering platform.1 DBpedia Ontology was selected as the set of types for
the ne-grained classi cation. We have found this choice natural, due to its wide
adoption and the fact that a mapping between its classes and the proprietary</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://gate.ac.uk/</title>
      <p>types output by many entity classi cation systems, such as OpenCalais2,
AlchemyAPI3 and Zemanta4 is provided by the NERD ontology.5</p>
      <p>The reminder of the paper is structured as follows. Section 2 reviews available
resources for wiki er evaluation. Section 3 describes the contributed News and
Tweets datasets. Section 4 presents the evaluation framework. Section 5 covers
availability and licenses. Finally, Section 6 provides a concluding summary.
2</p>
      <sec id="sec-2-1">
        <title>Datasets</title>
        <p>
          Currently, there is a lack of resources for wiki er evaluation, since those
previously created for benchmarking of NER systems cannot be directly used. While
some datasets have been recently contributed, in particular WEKEX6 or MSM7,
these do not contain all the necessary features for automated wiki er evaluation.
The WEKEX dataset contains manual evaluation of results obtained via the
NERD framework [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] from multiple common wiki ers. This provides a very
useful benchmark of the involved systems at the point in time when the assessment
was performed. However, the design of the dataset does not foster its
straightforward reuse for a new evaluation. Another recent dataset is the MSM challenge
dataset aiming at evaluation of coarse-grained type assignment, and the NIST
TAC Entity Linking contest dataset8, evaluating entity disambiguation. Only the
latter provides a comprehensive set of resources for the evaluation of wiki ers
(disambiguation only). Unfortunately, this dataset is available for purposes of
the TAC contest. None of the listed datasets supports evaluation of ne-grained
classi cation of entities.
        </p>
        <p>In this work, we take up two recently published datasets, the WEKEX dataset
and the MSM Challenge 2013 dataset and extend them to t the needs of
Wikipedia-based entity linking and classi cation, creating the News dataset, and
the Tweets dataset. The two datasets are complementary in that the WEKEX
dataset consists of a small number of standard-length news articles, while the
MSM datasets contains a large number of very short texts (tweets). Table 1 gives
an overview of the size of both datasets.
2.1</p>
        <sec id="sec-2-1-1">
          <title>WEKEX Dataset</title>
          <p>
            The 2011 paper [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] presents an evaluation of common entity recognition systems
using the NERD framework. In this evaluation there were two tasks. We
created the News dataset from the rst task's data. In this task, four participants
rated the entities output by the individual systems for ten English news articles
2 http://www.opencalais.com/
3 http://www.alchemyapi.com/
4 http://www.zemanta.com/
5 http://nerd.eurecom.fr/ontology/
6 http://nerd.eurecom.fr/ui/evaluation/wekex2011-goldenset.tar.gz
7 http://oak.dcs.shef.ac.uk/msm2013/ie_challenge/
8 http://www.nist.gov/tac/2013/KBP/EntityLinking/
selected from the on-line archives of the BBC and The New York Times. These
articles were from ve di erent categories. For each entity, a Wikipedia link, if
available, and the assigned type, were assessed.
          </p>
          <p>
            Since each entity recognition tool recognized a slightly di erent set of entities,
the set of all distinct entities identi ed by the benchmarked systems in [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] were
considered for the News dataset.
          </p>
          <p>A limitation of the original WEKEX dataset is a restricted copyright for the
underlying textual content. While this textual content is freely available from the
BBC's9 and NYTimes's10 o cial websites, it cannot be distributed along with
the annotations. The WEKEX dataset is released under the Creative Commons
BY-SA 3.0 license. Links to the original content are listed on the dataset website.
The Making Sense of Microposts (MSM) Challenge 2013 dataset aimed at
classifying entities in microposts (tweets). There are four entity types considered
corresponding to the standard CoNLL categories: Person, Organization,
Location, and Miscellaneous. The dataset is split into two parts { training and test
data. Both contained already recognized entities in the text. The entities in the
training dataset have the types already assigned, while the types are missing
in the test dataset. The organizers also published the goldstandard for the test
dataset containing the correct entity types. The original MSM dataset is
provided under the Creative Commons BY-NC-SA license.</p>
          <p>To construct the Tweets dataset, we used the tweets in the goldstandard that
contained at least one entity, resulting in 1044 tweets (1523 entities).
3</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>News and Tweets datasets</title>
        <p>The News and Tweets datasets were created by partial reannotation and
enrichment of the WEKEX and MSM datasets. The newly created datasets match the
needs of automated wiki er evaluation.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>9 http://www.bbc.com/ 10 http://www.nytimes.com/</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Annotation Guidelines</title>
        <p>The WEKEX and MSM datasets were reannotated to a (nearly) common set of
elds. The original version of both datasets already provided entity recognition.
The annotators thus worked on the same set of entities, providing for each entity:
{ URL to English Wikipedia: a URL of an article describing the entity,
{ Fine-grained type: a class from DBpedia Ontology 3.8,
{ Coarse-grained type: a CoNLL category (only for WEKEX)11
{ Most frequent sense ag: 1 if the correct Wikipedia page is found as the
rst hit of Wikipedia search for the entity name, otherwise empty.</p>
        <p>For the News dataset, we considered as entity candidates each entity output
by any of the systems that generated the original annotations in the WEKEX
dataset. To deal with this broader scope and lower quality, several speci c
annotation elds were added to the News dataset:
{ Common entity: 1 if the entity is not a named entity,
{ Full name: if this speci c entity is a part of a full entity name, which
appears in the article, then this elds lists the full entity name,
{ Partial: 1 if the recognized string is a part of the entity name, and this part
does not appear as a full standalone reference to the entity in the document.
The entities for which the result of the annotation process was \not an entity"
were removed.</p>
        <p>The MSM dataset contains high-quality recognition of entities (with the de
nition of entity being narrowed to the named entity), therefore entity recognition
can be reused in the Tweets dataset. There was just one problem related to
entity recognition, which was the frequent incorrect letter casing characteristic for
tweets. For the Tweets dataset, there was thus one dataset-speci c eld added:
{ Incorrect capitalization: 1 if there is at least one letter in the entity name
with incorrect case.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Annotation process</title>
        <p>Each entity was independently annotated by two annotators. If their annotations
matched in the speci c eld (e.g. link to English Wikipedia) this annotation was
automatically merged to the ground truth. The annotators were instructed to
provide explanation if they felt unsure. When there was no match, another
annotator, or in particularly spurious cases two annotators, resolved the con ict, using
the explanations from the rst-round annotators. The interannotator agreement
after the rst round (between the two annotators) for the core elds is given in
Table 2.</p>
        <p>The composition of annotators was as follows: one undergraduate computer
science student, two graduate computer science students, one post-doc
specializing on ontology alignment (all non-native English speakers). None of the authors
took part in the annotation process.
11 The MSM dataset already contains this information.
To facilitate the use of the newly created Tweets and News datasets, we have
developed three plugins for the GATE Text Engineering framework.</p>
      </sec>
      <sec id="sec-3-3">
        <title>The NewsCorpusBuilderPR and TweetsCorpusBuilderPR plugins load</title>
        <p>
          the datasets into GATE. It is assumed that the wiki er being benchmarked is
also wrapped as GATE plugin, creating GATE annotations on entities recognized
in the documents, with annotation features corresponding to the entity type and
Wikipedia URL. We provide a reference implementation of such a plugin for the
Targeted Hypernym Discovery12 entity classi cation system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>Once the wiki er has been run, the correct recognition of entities and of the
assignment of Wikipedia URLs can be evaluated by the standard GATE means.
We recommend using the GATE Corpus Quality Assurance tool. The evaluation
of the assigned DBpedia Ontology Type needs to be performed in an
ontologyaware fashion, which is not supported by this GATE tool. Consider the case,
when the ground truth type for an entity is dbpedia:VicePrimeMinister and
the benchmarked tool assigns dbpedia:Person. While the string-based
comparison performed by the Quality Assurance tool would mark such annotation as
incorrect, actually the dbpedia:Person is correct with respect to ground truth,
albeit more generic.</p>
        <p>The developed OntologyAwareFeatureDi PR performs comparison of
the assigned types taking into account the hierarchy of the DBpedia Ontology.
The plugin also assigns entities a new feature matchtype, which has either of
the following values:
{ exact match: ground truth and wiki er types match,
{ supertype: the ground truth annotation is a super-class of the assigned class,
{ subtype: the ground truth annotation is a sub-class of the assigned class,
{ nomatch: neither of the above.</p>
        <p>In case of supertype/subtype match, the plugin also uses the distance feature
to denote the length of the path between the subtype and supertype classes.
Note that the DBpedia Ontology, seen as a taxonomy tree, does not contain
cycles and the path length can be thus easily computed.</p>
        <p>The plugin also creates a new feature aligned-type, which is set to the
common supertype (if exists) of the ne-grained type assigned by the wiki er and
ground truth ne-grained type. Consider e.g. the following example. For a given
12 http://entityclassifier.eu/
entity, the ground truth annotation contains type dbpedia:VicePrimeMinister,
while the wiki er assigns type dbpedia:Person. The plugin will create the
aligned-type feature and set it to dbpedia:Person on both the ground truth
and wiki er annotations. This will allow the native GATE Corpus Quality
Assurance tool to evaluate this annotation as correct, while the matchtype feature
holds the detail type of the match.
5</p>
        <sec id="sec-3-3-1">
          <title>Availability and License</title>
          <p>The reannotated WEKEX (News) and the MSM (Tweets) datasets are available
online at http://entityclassifier.eu/datasets/evaluation/benchmark-datasets/.
The News dataset does not contain the source texts, which need to be obtained
from the BBC and NYTimes websites.</p>
          <p>Same licenses are used as for the original datasets: Creative Commons BY-SA
license for News and Creative Commons BY-NC-SA license for Tweets.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>The evaluation framework { the NewsCorpusBuilderPR, TweetsCor</title>
        <p>pusBuilderPR and the OntologyAwareFeatureDi PR plugin, together with
reference implementation of a GATE PR plugin providing annotations from the
entityclassifier.eu wiki er, are available online at http://entityclassifier.
eu/datasets/evaluation/tools/. These plugins are provided under the GNU
Lesser General Public License 3.0.
6</p>
        <sec id="sec-3-4-1">
          <title>Conclusion</title>
          <p>While there are several ground truth datasets for evaluation of \classic" NER
systems, a freely obtainable dataset for evaluation of Wikipedia-based NER
systems does not, to the best of our knowledge, exist. In this paper we proposed an
framework and two complementary datasets for benchmarking Wwiki ers, which
will hopefully signi cantly reduce the e ort required to perform the evaluation.
Sample results of wiki er benchmarks are available on the framework's website.
The authors will include links or information on additional results obtained with
the News and Tweets datasets if noti ed.</p>
          <p>Acknowledgements. This research was supported by the European Union's
7th Framework Programme via the LinkedTV project (FP7-287911) and CTU
in Prague grant (SGS13/100/OHK3/1T/18).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Dojchinovski</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Kliegr</surname>
          </string-name>
          .
          <article-title>Entityclassi er.eu: Real-time classi cation of entities in text with Wikipedia</article-title>
          . In H. Blockeel,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kersting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nijssen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Zelezny</surname>
          </string-name>
          , (eds.)
          <source>Machine Learning and Knowledge Discovery in Databases</source>
          , vol.
          <volume>8190</volume>
          of Lecture Notes in Computer Science, pp.
          <volume>654</volume>
          {
          <fpage>658</fpage>
          . Springer Berlin Heidelberg,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          . NERD:
          <article-title>Evaluating named entity recognition tools in the web of data</article-title>
          .
          <source>In ISWC'11, Workshop on Web Scale Knowledge Extraction (WEKEX'11)</source>
          ,
          <source>October 23-27</source>
          ,
          <year>2011</year>
          , Bonn, Germany.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>