<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Giving a Sense: A Pilot Study in Concept Annotation from Multiple Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roman Sudarikov</string-name>
          <email>sudarikov@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondrˇej Bojar</string-name>
          <email>bojar@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University in Prague Faculty of Mathematics and Physics Institute of Formal and Applied Linguistics Malostranské námeˇstí 25</institution>
          ,
          <addr-line>11800 Praha 1</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>88</fpage>
      <lpage>94</lpage>
      <abstract>
        <p>We present a pilot study in web-based annotation of words with senses coming from several knowledge bases and sense inventories. The study is the first step in a planned larger annotation of “grounding” and should allow us to select a subset of these “dictionaries” that seem to cover any given text reasonably well and show an acceptable level of inter-annotator agreement.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Annotated resources are very important for training,
tuning or evaluating many NLP tasks. Equipped with
experience in treebanking, we now move to resources for word
sense disambiguation (WSD) and entity linking (EL). By
EL, we mean the task of attaching a unique ID from some
database to occurrences of (named) entities in text [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Both entity linking and word-sense disambiguation have
been extensively studied, see for example [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2–4</xref>
        ]. Although
only a few researches consider several knowledge bases
and sense inventories at once [
        <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
        ], the convergence
between these two task is apparent, for example, the 2015
SemEval Task 13 promoted research in the direction of
joint word sense and named entity disambiguation [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>We understand the terms ontology, knowledge base and
sense inventory in the following way:
• Ontology is a formal representation of a domain of
knowledge. It is an abstract entity: it defines the
vocabulary for a domain and the relations between
concepts, but an ontology says nothing about how that
knowledge is stored (as physical file, in a database,
or in some other form), or indeed how the knowledge
can be accessed.
• Knowledge base is a database, a repository of
information that can be accessed and manipulated in some
predefined fashion. Knowledge is stored in
knowledge base according to an ontology.
• Sense inventory is a database, often build based on a
corpus, and providing clustered senses for the words
or expressions in the corpus.</p>
      <p>However, we recognize the blending of knowledge bases
and sense inventories, so we will use very generic terms
dictionary or resource interchangeably for either of them.</p>
      <p>
        In this pilot study, we examine several such dictionaries
in terms of their coverage and annotator agreement.
Unlike other works on “grounding”, which try to link only
the most important words in the sentence [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], we aim at
complete coverage of a given text, i.e. all content words or
multi-word expressions regardless their part of speech or
role in the sentence. Some of the examined resources have
a clear bias towards some parts of speech, for example,
valency dictionaries cover only verbs. We nevertheless ask
our annotators to annotate even across parts of speech if
the matching POS is not included in the resource. For
instance, verbs can get nominal entries in Wikipedia and
nouns get verb frames.1
      </p>
      <p>In Section 2, we describe the sense inventories included
in our experiment. Section 3 provides a unifying view on
these sources and introduces our annotation interface. We
conducted two experiments with English and Czech texts
using the interface, slightly adapting interface for the
second run. Details are in Section 4 and Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Resources Included</title>
      <p>Sense inventories and knowledge bases are plentiful and
they differ in many aspects including the domain coverage,
level of detail, frequency of update, integration of other
resources and ways of accessing them. Some of them
implement Resource Description Framework, the metadata data
model designed by W3C for the better data representation
in Semantic Web, while others are simply collections of
links in the web.</p>
      <p>We selected the following subset of general resources
for our experiment:</p>
      <p>
        BabelNet [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a multilingual knowledge base,
which combines several knowledge resources including
Wikipedia, Wordnet, OmegaWiki and Wiktionary. The
sources are automatically merged and accessible via
offline Java API or online REST API. An added benefit is
the multilinguality of BabelNet: the same resource can
be used for genuine (as opposed to cross-lingual)
annotation for both languages of our interest, English and Czech.
      </p>
      <p>
        1The conversion of nouns to predicates whenever possible is
explicitly demanded in some frameworks, e.g. in Abstract Meaning
Representation (AMR, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]).
      </p>
      <p>The main limitation is that BabelNet is not updated
continuously, so we also added both live Wikipedia and
Wiktionary as separate sources. BabelNet provides
information about nouns, verbs, adjectives and adverbs, but as
stated above, we are interested also in cross-POS
annotation.</p>
      <p>Wikipedia2 is currently the biggest online encyclopedia
with live updates from (hundreds of) thousands of
contributors so it can cover new concepts very quickly. Wikipedia
tries to nest all possible concepts as nouns. For
example, en.wikipedia.org/wiki/funny redirects to
the page “Humour”.</p>
      <p>Wiktionary3 is a companion to Wikipedia that covers
all parts of speech. It includes multilingual thesaurus,
phrase books, language statistics. Each word in
Wiktionary can have etymology, pronunciation, sample
quotations, synonyms, antonyms and translations, for better
understanding of the word.</p>
      <sec id="sec-2-1">
        <title>PDT-VALLEX and EngVallex (Valency lexicons for</title>
        <p>
          Czech and English): Valency or subcategorization
lexicons formally capture verb valency frames, i.e. their
syntactic neighborhood in the sentence [
          <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
          ]. We use the
valency lexicons for Czech and English in their offline
XML form as distributed with the tree editor TrEd 2.04.
        </p>
        <p>Google Search5 (GS): From our preliminary
experiments, we had the impression that no resource covers all
2http://wikipedia.org
3http://wiktionary.org
4http://ufal.mff.cuni.cz/tred/
5http://google.com
expressions seen in our data, but searching the web
provides some explanation almost always. We thus include
the top ten results returned by Google Search as a special
kind of dictionary, where the “concept” is a query string
and each result is considered to be its’ “sense”.</p>
        <p>Aside from coverage and frequency of updates, another
reason to include GS is that it provides “senses” at a very
different level of granularity than others. For instance, the
whole Wiktionary page can appear as one of the options in
GS “senses”. It will also often be a very sensible choice,
despite it actually covers several different meanings of the
word.</p>
        <p>We find the task of matching senses coming from
different ontologies and providing a different angle of view
or granularity very interesting. The current experiments
serve as a basis for its further investigation.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Annotation Interface</title>
      <p>To provide a unified view on the various resources, we use
the terms query, selection list and selection. Given an
expression in a text, which can be a word or a phrase, even a
non-continuous one, and a resource which should be used
to annotate it, the system construct a query. Querying the
resource, we get a selection list, i.e. a list of possible
senses.</p>
      <p>The process of extracting the selection list depends on
the resource. It is straightforward for Google Search (each
result becomes an option) and complicated for Wiktionary,
Source
Babelnet
Google Search
CS Vallex
EN Vallex
CS Wikipedia
EN Wikipedia
CS Wiktionary
EN Wiktionary
Babelnet
Google Search
EN Vallex
EN Wikipedia
EN Wiktionary
see Section 3.1 below. In principle and to include any
conceivable resource, even field-specific or ad hoc ones, the
annotator should be free to select the selection list prior to
the annotation.</p>
      <p>Our annotation interface allows to overwrite the query
for cases where the automatic construction does not lead
to a satisfactory selection list.</p>
      <p>Finally, the annotator is presented with the selection list
to make his choice (or multiple choices). Overall, the
annotator picks one of these options:
Whole Page means that the current URL is already a
good description of the sense and no selection list
is available on the page. The annotators were asked
to change the query and rather obtain a selection list
(e.g. a disambiguation page in Wikipedia) whenever
possible.</p>
      <p>Bad List means that the extraction of selection list failed
to provide correct senses. The annotators were
supposed to try changing the query to obtain a usable list
and resort to the “Bad List” option only if inevitable.
None indicates that the selection list is correct but that it
lacks the relevant sense.</p>
      <sec id="sec-3-1">
        <title>One or more senses selected is the desired annotation:</title>
        <p>The list, for the particular pair of selected word(s) and
selected resource, was correct and the annotator was
able to find the relevant sense(s) in the list.</p>
        <p>Our annotation interface (Figure 1) shows the input
sentence, tabs for individual sense inventories, the selection
list from the current resource and also the complete page
where the selection list comes from. The procedure is
straightforward: (1) select one or more words in the
sentence using checkboxes, (2) select a resource (we asked
our annotators to use them all, one by one), (3) check if
the selection list is OK and modify the query if needed,
(4) make the annotation choice by marking one or more of
the checkboxes in the selection list, and (5) save the
annotation.
3.1</p>
      </sec>
      <sec id="sec-3-2">
        <title>Queries and Selection Lists for Individual</title>
      </sec>
      <sec id="sec-3-3">
        <title>Resources</title>
        <p>This is how we construct queries and extract selection lists
for each of our dictionaries given one or more words from
the annotated sentence:
BabelNet We search BabelNet for the lemma of the
selected word (or the phrase of lemmas if more words
are selected). The selection list is the list of all
obtained BabelNet IDs.</p>
        <p>Google Search We search for the lemmas of the selected
words and return the snippets of the top ten results.</p>
        <p>The selection list is the list of snippets’ titles.</p>
        <p>Wikipedia We search for the disambiguation page for the
selected words and, if not found, we search for the
page with the title matching the lemmas of the
selected words. The selection list for disambiguation
pages is constructed by fetching hyperlinks
appearing within listings nested in particular HTML blocks.
For other pages we fetch links from the Table of
Contents and the first hyperlink from each listing item.
Wiktionary We search for the page with the title equal to
the lemmas of the selected words. The selection list
is created using the same heuristics as for Wikipedia.
Vallex We scan the XML file and return all the frames
belonging to the verb with the lemma matching the
selected word’s lemma.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>First Experiment</title>
      <p>The first experiment was held in March 2014. The 7
participating annotators (none of whom had any
experience in annotation tasks) were asked to annotate the
sentences from PCEDT 2.0 6 with Czech and English sources:
Wikipedia and Wiktionary for both languages, BabelNet,
6http://ufal.mff.cuni.cz/pcedt2.0/en/index.
html</p>
      <p>Google Search, and the Czech and English Vallexes. Each
annotator was given a set of sentences in English or Czech
and they were asked to annotate as many words or phrases
in each sentence as possible, with as many reasonable
meanings as they can. We required the annotators to
annotate across parts of speech if possible (for instance to
annotate the noun “teacher” with the corresponding verb “to
teach”). This requirement appeared because we wanted to
evaluate the possibility of using more abstract senses as
used, for instance, in works with AMR.
4.1</p>
      <sec id="sec-4-1">
        <title>Gathered Annotations</title>
        <p>In total, we collected 507 annotations for 158 units. 75 of
these units had more than one annotation.</p>
        <p>The upper part of Table 1 provides details on how
often each of the annotation options was picked for a given
source in the first annotation experiment. Note that in the
first experiment, we did not offer the “Whole Page” option.</p>
        <p>We see that the sources exhibit slightly different patterns
of use. Wikipedia has lots of “Bad List” options selected
due to the issue described in Section 4.2. GS is the most
ambiguous resource, the user has picked two or more sense
in about one half of GS annotations. The highest
number of “Bad Lists” was received by the English Wikipedia
(18 out of 40).</p>
        <p>Figure 2 shows the distribution of different POS per
source. Google Search seems to be the most versatile
resource, covering all parts of speech well. The relatively
low use of BabelNet was due to the web API usage limit.
Vallexes work well for verbs but cross-POS annotation is
only an exception. Wikipedia and Wiktionary are indeed
somewhat complementary in covered POSes.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Bad List vs. None Issue</title>
        <p>The “Bad List” annotations should be used in two cases:
(1) when the system fails to extract the selection list from
Source
Babelnet
GS
CS Vallex
EN Vallex
CS Wikipedia
EN Wikipedia
CS Wiktionary
EN Wiktionary
Total:
a good page, and (2) when the whole page is wrong, for
example when the system shows the Wikipedia page “South
Africa” for the word “south”. “None” was meant for
correct selection lists (matching domain, reasonable options)
but the right option missing. The guidelines for the first
experiment were not very clear on this so some
annotators marked problems with selection list as “Bad List” and
some used the label “None”.</p>
        <p>Manual revision revealed that only 10 out of 40 “Bad
List” annotations were indeed “Bad List” in one of the two
meanings described above. The right hand part of Table 2
shows IAA after changing wrongly annotated “Bad Lists”
into “None”.
4.3
Inter-annotator agreement is a measure of how well two
annotators can make the same annotation decision for
a certain item. In our case it is measured as the
percentage of cases when a pair (2-IAA) of annotators agree on
the (set of) senses for a given annotation unit. The
measurement was made pairwise for all the annotations, which
had more that one annotator. The results are presented in
Table 2, before and after fixing the “Bad List” issue.</p>
        <p>In general, the IAA estimates should be treated with
caution. Many units were assigned only to a single
annotator, so they weren’t taken into account while computing
IAA.</p>
        <p>The extremely low IAA for English Wikipedia was
caused by the following issue. For several units, one
annotator tried to select all the senses to show that the whole
page can be used, while others have picked one or only
a few senses. We resolved the issue by introducing a new
option “Whole page” in the second experiment.</p>
        <p>Interestingly, we see a negative correlation (Pearson
correlation coefficient of -0.37) between the number of
units annotated for a given source and the 2-IAA.</p>
        <p>
          We also report Cohen’s kappa [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], which reflects the
agreement when disregarding agreement by chance. In our
setting, we estimate the agreement by chance as one over
the length of the selection list plus two (for “None” and
“Bad List”). This is a conservative estimate, in principle
the annotators were allowed to select any subset of
selection list. We compute kappa using K = P1a−−PPee , where Pa
was the total 2-IAA and Pe was the arithmetical average of
agreements by chance for each annotation. Kappa for the
first experiment was 0.13.
        </p>
        <p>To assess the level of uncertainty for the estimates,
we use bootstrap resampling with 1000 resamples, which
gives us IAA of 0.25 ± 0.1 and kappa of 0.135 ± 0.115 for
95% of samples.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Second Experiment</title>
      <p>The second experiment was held in March 2015 with
another group of 6 annotators. One of the annotators had
experience in annotating tasks, while others had no such
experience. The setting of the experiment was slightly
different. The annotators were asked to annotate only English
sentences from QTLeap project7 using BabelNet, Google</p>
      <sec id="sec-5-1">
        <title>7http://qtleap.eu/</title>
        <p>Search, English Wikipedia, English Wiktionary and
ENGVALLEX. The guidelines were refined, asking the
annotators to mark the largest possible span for each concept
in the sentence, e.g. to annotate “mouse cursor” jointly as
one concept and not separately as “computer pointing
device” for the word “mouse” and “graphic representation of
computer mouse on the screen” for the word “cursor”. The
option “Whole page” was newly introduced to help users
indicate that the whole page can be used as a sense.
5.1</p>
        <sec id="sec-5-1-1">
          <title>Gathered Annotations</title>
          <p>We collected 570 annotations for 35 words, 32 of which
had annotations from more than one annotator. The
number of units here is lower that in the first experiment,
because all our annotators used the same sentences. Also,
for the second experiment we required the annotators to
use all the resources for each unit, so we have more results
per unit.</p>
          <p>During the second experiment, the system processed
147 unique (in terms of selected word(s) and selected
resource) queries. All the resources got nearly equal
number of queries (about 30), except for Vallex, which got
only 10 queries. The annotators changed the queries
59 times, but this also includes cases, when Wikipedia
used its own inner redirects, which our system did not
distinguish from users’ changes. BabelNet was changed
9 times, Google Search – 2, Vallex – 8, Wikipedia – 21
and Wiktionary – 19. Based on these numbers, GS may
seem more reliable but it is not necessary true. One reason
is that some of part of the changes for Wikipedia was made
automatically by Wikipedia itself. The other argument is
that users could limit their effort and after examining the
first 10 GS results for the query they just picked “Bad List”
option and moved on, not trying to change the query.</p>
          <p>The POS per source distribution (see Figure 3) for the
second experiment is similar to the first one, except for the
Source
Babelnet
GS
Vallex
Wikipedia
Wiktionary
Total content words
BabelNet, which did not reach any technical limit this time
and was therefore used more often across all POSes.
In Table 3, we show the coverage of content words in the
second experiment. By content words we mean all the
words in the sentence, except for auxiliary verbs,
punctuation, articles and prepositions. The instructions asked
to annotate all content words. Each annotator completed
a different number of sentences, so the number of words
annotated differs. The column Content words Attempted
shows the total number of words with some annotation at
all, while Labeled are words which received some sense,
not just “None” or “Bad List”. Both numbers are taken
from the union over all annotators. Babelnet get the best
coverage in terms of Labeled annotations. The right hand
side of the table shows how many words each annotator
has labeled. Since the union is considerably higher than
the most productive annotator, we need to ask an
important question: How many annotators do we need to have
a perfect coverage of the sentence.
5.3</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Inter-Annotator Agreement</title>
          <p>Results presented in Table 4 are overall better than in the
first experiment. The kappa was computed as in
Section 4.3 with the only one difference: we added 3 instead
of 2 options when estimating the local probability of the
agreement by chance (for the new “Whole Page” option).
Kappa for the second experiment was 0.40. Bootstrapping
showed IAA 0.39 ± 0.055 and kappa 0.32 ± 0.06 for 95%
central resamples. Again, the 2-IAA is negatively
correlated with the number of units annotated (Pearson
correlation coefficient -0.22).
6</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>Comparing first and second experiment, one can see, that
we managed to improve IAA by expanding the set of
available options and refining the instructions, but IAA is still
not satisfactory.</p>
      <p>For resources where IAA reaches 60% (Vallex and
Wikipedia), the coverage is rather low, 26% and 58%.
BabelNet gives the best coverage but suffers in IAA. Google
Search seems an interesting option for its versatility across
parts of speech, on par with established knowledge bases
like BabelNet in terms of inter-annotator agreement but
with much more ambiguous “senses”. The cross-POS
annotation does not seem very effective in practice, but
a more thorough analysis is desirable.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Comparison with Other Annotation Tools</title>
      <p>Several automatic systems for sense annotation are
available. Our dataset could be used to compare them
empirically on the annotations from the respective repository
used by each of the tools. For now we provide only an
illustrative comparison of these three systems: TAGME8,
DBpedia Spotlight9 ,and Babelfy10</p>
      <p>Figure 4 provides an example of our manually collected
annotations for the sentence “Move the mouse cursor to
the beginning of the blank page and press the DELETE
key as often as needed until the text is in the desired spot.”.</p>
      <p>For this sentence, the TAGME system with default
settings returned three entities (“mouse cursor”, “DELETE
key” and “text”). DBpedia Spotlight with default settings
(confidence level = 0.5) returned one entity (“mouse”).
Babelfy showed the best result among these systems in
terms of coverage, failing to recognize only the verb
“move” and adverbs “often” and “until”, but it also
provided several false meanings for found entities.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>In this paper, we examined how different dictionaries can
be used for entity linking and word sense disambiguation.</p>
      <sec id="sec-8-1">
        <title>8http://tagme.di.unipi.it/ 9http://dbpedia-spotlight.github.io/demo/ 10http://babelfy.org/</title>
        <p>bn:00062699n</p>
        <p>Wikipedia
Motion_(physics)
blank_page_(disambiguation)
blank_page_(disambiguation)
press_(disambiguation)
Delete_key, DELETE
Delete_key, key_(disambiguation)
often
Need_(disambiguation)
until
text_(disambiguation)
Desire_(disambiguation), desired
spot_(disambiguation)
TAGME
mouse
cursor
beginning
blank
page
press
DELETE
key
often
needed
until
text
desired
spot
Mouse_(computing)</p>
        <p>Mouse_(computing)
bn:00024529n,bn:00021487n
mouse_cursor, cursor_(disambiguation)
beginning, beginning_(disambiguation)
In our unifying view based on finding the best “selection
list” and selecting one or more senses from it, we tested
standard inventories like BabelNet or Wikipedia, but also
Google Search.</p>
        <p>We proposed and refined annotation guidelines in two
consecutive experiments, reaching average inter-annotator
agreement of about 46%, with Wikipedia and Vallex up
to 60%. Higher agreement seems to go together with lower
coverage, but further investigation is needed for
confirmation and to find the best balance of granularity, coverage
and versatility among existing sources.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>This research was supported by the grants
FP7-ICT-201310-610516 (QTLeap). This research was partially
supported by SVV project number 260 224. This work has
been using language resources developed, stored and
distributed by the LINDAT/CLARIN project of the Ministry
of Education, Youth and Sports of the Czech Republic
(project LM2010013).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Demartini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , et al.:
          <article-title>Zencrowd: leveraging probabilistic reasoning and crowdsourcing techniques for large-scale entity linking</article-title>
          .
          <source>In: Proceedings of the 21st international conference on World Wide Web, ACM</source>
          (
          <year>2012</year>
          )
          <fpage>469</fpage>
          -
          <lpage>478</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bennett</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          , et al.
          <source>: Report on the sixth workshop on exploiting semantic annotations in information retrieval (ESAIR'13)</source>
          .
          <source>In: ACM SIGIR Forum</source>
          . Volume
          <volume>48</volume>
          ., ACM (
          <year>2014</year>
          )
          <fpage>13</fpage>
          -
          <lpage>20</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ratinov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , et al.:
          <article-title>Local and global algorithms for disambiguation to wikipedia</article-title>
          .
          <source>In: Proc. of ACL/HLT</source>
          , Volume
          <volume>1</volume>
          . (
          <year>2011</year>
          )
          <fpage>1375</fpage>
          -
          <lpage>1384</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Word sense disambiguation: A survey</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>41</volume>
          (
          <issue>2</issue>
          ) (
          <year>February 2009</year>
          )
          <volume>10</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          :
          <fpage>69</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Entity linking with multiple knowledge bases: An ontology modularization approach</article-title>
          .
          <source>In: The Semantic Web-ISWC 2014</source>
          . Springer (
          <year>2014</year>
          )
          <fpage>513</fpage>
          -
          <lpage>520</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Moro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.: SemEval-2015
          <source>Task</source>
          <volume>13</volume>
          :
          <string-name>
            <given-names>Multilingual</given-names>
            <surname>All-Words Sense</surname>
          </string-name>
          Disambiguation and
          <string-name>
            <given-names>Entity</given-names>
            <surname>Linking</surname>
          </string-name>
          .
          <source>In: Proc. of SemEval-2015</source>
          . (
          <year>2015</year>
          ) In press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scaiella</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Tagme: on-the-fly annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In: Proc. of CIKM</source>
          , ACM (
          <year>2010</year>
          )
          <fpage>1625</fpage>
          -
          <lpage>1628</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rettinger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Färber</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Tadic´,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>A comparative evaluation of cross-lingual text annotation techniques</article-title>
          .
          <source>In: Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Visualization. Springer (
          <year>2013</year>
          )
          <fpage>124</fpage>
          -
          <lpage>135</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Banarescu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , , et al.:
          <source>Abstract Meaning Representation for Sembanking</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.:</given-names>
          </string-name>
          <article-title>BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>193</volume>
          (
          <year>2012</year>
          )
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopatková</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Valency information in VALLEX 2.0: Logical structure of the lexicon</article-title>
          .
          <source>The Prague Bulletin of Mathematical Linguistics</source>
          (
          <volume>87</volume>
          ) (
          <year>2007</year>
          )
          <fpage>41</fpage>
          -
          <lpage>60</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Lopatková</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ketnerová</surname>
          </string-name>
          , V.:
          <article-title>Valencˇní slovník cˇeských sloves</article-title>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A Coefficient of Agreement for Nominal Scales</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          <volume>20</volume>
          (
          <issue>1</issue>
          ) (
          <year>1960</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>