<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IRISA System for Entity Detection and Linking at CLEF HIPE 2020</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cheikh Brahim El Vaigh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guillaume Le Noe-Bienvenu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guillaume Gravier</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pascale Sebillot</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>INRIA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IRISA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France cheikh-brahim.el-vaigh@inria.fr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IRISA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France guillaume.le-noe-bienvenu@irisa.fr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IRISA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France guig@irisa.fr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>INSA Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IRISA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France pascale.sebillot@irisa.fr</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This note describes IRISA's system for the task of named entity processing on historical newspapers in French. Following a standard entity detection and linking pipeline, our system implements three steps to solve the named entity linking task. Named Entity Recognition (NER) is rst performed to identify the entity mentions in a document based on a Conditional Random Fields classi er. Candidate entities from Wikidata are then generated for each mention found, using simple search. Finally, every mention is linked to one of its candidate entities in a so-called linking step leveraging various string metrics and the semantic structure of Wikidata to improve on the linking decisions.</p>
      </abstract>
      <kwd-group>
        <kwd>Named entity recognition CRF Collective entity linking WRSM entity relatedness measure</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Entity linking is a core task in textual document processing, which consists in
identifying the entities of a knowledge base (KB) that are mentioned in a text.
For instance, approaches from the literature implement three stages to solve
mention ambiguity in texts. The rst stage consists in the detection of named entities
within the text and is known as named entities recognition (NER). To further link
the mention found in the text, candidate entities are generated for each mention
detected in the rst stage. Finally, every mention is linked to one of its candidate
entities in a so-called linking step. This last step can be performed independently
for each individual mention, or collectively for all mentions at once. In the rst
case, every mention in a text is assumed to be independent from other mentions
and is linked to a candidate entity on sole basis of some similarity between the
mention and the candidate entities, so-called local scores. By contrast, for
collective entity linking, entity mentions and the corresponding entities are not assumed
independent one from another but somehow semantically related within a
(coherent) document, i.e., mention-to-entity linking decisions are interdependent. In this
case, the local mention-entity scores are complemented with global scores re ecting
to which extent the candidate entities chosen for the mentions under consideration
are related in the KB, according to a so-called entity relatedness measure. The last
two stages of the pipeline are also known as named entity linking (NEL).</p>
      <p>
        In the context of the shared task CLEF HIPE 2020 |Identifying Historical
People, Places and other Entities|which is a named entity processing on historical
newspapers in French, German and English [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], entity linking techniques can be
used to retrieve entities from text. CLEF HIPE 2020 is organised as a CLEF 2020
evaluation Lab. However, the historical context makes the linking task harder since
texts considered are the results of an optical character recognition (OCR)
algorithm which introduces noise. Therefore, we leveraged various features to reduce
the impact of the OCR *noise* on named entity processing.
      </p>
      <p>
        Our system for CLEF HIPE 2020 follows a standard pipeline for entity linking
and implements three separate stages:
1. We devise a NER stage on top of the baseline provided by CLEF HIPE 2020
organizers. This system used Conditional Random Fields (CRFs) to detect and
classify named entities. We added several features that we found e ective for
the task of NER.
2. The generation step consists in looking to Wikidata directly when searching
entities similar to a given mention. As a lookup in the heavy database (CLEF
HIPE 2020 Wikidata dump) is costly in time, we performed automatic searches
for the entity mentions using online Wikidata. Note that this search is based
on Wikidata indexing algorithm.
3. The linking step is to decide which candidate should be retained for each
mention within a document. We tried to link the mentions separately or collectively,
training a classi er to predict if a mention is related to one of its candidate
entities. The former is based solely on the similarity between a given mention and
its candidate entities. The latter which performs the linking collectively for all
the mentions at once, beside the previous similarity metrics, makes use of the
entity relatedness measure WSRM that we have proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4. The collective linking setup gave the best results and was ranked second for the
bundle2 of the shared task CLEF HIPE 2020.
      </p>
      <p>Our source code, datasets and experimental results are made available online for
reproducibility purposes5.</p>
      <p>The note is organized as follows. We give the description of our method in Sec. 2.
Then we group the experimental results in Sec. 3 before discussing the perspectives
and conclude in Sec. 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System Architecture</title>
      <p>
        This section gives the description of our system. We distinguish two independent
tasks for named entity processing, namely the NER and the NEL. Our solutions
for the NER and the NEL are described respectively in Sec. 2.1 and Sec. 2.2.
5 https://gitlab.inria.fr/celvaigh/hipe2020
2.1
The NER task aims at detecting the surface forms in a text that correspond to
named entities and at classifying those forms as a type (PER, LOC, ORG, TIME,
PROD). The NER system that we developed originally came from the NER
baseline provided by the organizing team of the evaluation campaign CLEF HIPE
2020. This system used Conditional Random Fields (CRFs) based on a Python
implementation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to detect and classify named entities. The features used in this
system, as well as the ones we have chosen are described in Table 1.
2.2
      </p>
      <p>
        NEL
In the NEL stage, we assume that the annotations are known for the mentions
(person, organization, location, etc.) for each document. Those annotations are
provided by the NER system described in Sec. 2.1 or by an oracle NER. For the
candidate generation stage we rely on a simple Wikidata web search. The
candidate selection stage accounts for the WSRM entity relatedness measure between
candidate entities within the document in an e cient manner, relying here on
Wikidata, the KB provided by CLEF HIPE 2020 for named entity processing
(see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for details on the measure). These di erent steps are described below.
Candidate Entities Generation To generate candidate entities from the KB
for each mention in a document, we chose a simple yet e cient method exploiting
the index of Wikidata. For each mention found by the NER phase, we perform
online search using Wikidata web pages. We limit ourselves to the top 10 ranked
candidate entities. The motivation behind our choice is to speed up the candidate
entities generation step as a lookup in the heavy Wikidata dump is costly in time
compared to simple web search.
      </p>
      <p>Local Scores The local scores depict the similarity between a mention and its
candidate entities. If we assume the mentions to be independent in text, the linking
problem can be formalized as
e^= argmax (m;ei) (1)</p>
      <p>
        ei
where ei is a candidate entity, m is an entity mention, and is the local score
function. We tried several metrics for . Beside the longest contiguous matching
sub-sequence, we tried a Levenshtein distance to handle the OCR noise, Wikipedia
popularity [
        <xref ref-type="bibr" rid="ref2 ref5">2,5</xref>
        ] and the cosine similarity based on a word embedding model,
similar to the Skip-gram embedding model [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Collective Entity Linking In a collective NEL setup, the local score is
complemented with a global score accounting for the intricate interrelationships that
candidate entities of the di erent mentions may share. The latter is known as an
entity relatedness measure and used to assess entity relationships in the KB, which
will allow to estimate the interdependence of the mentions in the text. The CEL
problem can be thus formalized as</p>
      <p>
        0 n n
(e^1;:::;e^n) = argmax@X (mi;ei)+X
e1;:::;n
(2)
i=1 i=1j=1;j6=i
where n is the total number of mentions in a text and (ei;ej ) donates the
entity relatedness measure. In the collective linking version of our system, we used
the semantic entity relatedness measure WSRM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which weights the relation
between entities, where the more relations between the entities, the stronger their
relationship. Formally, is de ned between two entities ei and ej as
(ei;ej ) = Xjfjrfjr(0eji(;eri;;erj0);e20)K2BKgBjgj ; (3)
      </p>
      <p>e02E
where E denotes the set of entities in the KB and jSj the cardinality of the set S.</p>
      <p>Because the directions of the relations are somewhat arbitrary in KBs,
depending on how the relation vocabulary was designed (think about the publishes and
publishedBy symmetric RDF properties), we use a symmetric version of WSRM
de ned as</p>
      <p>
        1
(ei;ej ) = (WSRM(ei;ej )+WSRM(ej ;ei)) :
2
(4)
Using the NEL Output to Correct the NER Predictions We also exploited
the output of the NEL in order to enhance the NER results. First, we used the type
(obtained from Wikidata) of the entities retrieved by the NEL and forwarded it to
the NER stage, which can be updated accordingly. Then we leveraged WSRM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
to retrieve, for each entity found by the NEL, a list of potential related entities
from the KB. We argue that if an entity is mentioned in the text, its related
entities in the KB should be also mentioned in that text. Our aim is to exploit the
semantics of the KB for the NER task. Those information provided by the NEL
are used as pseudo-labels or features in the CRF to supervise the NER.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>Experimental validation was conducted on the CLEF HIPE 2020 French corpus
to assess the quality of our system. The dataset is described in Sec. 3.1. Results
for the NER are provided in Sec. 3.2, and in Sec. 3.3 for the NEL.</p>
      <p>Due to the high number of results given by the CLEF HIPE 2020 scorer, we
decided to focus only on a couple of them, that were given in the produced json
le: NE-COARSE-LIT - ALL - strict - F1 micro and NE-COARSE-LIT - ALL
ent type - F1 micro.
3.1</p>
      <sec id="sec-3-1">
        <title>Dataset</title>
        <p>The evaluation corpus is composed of newspaper articles sampled among several
Swiss, Luxembourgish and American historical newspapers on a diachronic basis.
This corpus is digitised based on an OCR algorithm which hightails the historical
context of the evaluation campaign. The time-span of the whole corpus goes from
1798 until 2018. We used only the French version of the corpus composed of a
train, a validation and a test sets 6.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results of the NER</title>
        <p>The NER classi er described in Sec. 2.1 is trained on the CLEF HIPE 2020 dataset.
We added several features to the ones of the baseline. We performed a random
search to select the best features while controlling the over tting. We provide the
list of the features used in Tab. 1 and the list of the best hyper-parameters in
Tab. 2. The system was trained on the train le and then tested on the dev and
test les provided by the organizers.</p>
        <p>We compared our NER system with the baseline provided by the CLEF HIPE
2020 organizers on the validation set (dev le). The results are gathered in Tab. 3.
We can see that our NER system outperforms the baseline. We believe that its
good performance is due to the choice of the selected features, e.g., the use of
the tokens present in the text as features for the classi er. The ne tuning of the
hyper-parameters of our CRF also partly explains the results better than those of
the baseline. The results of our system on the test le are gathered in Tab. 3
6 Details statistics about the data can be found at https://impresso.github.io/</p>
        <p>CLEF-HIPE-2020/datasets.html
Parameter
c1, the coe cient for L1 regularization, between 0 and 1
c2, the coe cient for L2 regularization, between 0 and 1
min freq, cut-o threshold for occurrence frequency
of a feature, between 0 and 1
max iterations, the maximum number of iterations
for optimization algorithm, between 100 and 1000
all possible transitions
num memories, the number of limited memories for
approximating the inverse Hessian matrix, between 4 and 8</p>
        <p>Table 2. The best parameters for the CRF.</p>
        <p>Best value found
0.1798
0.0551
192
false
0
4
Results of the entity linking process evaluated in terms of micro-averaged F1
classi cation scores are reported in Tab. 4. The three systems that we submitted to
CLEF HIPE 2020 were ranked second (team7 results). We rst evaluated the entity
linking based on the sole use of the local scores donated by team7 bundle2 fr 1.
Second, we added the global score devising a collective entity linking which we
named team7 bundle2 fr 2. And nally, we changed the collective linking system
to lter the non-linkable mentions (NIL) based on a threshold, meaning we only
link a mention to a candidate entity if the prior probability is below a xed
threshold (here 0.5). We can see that the collective linking gave the best results, while
the collective linking with a xed threshold is worse than the non-collective one.
These results show the bene t of the collective linking.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Supervising the NER with the NEL</title>
        <p>A few experiments have been carried out to exploit the outputs of the NEL in
order to enhance the NER results. The rst one consisted in using the types of the
entities found by the NEL to change the NER labels; e.g., if the NER detects the
entity 'Europe' and classi es it as 'PERS', the NEL links it to 'Q46' and gives the
information that the type of 'Q46' is 'LOC'. The second consisted in generating
closely related entities to the ones found by the NEL. We found that the output
of the NEL stage can correct the NER, but can also introduce too much noise.
Despite not being able to directly incorporate the output of the NEL with the
existing features, we believe that applying a major vote between the di erent
versions of the NER|with and without the NEL output|can lead to an increase of
the accuracy of the NER. Nonetheless, our system opened the door to incorporate
the semantics of the KB into the NER task.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>
        We built an entity processing system based on a CRF classi er for the NER task,
and a collective entity linking system for the NEL one, exploiting the WSRM
entity relatedness measure that we have proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Our system was evaluated
on the CLEF HIPE 2020 French dataset. Though initially expected, we did not
succeed in incorporating the output of the NEL to correct the NER step, but we
paved the way to fully use the KB semantics in the NER task.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Buitinck</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Louppe</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niculae</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grobler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Layton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , VanderPlas, J.,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holt</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
          </string-name>
          , G.:
          <article-title>API design for machine learning software: Experiences from the scikit-learn project</article-title>
          .
          <source>In: European Conference on Machine Learning and Principles and Practices of Knowledge Discovery in Databases Workshop: Languages for Data Mining and Machine Learning</source>
          . pp.
          <volume>108</volume>
          {
          <issue>122</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Durrett</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A joint model for entity analysis: coreference, typing, and linking</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>2</volume>
          ,
          <issue>477</issue>
          {
          <fpage>490</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ehrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romanello</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Fluckiger,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Clematide</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <source>Overview of CLEF HIPE</source>
          <year>2020</year>
          :
          <article-title>Named entity recognition and linking on historical newspapers</article-title>
          .
          <source>In: CLEF HIPE</source>
          <year>2020</year>
          . pp.
          <volume>1</volume>
          {
          <issue>25</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>El</given-names>
            <surname>Vaigh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.B.</given-names>
            ,
            <surname>Goasdoue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Gravier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Sebillot</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Using knowledge base semantics in context-aware entity linking</article-title>
          .
          <source>In: ACM Symposium on Document Engineering</source>
          <year>2019</year>
          . pp.
          <volume>8</volume>
          :
          <issue>1</issue>
          {8:
          <fpage>10</fpage>
          . Berlin, Germany (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Francis-Landau</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Durrett</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Capturing semantic similarity for entity linking with convolutional neural networks</article-title>
          .
          <source>In: 15th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          . pp.
          <volume>1256</volume>
          {
          <issue>1261</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>