<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>REDIT: A Tool and Dataset for Extraction of Personal Data in Documents of the Public Administration Domain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Teresa Paccosi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trento (Italy) [tpaccosi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>aprosio]@fbk.eu</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>English. New regulations on transparency and the recent policy for privacy force the public administration (PA) to make their documents available, but also to limit the diffusion of personal data. The present work displays a first approach to the extraction of sensitive data from PA documents in terms of named entities and semantic relations among them, speeding up the process of extraction of these personal data in order to easily select those which need to be hidden. We also present the process of collection and annotation of the dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Le nuove regolamentazioni sulla
trasparenza e la recente legislazione sulla
privacy hanno spinto la pubblica
amministrazione a rendere i loro documenti
pubblicamente consultabili limitando pero` la
diffusione di dati personali. Presentiamo
qui un primo approccio all’estrazione di
questi dati da documenti amministrativi in
termini di named entities e relazioni
semantiche tra di esse, in modo da facilitare
la selezione dei dati che devono rimanere
privati. Presentiamo inoltre il processo di
collezione e annotazione del dataset.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>In recent years, public administrations (PA) in
the Italian government have been forced to
publish a huge amount of documents, to make them
available to citizens, organisations, and
authorities. This is the result of the recent legislation</p>
      <p>Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
about the transparency. For instance,
municipalities have to share their documents in a virtual place
called Albo Pretorio. In most cases, the online
publication of these acts is a necessary condition
for their purposes to become effective.1</p>
      <p>On the other side, the General Data Protection
Regulation (GDPR), approved in 2016 by the
European Union, enhances individuals’ control and
rights over their personal data, limiting its
diffusion over any medium (especially including online
platforms such as websites and social networks).</p>
      <p>In this context, it is important for the public
servants within the PA to amend some documents by
hiding the data that cannot be publicly published.
Nowadays, most of this work is done manually,
hiding the sensitive information document by
document. This procedure is clearly time-consuming,
non-scalable, and error-prone.</p>
      <p>Natural Language Processing (NLP) techniques
can be seen as a watershed between a manual
management of the PA documents and a new
generation of instruments that will finally speed up the
process, leaving manual effort as the sole final
check just before the publication of the data.</p>
      <p>
        This is not the first time this problem is tackled
using NLP, but past works are mainly focused on
English and limited to the entity extraction task
        <xref ref-type="bibr" rid="ref11">(Guo et al., 2021)</xref>
        .
      </p>
      <p>
        Our approach to the extraction of personal data
from documents focuses on a combination of three
NLP instruments:
• Named-entity Recognition (NER). This
task consists in seeking texts in natural
language to locate and classify named entities
(NE) mentioned in them. This search is
usually limited to a few needed categories: the
most common are persons, locations, and
organisations. Several approaches have been
1In the Italian legislation, this is called “referto di
pubblicazione”. See also: http://qualitapa.gov.it/
used in literature, between completely
rulebased
        <xref ref-type="bibr" rid="ref1 ref4">(Appelt et al., 1993; Budi and
Bressan, 2003)</xref>
        and machine learning-based
        <xref ref-type="bibr" rid="ref19 ref6 ref8">(Chiu
and Nichols, 2016; Strubell et al., 2017;
Devlin et al., 2019)</xref>
        , including some hybrid
approaches, for example using gazettes of
known entities belonging to a particular
category
        <xref ref-type="bibr" rid="ref9">(Finkel et al., 2005)</xref>
        . In this paper, we
use the last approach, mixing a Conditional
Random Fields (CRF) algorithm
        <xref ref-type="bibr" rid="ref13">(Lafferty et
al., 2001)</xref>
        with the addition of a list of
entities, extracted from various knowledge bases,
that describe persons, companies, and
locations. We describe this process in detail in
Section 5.
• Structured-entity Identification. A
parallel rule-based task is used to extract
entities that can easily be recognised without the
need of training data. Among them: dates
and times, numbers, email addresses,
Italian “codice fiscale”, that are based on textual
patterns; roles and document types, that are
based on prepacked lists.
• Relation extraction (RE). It is the task of
extracting semantic relationships from text.
Extracted relationships usually occur between
two or more entities of a certain type (for
example persons, locations, etc., see
previous points), and fall into a number of
semantic categories (such as birth location, role
in a company, etc.). Relation extraction is
widely used also in specific domains such as
medicine
        <xref ref-type="bibr" rid="ref10">(Giuliano et al., 2007)</xref>
        and finance
        <xref ref-type="bibr" rid="ref21">(Vela and Declerck, 2009)</xref>
        . Successful
experiments made use of Conditional Random
Fields
        <xref ref-type="bibr" rid="ref20">(Surdeanu et al., 2011)</xref>
        ,
Dependencybased Neural Networks
        <xref ref-type="bibr" rid="ref14">(Liu et al., 2015)</xref>
        , and
transformers like BERT
        <xref ref-type="bibr" rid="ref3">(Baldini Soares et
al., 2019)</xref>
        .
      </p>
      <p>
        In this paper we present REDIT (Relation and
Entities Dataset for Italian with Tint), a complete
framework that aims to solve the personal data
identification in textual documents. The software
is mainly based on Tint
        <xref ref-type="bibr" rid="ref18">(Palmero Aprosio and
Moretti, 2018)</xref>
        , an NLP pipeline specifically
designed for Italian and based on Stanford CoreNLP
        <xref ref-type="bibr" rid="ref15 ref5">(Manning et al., 2014)</xref>
        . REDIT includes part of the
annotated dataset (Section 4), the compiled model,
and the supporting Java code. It is available for
free on Github (see Section 6).
      </p>
      <p>The content is structured as follows. Section 2
presents in detail how we collected the documents
that are annotated and how we used fictitious data
to make the resource available for download. In
Section 3, we describe the process used to
annotate the data. Section 4 illustrates the dataset,
giving some statistics on the entities and relations
included in it. In Section 5 we give some results on
the performance of the resulting entity extraction
and relation extraction system. The downloadable
package (that contains the dataset, the model and
the Java code) is finally described in Section 6.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Data Collection</title>
      <p>The corpus is composed of documents taken from
different institutions of the public administration.
The documents with which we have worked are
different types of forms, varying from license for
parking to adoption forms, school enrollments,
marriage licenses and so on.</p>
      <p>Starting from this set, we create two datasets.
One is composed of documents compiled with real
data and one with documents compiled by us with
ifctitious data, using lists of all the Italian streets
and surnames in order to guarantee the
diversification of the data in the compiled forms, and to not
exclusively rely on the annotators’ fantasy. The
ifctitious compilation aims to avoid using
sensitive data in terms of privacy issues, leading to the
possibility of publicly releasing the dataset. The
documents which contain real data are indeed not
included in the public dataset. For instance, a
sentence such as Il sottoscritto Gianluca Freschi, nato
a Pesaro il 12/12/1990 e residente in Pesaro, Via
Virgilio n.76 presents data whose association was
invented by the annotator. It could be possible
that a person called Gianluca Freschi exists in real
world but it is almost impossible that he would fit
with the rest of the data since they all derive by
annotator’s fantasy. However, as we can see from the
example, while the data are fictitious the structure
of the document is identical to that of real ones.
3</p>
    </sec>
    <sec id="sec-4">
      <title>The Annotation</title>
      <p>Each document in the set is annotated both with
entities and relations between them.</p>
      <p>
        For the annotation of entities we adopt the
guidelines already used for KIND
        <xref ref-type="bibr" rid="ref11 ref17">(Paccosi and
Palmero Aprosio, 2021)</xref>
        , a corpus containing NE
on documents taken from Wikinews. The named
entities included in KIND belong to the standard
NE classes and are of three types: LOC, PER,
and ORG. As already noticed by (Passaro et al.,
2017), these categories are quite unsatisfactory
to deal with the information contained in the PA
documents, since the model is not designed at
capturing information such as laws or protocols.
In REDIT we then distinguish different types of
ORGs, differentiating public offices and
municipality and companies: the former is annotated as
ENTE, while the latter as usual (ORG). Finally,
we add a label to mark laws and protocols, LEX,
so that in the present work there are vfie types
of annotated entities: LOC, PER, ORG, LEX,
and ENTE. The original guidelines used in KIND
have therefore been slightly modified to meet our
needs (see Section 5.1 for more details).
      </p>
      <p>In addition to the NE annotation, we are
interested in annotating the relations among them. In
particular, we need to develop a system of
relations which links the person with its personal data
or with its role in terms of responsibility of the
company/public administration or in terms of
relative/family relationships.</p>
      <p>Since for the annotation task a relation must
connect two entities, some additional entity types
are annotated only when involved in a relation (see
below). The list of additional entities includes
ROLE for personal and organisation roles (for
example, words such as “responsabile”, “titolare”,
“genitore”, and so on, representing the role of a
person in a company, in the PA domain, or in a
family), DOCTYPE for document types (such as
“passaporto”, “patente”), EMAIL for e-mail
addresses, DATE for dates, NUMBER for generic
numbers (such as VAT), CF for the Italian “codice
ifscale” sequence of chars.</p>
      <p>Regarding relations, address is used for
instance to link a LOC entity representing an address
to the person or company to which the address
belongs, while birthDate, birthLoc link
respectively the date and location of birth.</p>
      <p>Table 1 shows the complete list of the relations
included in the dataset.</p>
      <p>
        The annotation is performed by a domain
expert using INCEpTION
        <xref ref-type="bibr" rid="ref12">(Klie et al., 2018)</xref>
        , a
webbased text-annotation environment which allows
users to: (i) select a group of tokens and assign
a label to it (entities); (ii) connect two entities
among them and assign a label to the link
(relations).
      </p>
      <p>This is an example of NER annotation:
Al [Comune di Alessandria]ENTE.
[Casale Monferrato]LOC, 20 settembre
2021.</p>
      <p>Il sottoscritto [Davide Aiello]PER, nato
a [Milano]LOC il [31/07/1985]DATE,
[titolare]ROLE della ditta [Aiello
Ceramiche S.r.l.]ORG, ai sensi dell’ [art.
76 del D.P.R. n. 445/2000]LEX, dichiara
di voler partecipare all’evento “Il
mercante in Fiera”.</p>
      <p>These are the corresponding relations:
• birthLoc (Davide Aiello, Milano)
• birthDate (Davide Aiello, 31/07/1985)
• companyRole (Davide Aiello, titolare)
• personInOrg (Davide Aiello, Aiello
Ceramiche S.r.l.)</p>
      <p>In the example, “31/07/1985” is tagged as
DATE, since it is involved in the birthDate
relation. On the contrary, since no relations include
“20 settembre 2021”, it’s not mandatory, for the
annotator, to mark it as DATE.</p>
      <p>The system uses two different approaches to
identify entities. Entities such as DATE or ROLE
are annotated only when involved in a relation
because they are labels identified through a
rulebased approach which can be easily recognised
without the need of training data. For what
concerns instead PER, LOC, ORG, ENTE and LEX
the identification occurs using a machine-learning
technique and they need to be always annotated.
4</p>
    </sec>
    <sec id="sec-5">
      <title>The Dataset</title>
      <p>As we have seen in Section 2, the complete dataset
consists of two parts: the first one presents the
documents fictitiously compiled and it is publicly
released; the latter, on the contrary, comprehends
instances compiled with real data and is not
released. Nevertheless, we consider also the
unreleased dataset in training the model, so that the
amount of annotated relations in the final dataset
is 7,821, while that of annotated entities is 21,307.
The released one presents 1,439 annotated entities
and 1,476 annotated relations.</p>
      <p>Looking at the data in Table 1, it is possible to
notice that the amount of annotations referring to
some relations (marked with *) are considerably
fewer than others. Despite the small amount, we
have already annotated them in the view of future
works on these relations but we do not consider
them in the experiments.</p>
      <p>Tint processing</p>
      <p>Tokenization
Named-entities recognition</p>
      <p>Dependency parsing
REDIT processing
Relation extraction
REDIT
dataset
Part-of-speech tagging
Rule-based extractors
Lemmatization
ENTE classifier
To work properly, REDIT relies on a complex
pipeline that includes various steps, very different
in structure and management (see Figure 1). Most
of the steps are performed using well-known tools
and algorithms (sometimes not reaching
state-ofthe-art accuracy), so that the whole program does
not need particular hardware (such as the GPUs
needed in environments using deep learning and
transformers) and is easy to run on almost every
common software environment.</p>
      <p>
        1. First, the input text is parsed with Tint
        <xref ref-type="bibr" rid="ref18">(Palmero Aprosio and Moretti, 2018)</xref>
        using
these annotators: tokenizer, sentence splitter,
truecaser, part-of-speech tagger, lemmatizer,
dependency parser.
2. Named-entities are extracted using the CRF
implementation included in Stanford NER
        <xref ref-type="bibr" rid="ref9">(Finkel et al., 2005)</xref>
        and the model trained on
the annotated dataset (see Subsection 5.1).
3. A second run on named-entities, with the
rule-based Stanford TokensRegex software
        <xref ref-type="bibr" rid="ref15 ref5">(Chang and Manning, 2014)</xref>
        , is performed
(see Subsection 5.2)
4. ORG and LOC entities are passed into a
Support Vector Machines classifier
        <xref ref-type="bibr" rid="ref7">(Cortes and
Vapnik, 1995)</xref>
        to extract ENTE entities (see
Subsection 5.3).
5. Finally, the Stanford Relation Extractor
        <xref ref-type="bibr" rid="ref20">(Surdeanu et al., 2011)</xref>
        is used to find relations
between entities in the text (see Subsection 5.4).
Since the sole REDIT dataset is not sufficient to
train a robust NER tagger, we use it in
combination with KIND (see Section 3). Guidelines for
the two datasets are, of necessity, slightly
different, therefore we need to use some precautions in
merging them.
      </p>
      <p>Sometimes, the entities annotated as ORG in
KIND (such as “Unione Europea”) should have
been annotated as ENTE in REDIT. We then
decided, in the training phase, to merge all ENTE
entities into ORG. We then trained a classifier
dedicated to the ENTE tag (see Subsection 5.3),
trained on REDIT dataset only, that performs the
sole disambiguation between ORG and ENTE.</p>
      <p>
        To enhance the classification, Stanford NER
also accepts gazettes of names labelled with the
corresponding tag. We collect a list of
persons, organizations and locations from the Italian
Wikipedia using some classes in DBpedia
        <xref ref-type="bibr" rid="ref2">(Auer
et al., 2007)</xref>
        : Person, Organisation, and
Place, respectively. In addition to this, we
collect the list of streets from OpenStreetMap
        <xref ref-type="bibr" rid="ref16">(OpenStreetMap contributors, 2017)</xref>
        , limiting the
extraction to Italian names. Table 3 shows statistics
about the gazettes.
      </p>
      <p>The evaluation is performed by randomly
splitting the dataset into train/dev/test using 80/10/10
ratio. During training phase, we tried some
sets of features choosing among the ones
available in Stanford NER. We obtained the best
results (considering also a good balance between
training/testing time and performances) with word
shapes, n-grams with length 6, previous, current,
and next token/lemma/class. Table 4 displays the
results of the NER module.
5.2</p>
      <sec id="sec-5-1">
        <title>The Rule-Based Named-Entities Tagger</title>
        <p>As said in Section 3, there is the need for more
entity types, because in the training phase we need
to have both arguments of a relation annotates as
an entity (of any type). For this reason, we use
Relation
LEX
LOC
ORG
PER
Total (micro)
Total (macro)
a rule-based approach to annotate DATE, ROLE,
DOCTYPE, EMAIL, NUMBER, and CF.
• Tint TIMEX annotator is used to tag DATE
entities.
• ROLE and DOCTYPE entities are extracted
given a list of roles taken from the annotated
training set.
• Numbers, e-mail addresses and Italian codice
ifscale are tagged using regular expressions.
5.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>The SVM Classifier for ENTE Entities</title>
        <p>After the previous steps, the entities that should
be marked with ENTE now falls into the ORG or
LOC entity sets. We then use a simple SVM
classifiers (using shallow features, such as words,
bigrams, previous and following content words, etc.)
that, given an entity tagged as LOC or ORG,
return whether it should be annotated as ENTE. The
training set used by the classifier consists in
entities taken from REDIT and annotated as ORG,
LOC, and ENTE. The first two categories
represent the zero class, while entities tagged with
ENTE represent the other class. It is therefore a
binary classifier. In a 10-fold cross-validation
environment, results shows a F-score equals to 0.978
(precision 0.981, recall 0.974).
5.4</p>
      </sec>
      <sec id="sec-5-3">
        <title>Relation Extractor Module</title>
        <p>
          The last module in REDIT is Stanford Relation
Extractor
          <xref ref-type="bibr" rid="ref20">(Surdeanu et al., 2011)</xref>
          , used to train and
extract relations in the text.
        </p>
        <p>Similarly to the NER training, we test
approaches with different sets of features, obtaining
the best results with unigrams/bigrams, adjacent
words, argument words, argument class,
dependency path between the arguments, entities and
concatenation of POS tags between arguments.</p>
        <p>
          Table 5 shows the results on the relation
extractor (the evaluation is performed using gold-labeled
entities).
All parts of REDIT (except part of the annotated
dataset, see Section 4) are released for free under
the CC BY 4.0 license,2 and can be downloaded
on Github.3 These include the annotations, in
WebAnno format
          <xref ref-type="bibr" rid="ref22">(Yimam et al., 2013)</xref>
          , the gazettes,
both the NER and the RE models (created using
the whole corpus), and the source code, written in
Java, used to parse the files and run the classifiers.
        </p>
        <p>A working demo of the tool is available online
(See Figure 2).4 Its web interface is written with
VueJS/Boostrap and it is available for download in
the Github project page.
7</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>In this paper we present a completely automatic
approach to extract personal data (view as entities)
and relations between them from documents of the
public administration written in Italian texts. The
pipeline relies on a mix of rule-based and machine
learning-base modules. The latter are trained
using a manually annotated dataset, which is in part
available for download. All the source code,
instead, is released and available for download.</p>
      <p>In the future, we plan to enhance the coverage of
our system by adding more examples on relations
that are less represented (see Table 1).</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The research leading to this paper was partially
supported by Wemapp Srl, Potenza, Italy.5
2https://bit.ly/cc-by-40-intl
3https://github.com/dhfbk/redit
4https://bit.ly/relation-extraction
5https://wemapp.eu/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Douglas</given-names>
            <surname>Appelt</surname>
          </string-name>
          , Jerry Hobbs, John Bear, David Israel,
          <string-name>
            <given-names>and Mabry</given-names>
            <surname>Tyson</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>FASTUS: A Finite-state Processor for Information Extraction from Realworld Text</article-title>
          .
          <source>In Proceedings of the International Joint Conference on Artificial Intelligence</source>
          , pages
          <fpage>1172</fpage>
          -
          <lpage>1178</lpage>
          ,
          <fpage>01</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>So¨ren Auer, Christian Bizer</article-title>
          , Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and
          <string-name>
            <given-names>Zachary</given-names>
            <surname>Ives</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>In Karl Aberer</source>
          ,
          <string-name>
            <surname>Key-Sun</surname>
            <given-names>Choi</given-names>
          </string-name>
          , Natasha Noy,
          <string-name>
            <surname>Dean Allemang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kyung-Il</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Lyndon Nixon, Jennifer Golbeck,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Mika</surname>
          </string-name>
          , Diana Maynard, Riichiro Mizoguchi, Guus Schreiber, and
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Cudre</surname>
          </string-name>
          ´- Mauroux, editors,
          <source>The Semantic Web</source>
          , pages
          <fpage>722</fpage>
          -
          <lpage>735</lpage>
          , Berlin, Heidelberg. Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Livio</given-names>
            <surname>Baldini</surname>
          </string-name>
          <string-name>
            <surname>Soares</surname>
          </string-name>
          , Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski.
          <year>2019</year>
          .
          <article-title>Matching the blanks: Distributional similarity for relation learning</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>2895</fpage>
          -
          <lpage>2905</lpage>
          , Florence, Italy, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>I.</given-names>
            <surname>Budi</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bressan</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Association rules mining for name entity recognition</article-title>
          .
          <source>In Proceedings of the Fourth International Conference on Web Information Systems Engineering</source>
          ,
          <year>2003</year>
          .
          <source>WISE</source>
          <year>2003</year>
          ., pages
          <fpage>325</fpage>
          -
          <lpage>328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Angel</surname>
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
            and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>TokensRegex: Defining cascaded regular expressions over tokens</article-title>
          .
          <source>Technical Report CSTR 2014-02</source>
          , Department of Computer Science, Stanford University.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Jason P. C.</surname>
          </string-name>
          <article-title>Chiu</article-title>
          and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Nichols</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Named entity recognition with bidirectional lstm-cnns.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Corinna</given-names>
            <surname>Cortes</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Supportvector networks</article-title>
          .
          <source>Machine learning</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ):
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Jenny</given-names>
            <surname>Rose</surname>
          </string-name>
          <string-name>
            <surname>Finkel</surname>
          </string-name>
          , Trond Grenager, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Incorporating non-local information into information extraction systems by Gibbs sampling</article-title>
          .
          <source>In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05)</source>
          , pages
          <fpage>363</fpage>
          -
          <lpage>370</lpage>
          , Ann Arbor, Michigan, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Giuliano</surname>
          </string-name>
          , Alberto Lavelli, Daniele Pighin, and
          <string-name>
            <given-names>Lorenza</given-names>
            <surname>Romano</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>FBK-IRST: Kernel methods for semantic relation extraction</article-title>
          .
          <source>In Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007)</source>
          , pages
          <fpage>141</fpage>
          -
          <lpage>144</lpage>
          , Prague, Czech Republic, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Yongyan</given-names>
            <surname>Guo</surname>
          </string-name>
          , Jiayong Liu,
          <string-name>
            <given-names>Wenwu</given-names>
            <surname>Tang</surname>
          </string-name>
          , and Cheng Huang.
          <year>2021</year>
          .
          <article-title>Exsense: Extract sensitive information from unstructured data</article-title>
          .
          <source>Comput. Secur.</source>
          ,
          <volume>102</volume>
          :
          <fpage>102156</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Jan-Christoph</surname>
            <given-names>Klie</given-names>
          </string-name>
          , Michael Bugert, Beto Boullosa, Richard Eckart de Castilho, and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The inception platform: Machine-assisted and knowledge-oriented interactive annotation</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations</source>
          , pages
          <fpage>5</fpage>
          -
          <lpage>9</lpage>
          . Association for Computational Linguistics, Juni.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>John D. Lafferty</surname>
          </string-name>
          ,
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          , and
          <string-name>
            <surname>Fernando</surname>
            <given-names>C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data</article-title>
          .
          <source>In Proceedings of the Eighteenth International Conference on Machine Learning, ICML '01</source>
          , pages
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          , San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Yang</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Furu Wei,
          <string-name>
            <given-names>Sujian</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Heng</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming Zhou</surname>
            , and
            <given-names>Houfeng</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A dependency-based neural network for relation classification</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)</source>
          , pages
          <fpage>285</fpage>
          -
          <lpage>290</lpage>
          , Beijing, China, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Christopher D. Manning</surname>
            , Mihai Surdeanu, John Bauer, Jenny Finkel,
            <given-names>Steven J.</given-names>
          </string-name>
          <string-name>
            <surname>Bethard</surname>
          </string-name>
          , and
          <string-name>
            <surname>David McClosky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations</article-title>
          , pages
          <fpage>55</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>OpenStreetMap contributors</source>
          .
          <year>2017</year>
          .
          <article-title>Planet dump retrieved from https://planet</article-title>
          .osm.org . https://ww w.openstreetmap.org.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Teresa</given-names>
            <surname>Paccosi</surname>
          </string-name>
          and Alessio Palmero Aprosio.
          <year>2021</year>
          .
          <article-title>KIND: an Italian Multi-Domain Dataset for Named Entity Recognition</article-title>
          .
          <source>In arXiv preprint.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Palmero</surname>
          </string-name>
          Aprosio and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Tint 2.0: an all-inclusive suite for nlp in italian</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it</source>
          , volume
          <volume>10</volume>
          , page 12.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Emma</given-names>
            <surname>Strubell</surname>
          </string-name>
          , Patrick Verga, David Belanger, and
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Fast and accurate entity recognition with iterated dilated convolutions</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Surdeanu</surname>
          </string-name>
          ,
          <string-name>
            <surname>David McClosky</surname>
            ,
            <given-names>Mason</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>Andrey</given-names>
          </string-name>
          <string-name>
            <surname>Gusev</surname>
            , and
            <given-names>Christopher</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Customizing an information extraction system to a new domain</article-title>
          .
          <source>In Proceedings of the ACL 2011 Workshop on Relational Models of Semantics</source>
          , pages
          <fpage>2</fpage>
          -
          <lpage>10</lpage>
          , Portland, Oregon, USA, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Mihaela</given-names>
            <surname>Vela</surname>
          </string-name>
          and
          <string-name>
            <given-names>Thierry</given-names>
            <surname>Declerck</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Concept and relation extraction in the finance domain</article-title>
          . In H. Bunt,
          <string-name>
            <given-names>V.</given-names>
            <surname>Petukhova</surname>
          </string-name>
          , and S. Wubben, editors,
          <source>Proceedings of the Eighth International Conference on Computational Semantics (IWCS-8)</source>
          .
          <source>International Conference on Computational Semantics (IWCS-8)</source>
          , January 7-9, Tilburg, Netherlands, pages
          <fpage>346</fpage>
          -
          <lpage>351</lpage>
          . Tilburg University, 1.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Seid</given-names>
            <surname>Muhie</surname>
          </string-name>
          <string-name>
            <surname>Yimam</surname>
          </string-name>
          , Iryna Gurevych, Richard Eckart de Castilho, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Biemann</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>WebAnno: A flexible, web-based and visually supported system for distributed annotations</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , Sofia, Bulgaria, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>