<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extraction of Medical Concepts from Italian Natural Language Descriptions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrizia Agnello</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia Maria Ansaldi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Azzalini</string-name>
          <email>fabio.azzalini@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Gangemi</string-name>
          <email>giovanni.gangemi@mail.polimi.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Piantella</string-name>
          <email>davide.piantella@polimi.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Rabosio</string-name>
          <email>emanuele.rabosio@fht.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Letizia Tanca</string-name>
          <email>letizia.tanca@polimi.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>INAIL - Dipartimento Innovazioni Tecnologiche</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Human Technopole - Center for Analysis, Decisions and Society</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present a Natural Language Processing (NLP) pipeline to automatically extract medical concepts from a free text written in a language other than English. To do so, we use common NLP techniques and the metathesaurus of Unified Medical Language System (UMLS). Specifically, our goal is to automatically extract ontological concepts representing which part of the human body is injured and what is the nature of the injury, given an Italian textual description of a work accident. We start by partitioning the text into tokens and assigning to each token its part-of-speech, and then use an appropriate tool to extract relevant concepts to be searched within UMLS. We tested our system on a public large repository containing textual descriptions of work accidents produced by INAIL. Experimental results confirm that our system is able to correctly extract relevant medical concepts from texts written in Italian.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Ontology</kwd>
        <kwd>EHR</kwd>
        <kwd>NLP</kwd>
        <kwd>Work accident</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The term Electronic Health Records (EHRs) describes the concept of a comprehensive,
crossinstitutional, and longitudinal collection of healthcare data, trying to group the entire clinical
life of a patient [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. EHRs store information both in structured (e.g. diagnosis codes, laboratory
results, etc) and unstructured (e.g. clinical notes, discharge summaries, etc.) formats.
Unstructured data usually contain a more complete, and broader, view of the patient’s conditions, as well
as additional valuable information that would be dificult to represent in a structured manner
(e.g. social history, special conditions, etc.). To leverage all the advantages of a systematic
adoption of EHRs, many technical and non-technical requirements must be fulfilled [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], such as
privacy, data security, portability, performance, maintainability, reliability, interoperability, and
usability.
      </p>
      <p>In this work we focus on the natural-language texts included in EHRs, since unstructured
data analytics is one of the most challenging task of EHRs automated analysis. Our main
contributions are the development of a system capable of automatically extracting medical
ontological concepts from an Italian natural language text, and the experimental evaluation of
our system using a real-world dataset regarding work injuries.</p>
      <p>The paper is organized as follows. Section 1 gives an overview of the problem, Section 2
reviews the state of the art regarding medical concept extraction, Section 3 describes our
methodology, Section 4 presents the experiments and, finally, Section 5 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. State of the Art</title>
      <p>
        Concept extraction from natural language texts related to clinical information consists of three
phases [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]: (1.) Identification of concept mentions such as medications, drugs, anatomical parts,
and diseases; (2.) Coreference resolution regarding relationships between diferent mentions
referring to the same entity; and (3.) Extraction of relationships between concepts.
      </p>
      <p>We now briefly describe one of the most complete medical ontology system (UMLS) and two
state-of-the-art frameworks for clinical concept extraction: cTAKES and MetaMap.</p>
      <sec id="sec-2-1">
        <title>2.1. Unified Medical Language System</title>
        <p>
          The Unified Medical Language System (UMLS) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is a compendium of many controlled
vocabularies in the biomedical sciences, produced and distributed by the National Library of Medicine
(NLM). It also provides a mapping structure among these vocabularies and thus allows to
translate among the various terminology systems. It can be therefore considered a comprehensive
thesaurus and ontology of biomedical concepts.
        </p>
        <p>UMLS is composed by three modules: Metathesaurus, Semantic Network, and Specialist
Lexicon. We now provide a brief explanation of each of these modules.</p>
        <sec id="sec-2-1-1">
          <title>2.1.1. Metathesaurus</title>
          <p>The Metathesaurus of UMLS includes over one million biomedical concepts and five million
concept names, enclosing many vocabularies such as ICD-10, SNOMED CT, MeSH, and more.
The Metathesaurus is structured to facilitate the identification of synonyms between concepts,
also in diferent languages ensured by leveraging hierarchical concept identifiers, in turn linked
to:
• Concepts (CUI): identifying the meaning to be expressed, it never changes over time, no
matter the updates in the vocabularies.
• Strings (SUI): each string representing a concept is assigned with a permanent string
identifier. Any character variation (e.g. case sensitivity, punctuation, etc.) will result in a
diferent SUI, for each language. A SUI can in principle be linked to more than one CUI.
• Atoms (AUI): being building blocks of the Metathesaurus, atoms represent specific entries
in the vocabularies included in UMLS. An AUI is linked to one and only one CUI.
• Lexical terms (LUI): a lexical term comprises diferent strings (i.e. SUI) that are lexical
variants or minor variants. This layer is often used to reduce the computational complexity
of exploration and for a more efective concept lookup. It is currently available for all the
English strings, and only partially for other languages.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.1.2. Semantic Network</title>
          <p>The Semantic Network provides a consistent categorization of all concepts represented in the
Metathesaurus along with a set of useful relationships between these concepts. The network
contains 133 semantic types and 54 relationships. Each concept in the Metathesaurus is assigned
one or more semantic types, which are linked to one another through semantic relationships. The
major semantic types are organisms, anatomical structures, biologic function, chemical, events,
physical objects, etc.</p>
          <p>The possible relationships between semantic types range from simple ⟨is-a⟩ hierarchies
to complex associations, such as ⟨physically related to⟩, ⟨spatially related to⟩, ⟨co-occurs with⟩.
Relationships can be derived from associations already present in the vocabularies
(intrasource relationships) or they can connect concepts from diferent vocabularies ( inter-source
relationships), including not only synonyms. Inter-source relationships enhance the integration
of all the vocabularies present in UMLS and enable an easy exploration of the resulting ontology.
A subset of the Semantic Network with ⟨is-a⟩ relationships is shown in Figure 2.</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>2.1.3. Specialist Lexicon</title>
          <p>Specialist Lexicon is a module of UMLS that addresses the high degree of variability in natural
language words, allowing the abstraction of lexical variants. It covers general English lexicon
and many biomedical terms, including syntactic, morphological, and orthographic information.
Since only English words are covered by this module, we adopted a diferent approach for Italian
Natural Language Processing, which we will describe in Section 3.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. cTAKES</title>
        <p>
          Apache clinical Text Analysis and Knowledge Extraction System [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] (cTAKES, for short) is an
open-source framework for knowledge extraction from clinical texts, exploiting NLP techniques
including rule-based and machine learning approaches. It leverages a pipeline of six components.
First of all the text is divided into sentences by the sentence boundary detector, a component
which extends OpenNLP sentence detector [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Each sentence is then tokenized taking into
consideration also context-specific occurrences (e.g. dates, time intervals, etc.). Each token is
then normalized leveraging tools part of UMLS Specialist Lexicon (described in Section 2.1.3),
in order to map each token in a normalized form with respect to many lexical properties
(e.g. inflection, diacritics, symbols, stop words, etc.). Both the normalized and the original
occurrences are maintained for further analysis. After a part-of-speech tagging, the named
entity recognition annotator component performs a terminology-agnostic dictionary lookup on
a subset of UMLS Metathesaurus (described in Section 2.1.1), searching all the noun-phrases
identified and their respective unnormalized occurrences.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. MetaMap</title>
        <p>
          MetaMap was developed by the National Library of Medicine (NLM) with the goal of mapping
biomedical text to the UMLS Metathesaurus [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. It relies on a pipeline similar to cTAKES, apart
from the leveraging of relationships and hierarchical identifiers, present in UMLS Metathesaurus,
to better identify synonyms and lexical variants of the tokenized texts.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Comparison between cTAKES and MetaMap</title>
        <p>
          A comparison between cTAKES and MetaMap has already been investigated in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]: the results
of the experiments proved that cTAKES slightly outperforms MetaMap, with the exception
of texts in which abbreviations are present. It has been shown that abbreviations are quite
common in natural language texts written by doctors and both tools have dificulties in correctly
identify their correct meanings. With MetaMap, however, it is possible to partially solve this
problem, specifying a list of strings that will be treated as special tokens. This possibility is
not particularly investigated in the cited experimental comparison, and could be the subject of
future studies.
        </p>
        <p>The main disadvantage of both cTAKES and MetaMap, with respect to our study, is that they
are strongly English-centric, since they both rely on UMLS Specialist Lexicon which, as we
already described in Section 2.1.3, fully covers only the English language.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>We now present the methodology of our system. The final goal is to automatically extract
ontological concepts representing which part of the human body is injured and what is the
nature of the injury, given an Italian textual description of a work accident. Figure 3 displays
the three phases of our workflow: Part-of-Speech (POS) Tagging, Keyphrase Extraction and
Dictionary Lookup.</p>
      <p>
        The first phase receives as input a textual description representing the dynamic of an accident
and gives as output a preprocessed and enriched representation of the input text. Specifically,
Tint takes as input a raw text in Italian and performs a series of natural language processing
operations. Tint (The Italian NLP Tool) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is an open-source Java-based pipeline for Natural
Language Processing (NLP) in Italian. It is very fast and accurate, and implements most of the
common linguistic tools, such as part-of-speech tagging and dependency parsing. This first
phase is necessary since the next stage needs the text divided into tokens, lemmas and parts of
speech in order to continue the execution.
      </p>
      <p>Example 1: Given the following accident description: “Erano in corso attività di produzione
di acciaio. Mentre un agganciamento del nastro trasportatore vibrovaglio alla motopala, a causa
di una manovra pericolosa rimaneva con le braccia in contrasto tra le due macchine decedendo
per contusione al fegato”, the pipeline produces the tagged text visible in Table 1.</p>
      <p>
        The second phase receives as input the tagged text with lemmas and POS, and returns as
output a new series of keyphrases ordered by importance and frequency. This step uses a tool
called Keyphrase Digger (KD) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] which analyzes the text file with tokens, lemmas and pos
and returns as a result a new text file with a series of keyphrases ordered by importance and
frequency. Keyphrases are n-grams of diferent length, both single and multi-token expressions,
which capture the main concepts of a given document [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Keyprhase extraction is essential to
understand the topic covered in long text and has many applications, especially when integrated
into pipelines, like this one, that perform more complex tasks.
      </p>
      <p>Example 2: After the second phase the pipeline extracted the following concepts: “produzione
di acciaio”, “nastro trasportatore vibro”, “contusione al fegato”, “vaglio alla motopala” and
“manovra pericolosa”.</p>
      <p>As can be seen from Example 2, not all the concepts extracted from the accident description
contain information regarding the part of the body and the nature of the incident. The third and
last phase, exploiting the Rest API provided by the UMLS, allows the system to query various
databases and to discard all the concepts that do not relate to the medical field. This phase can
be divided into two distinct sub-phases:
• Concept lookup: we create a query that queries UMLS to get the medical concept. We use
as input every possible combination of the keyphrases obtained in the previous step.
• Semantic type lookup: after obtaining the medical concept, we check if it belongs to one
of these four semantic types, which represent nature and location of an injury.</p>
      <p>Example 3: After the third phase the only keyphrase not discarded is “contusione al fegato”.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>
        We tested our system on the accident descriptions contained in the InforMo dataset [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] made
available by INAIL, a repository containing the results of a survey on mostly fatal accidents
occurring during work time. This dataset contains 636 entries, each with detailed information
on the incident. Concepts are extracted from the description written in natural language in the
questionnaire (called dynamic) by those who compiled it. In the original questionnaire there
is not always consistency between what is written in the dynamic section and the nature and
location of the lesion’s attributes. As an example, we may find the concept “skull injury" in the
dynamic of the accident, and “contusion" manually written as the nature of the injury. These two
concepts might be considered as synonyms for someone who is filling out the questionnaire, but
in an ontology they are two diferent concepts. To solve this problem, we decided to manually
create a Golden Truth, analyzing all the dynamics to understand what could be extracted from
them.
      </p>
      <p>For each text, our system must extract two concepts that will form a pair constituting the
nature and location of an injury. What is extracted is compared with the golden truth to
assess the accuracy of the framework. What we want to achieve is an exact extraction of the
nature-location pair directly from the textual description of the accident.</p>
      <p>After analyzing each textual description, we have evidence to say that most of the times when
a nature-location pair is present in the text, it is in the same period, so we analyze each period.
If a couple is present in a period of the text we keep it, otherwise we delete the couple. This
allowed to greatly reduce the extracted concepts while keeping the performance unchanged.</p>
      <p>To evaluate the performance of our pipeline we used recall, precision and F1-score, whose
definitions are reported reported here:</p>
      <p>Recall :=</p>
      <p>+</p>
      <p>Precision :=</p>
      <p>+</p>
      <p>Precision · Recall
F1-Score := 2 · Precision + Recall</p>
      <p>TP represents a correctly guessed nature-location pair, the FP instead are all those incorrect
but still extracted, FN are those that are mistakenly not recognized as a match.</p>
      <p>In the UMLS query we can specify a parameter called “searchType" which can take two
diferent values: “words" (by default) or “exact". With the first, a similarity search is carried out,
resulting in the list of concepts most similar to the one given in input, ordered by decreasing
similarity. With the second, on the other hand, a result is obtained only if the input word
really exists in the database. We tested the system checking both of these parameters so as to
understand which is the best one. In the Table 2 we list the results in the two cases. The results
show that the “exact" case is better than the second. A deeper analysis highlights that this it is
due to the fact that in the second case many more concepts are extracted, most of which are
quite diferent from the original one.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>In this paper we presented a methodology for extracting medical concepts from accident
descriptions written in natural language, specifically tailored for the Italian language. The system,
still being in a preliminary phase, sufers from some limitations: (i) there is a strong dependence
on UMLS and its provided APIs, this often makes the system pretty slow in its computation
(ii) the experimental campaign carried out is pretty limited, this may cause problems in the
applicability of the framework to other input sources.
This work has been funded by INAIL within the BRiC 2018, ID09 framework, project RECKON.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Shickel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Tighe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bihorac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rashidi</surname>
          </string-name>
          ,
          <article-title>Deep ehr: a survey of recent advances in deep learning techniques for electronic health record (ehr) analysis</article-title>
          ,
          <source>IEEE journal of biomedical and health informatics 22</source>
          (
          <year>2017</year>
          )
          <fpage>1589</fpage>
          -
          <lpage>1604</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hoerbst</surname>
          </string-name>
          , E. Ammenwerth, Electronic health records,
          <source>Methods Inf Med</source>
          <volume>49</volume>
          (
          <year>2010</year>
          )
          <fpage>320</fpage>
          -
          <lpage>336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rastegar-Mojarad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Moon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Afzal</surname>
          </string-name>
          , S. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehrabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sohn</surname>
          </string-name>
          , et al.,
          <article-title>Clinical information extraction applications: a literature review</article-title>
          ,
          <source>Journal of biomedical informatics 77</source>
          (
          <year>2018</year>
          )
          <fpage>34</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          ,
          <article-title>The Unified Medical Language System (UMLS): integrating biomedical terminology</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <year>2004</year>
          )
          <fpage>D267</fpage>
          -
          <lpage>D270</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkh061.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] National Library of Medicine, UMLS reference manual</article-title>
          , https://www.ncbi.nlm.nih.gov/ books/NBK9676/,
          <year>2021</year>
          . Online; accessed 24-April-
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Masanz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Ogren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Kipper-Schuler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. G.</given-names>
            <surname>Chute</surname>
          </string-name>
          ,
          <article-title>Mayo clinical text analysis and knowledge extraction system (ctakes): architecture, component evaluation and applications</article-title>
          ,
          <source>Journal of the American Medical Informatics Association</source>
          <volume>17</volume>
          (
          <year>2010</year>
          )
          <fpage>507</fpage>
          -
          <lpage>513</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Software</surname>
          </string-name>
          <string-name>
            <surname>Foundation</surname>
          </string-name>
          , openNLP website, https://opennlp.apache.org/,
          <year>2021</year>
          . Online; accessed 28-April-
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Aronson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.-M.</given-names>
            <surname>Lang</surname>
          </string-name>
          ,
          <article-title>An overview of metamap: historical perspective and recent advances</article-title>
          ,
          <source>Journal of the American Medical Informatics Association</source>
          <volume>17</volume>
          (
          <year>2010</year>
          )
          <fpage>229</fpage>
          -
          <lpage>236</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Reátegui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ratté</surname>
          </string-name>
          ,
          <article-title>Comparison of metamap and ctakes for entity extraction in clinical notes, BMC medical informatics and decision making 18 (</article-title>
          <year>2018</year>
          )
          <fpage>13</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Palmero Aprosio</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Moretti, Tint 2.0: an all-inclusive suite for nlp in italian</article-title>
          ,
          <source>Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 10</source>
          (
          <year>2018</year>
          )
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Moretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tonelli</surname>
          </string-name>
          ,
          <article-title>Digging in the dirt: Extracting keyphrases from texts with kd</article-title>
          ,
          <source>CLiC it 198</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Turney</surname>
          </string-name>
          ,
          <article-title>Learning algorithms for keyphrase extraction</article-title>
          ,
          <source>Information retrieval 2</source>
          (
          <year>2000</year>
          )
          <fpage>303</fpage>
          -
          <lpage>336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>INAIL</surname>
          </string-name>
          ,
          <string-name>
            <surname>Informo</surname>
            <given-names>dataset</given-names>
          </string-name>
          , https://www.inail.it/sol-informo/,
          <year>2021</year>
          . Online; accessed 28-
          <fpage>April2021</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>