<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrico Mensa?</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gian Manuel Marinoy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Colla?</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Delsanto?</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele P. Radicioni?</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita` di Torino</institution>
          ,
          <addr-line>Dipartimento di Informatica</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. In this paper we propose a method for collecting a dictionary to deal with noisy medical text documents. The quality of such Italian Emergency Room Reports is so poor that in most cases these can be hardly automatically elaborated; this also holds for other languages (e.g., English), with the notable difference that no Italian dictionary has been proposed to deal with this jargon. In this work we introduce and evaluate a resource designed to fill this gap.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Italiano. In questo lavoro illustriamo un
metodo per la costruzione di un dizionario
dedicato all’elaborazione di documenti
medici, la porzione delle cartelle cliniche
annotata nei reparti di pronto soccorso.
Questo tipo di documenti e` cos`ı
rumoroso che in genere le cartelle cliniche
difficilmente posono essere direttamente
elaborate in maniera automatica. Pur
essendo il problema di ripulire questo tipo
di documenti un problema rilevante e
diffuso, non esisteva un dizionario completo
per trattare questo linguaggio settoriale.
In questo lavoro proponiamo e valutiamo
una risorsa finalizzata a condurre questo
tipo di elaborazione sulle cartelle cliniche.
Noise in textual data is a very common
phenomenon afflicting text documents, especially
1Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
when dealing with informal texts such as chats,
SMS and e-mails. This kind of text inherently
contains spelling errors, special characters,
nonstandard word forms, grammar mistakes, and so
on
        <xref ref-type="bibr" rid="ref17 ref27">(Liu et al., 2012)</xref>
        . In this work we focus on a
type of text which can also be very noisy:
emergency room reports. In the broader frame of a
project aimed at detecting injuries stemming from
violence acts in narrative texts contained in
emergency room reports, we recently developed the
VIDES, so dubbed after ‘Violence Detection
System’
        <xref ref-type="bibr" rid="ref21 ref6 ref7">(Mensa et al., 2020)</xref>
        . This system is
concerned with categorizing textual descriptions as
containing violence-related injuries (V) vs.
nonviolence-related injuries (NV), which is a
relevant task to the ends of devising alerting
mechanisms to track and prevent violence episodes.
VIDES combines a neural architecture which
performs the categorization step (thus
discriminating V and NV records) and a Framenet-based
approach, whereby semantic roles are represented
through a synthetic description employing a set of
word embeddings.2 More specifically, a model of
violent event has been devised: records that are
recognized as containing violence-related injuries
are further processed by an explanation module,
which is charged to individuate the main elements
corroborating that categorization (V) by
identifying the involved agent, the type of injury, the
involved body district etc.. Explaining the
categorization ultimately involves filling the semantic
components of the violence frame. All such
ele2Related approaches have been designed as Semantic
Role Labeling tasks
        <xref ref-type="bibr" rid="ref10 ref11 ref28 ref34">(Gildea and Jurafsky, 2002; Zapirain et
al., 2013)</xref>
        , but also frame-based approaches have been
proposed, paired to deep syntactic analysis, to extract salient
information through a template-filling approach
        <xref ref-type="bibr" rid="ref10 ref16 ref28 ref34">(Lesmo et al.,
2009; Gianfelice et al., 2013)</xref>
        .
ments contribute to recognizing a violent event as
the source of the injuries complained by ER
patients.
      </p>
      <p>During the development of VIDES we realized
that in order to run sophisticated algorithms for
the detection and extraction of such violent traits
we needed to cope with the noise contained in the
input medical records. Some efforts have been
invested to deal with different sorts of linguistic
phenomena menacing the comprehension of texts;
however, most existing works are focused on the
English language, and rely on dictionaries that
cannot be directly employed on Italian text
documents.</p>
      <p>
        In this preliminary work we start to tackle the
issue of noisy words in medical records for
Italian texts, by specifically focusing on misspellings.
Our contribution is twofold: we first manually
explore the dataset by analyzing a small sample of
records in order to determine whether the main
traits and issues present in other languages are also
shared by Italian reports; secondly, we collect,
merge and evaluate a set of Italian dictionaries,
which constitute a brick fundamental to build any
domain specific spell-checking algorithm
        <xref ref-type="bibr" rid="ref18 ref3 ref32">(Lo´pezHerna´ndez et al., 2019)</xref>
        .
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Literature shows a limited but significant interest
on the issue of detecting and correcting noisy
medical text documents; nonetheless, some
commonalities underlying this sort of text can be drawn.</p>
      <p>
        Medical texts are often very noisy; among the
most common mistakes we mention mistyping,
lack or improper use of punctuation,
grammatical errors and domain-specific abbreviations and
Latin medical terminology
        <xref ref-type="bibr" rid="ref10 ref28 ref34">(Siklo´si et al., 2013)</xref>
        .
This is mainly due to the nature of the records
themselves, and to the fact that the medical
personnel compiling the entries is often under
pressure and in a hurry.
      </p>
      <p>
        Most of the spelling correction approaches have
been carried out for English, with the exception
of research in Swedish
        <xref ref-type="bibr" rid="ref9">(Dziadek et al., 2017)</xref>
        and
Hungarian
        <xref ref-type="bibr" rid="ref10 ref28 ref34">(Siklo´si et al., 2013)</xref>
        , while no work
has been found dealing with the Italian language.
Regarding the methodologies, most works focus
on non-word errors, while disregarding
grammatical and real word mistakes. Non-word mistakes
occur when a misspelling error produces a word
that does not exist, such as ‘patienz’ instead of
’patient’, while real word mistakes occur when
a word is mistakenly replaced with another –
existing– one, like the substitution of ‘abuse’ with
‘amuse’. The adopted algorithms are diverse, with
the prevalence of approaches relying on
embeddings
        <xref ref-type="bibr" rid="ref13 ref15 ref18 ref19 ref3 ref32">(Kilicoglu et al., 2015; Workman et al.,
2019)</xref>
        or regular expressions and rule-based
systems
        <xref ref-type="bibr" rid="ref13 ref15 ref17 ref19 ref24 ref27">(Patrick et al., 2010; Sayle et al., 2012;
Lai et al., 2015)</xref>
        . However, basically all
contributions adopt a preliminary dictionary look-up
step
        <xref ref-type="bibr" rid="ref18 ref3 ref32">(Lo´pez-Herna´ndez et al., 2019)</xref>
        . To this
purpose, besides the general dictionaries provided in
toolkits such as Aspell and Google Spell Checker,3
authors often rely on (medical) domain-specific
dictionaries, such as The Unified Medical
Language System (UMLS)
        <xref ref-type="bibr" rid="ref2">(Aoki et al., 2004)</xref>
        , the
Systematized Nomenclature of Medicine-Clinical
Terms
        <xref ref-type="bibr" rid="ref29">(SNOMED-CT, 2020)</xref>
        and The
SPECIALIST Lexicon
        <xref ref-type="bibr" rid="ref4">(Browne et al., 2000)</xref>
        . It is thus
evident that the development of analogous resources
for the Italian language is a crucial step for the
design of tools and systems aimed at dealing with the
spell-checking of Italian medical text documents.
      </p>
      <p>
        Besides the treatment of misspellings, there are
also works specifically focused on abbreviations.
For instance, in
        <xref ref-type="bibr" rid="ref33">(Wu et al., 2011)</xref>
        the authors
present a corpus-based method to create a lexical
resource of English clinical abbreviations via
several machine learning algorithms. The resource
has been used to automatically detect and
expand abbreviations, and obtained interesting
experimental results. More recently, another
approach proposed in
        <xref ref-type="bibr" rid="ref14">(Kreuzthaler et al., 2016)</xref>
        focuses on abbreviations ending with a period
character; the proposed technique puts together
statistical and dictionary-based strategies to detect
abbreviations in German clinical narratives.
      </p>
      <p>In the present work we are not proposing a
specific technique for dealing with abbreviations, we
are rather concerned with misspellings. However,
the approaches already proposed for other
languages will be considered in future work to also
treat Italian abbreviations in our dataset.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data Analysis</title>
      <p>
        We analyze real data coming from a set of
emergency room reports collected in Italian hospitals
by the Italian National Institute of Health in the
3https://git.savannah.gnu.org/git/
aspell.git and https://languagetool.org,
respectively.
frame of the SINIACA project
        <xref ref-type="bibr" rid="ref25">(Pitidis et al.,
2014)</xref>
        . The SINIACA project, so dubbed after
‘Sistema Informativo Nazionale sugli Incidenti in
Ambiente di Civile Abitazione’ (National
Information System on Accidents in Civil Housing
Environment), is the Italian branch of the European
Injury Database (EU-IDB)
        <xref ref-type="bibr" rid="ref13 ref15 ref19">(Lyons et al., 2015)</xref>
        , an
EU-wide surveillance system concerned with
accidents, collecting data from hospital emergency
departments according to the EU recommendation
no. C 164/2007/01, aimed at injury prevention and
safety promotion.
      </p>
      <p>Dataset. The whole dataset amounts to 136; 144
non-empty entries, 592 of which were randomly
selected for the manual analysis. Table 1
reports some figures describing the dataset.
Double spaces and punctuation redundancy have been
fixed through regular expressions, while tokens
have been extracted by splitting the sentences
based on spaces. Also, tokens containing numbers
are presently discarded.</p>
      <p>Analysis result. We performed a manual
analysis on the subset of the original dataset: the 592
randomly selected entries herein were manually
examined, and for each entry we looked for noisy
words. Three main types of words were
annotated: i) misspellings: a wrongly typed word, e.g.,
fratura instead of frattura – fracture; ii)
abbreviations: a shortened form of a word or phrase,
e.g., dx instead of destra – right; iii) acronyms:
a word formed from the initial letters of other
words, e.g., ps instead of pronto soccorso –
emergency room. Interestingly enough, both
abbreviations and acronyms can be at least partly
considered as domain dependent: for example, in
different settings, ps may denote post scriptum
(something added at a later time, likely a letter, after
the signature), but also ‘Polizia di Stato’ (Police)
or ‘previdenza sociale’ (social security).
Dealing with such phenomena thus involves
accessing a context dependent knowledge base that
allows selecting the utterance appropriate for the
context at hand. We are presently concerned
with misspellings, acronyms and abbreviations as
noise, but only the first category can be actually
considered as an error. More specifically, while
misspellings are actual errors, abbreviations and
acronyms belong to a domain-specific language,
and these are way too specific to be recognized as
legitimate words through a general-purpose
dictionary. As seen in literature, misspellings and
abbreviations/acronyms must be treated with
different techniques, and in this work we mainly focus
on tackling the first category, while also obtaining
interesting insights regarding the second one.</p>
      <p>Table 2 illustrates the results of the annotation
process. We discovered that the dataset contains
a lot of noise, amounting to almost the 10% of
the tokens, on average 2 noisy tokens per record.
By looking separately at the different typologies
of noise we observe that misspellings are more
scattered and diverse, while the usage of
abbreviations and acronyms seems to be more coherent:
we have 670 instances of abbreviations but only
76 unique abbreviations, while 304 out of the 433
instances of misspellings are unique. This
phenomenon is also depicted in Figure 1, where we
provide the log-log plot of the frequency of each
misspelling, abbreviation and acronym ordered by
rank. We observe that the distribution of
abbreviations and acronyms has a different magnitude,
but is very similar in shape; on the other side, the
misspellings are clearly more scattered with a very
long tail of items appearing only once.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Dictionaries Creation and Evaluation</title>
      <p>
        The manual analysis uncovered characteristics and
features that are in line with those found in
literature for English datasets
        <xref ref-type="bibr" rid="ref18 ref3 ref32">(Lo´pez-Herna´ndez et
al., 2019)</xref>
        . However, to allow the development
of spell-checkers for Italian medical texts, another
key component is still missing: most approaches
aimed at error detection rely on dictionaries to
determine if a token is a legitimate word or not. In
fact, the simplest implementation of misspellings
detection is as follows: if we have at our disposal
the set W containing all of the terms of a given
language, joined to all terms pertaining the
specific domain at hand, any word w 2= W can be
likely considered as a misspell. To the best of our
knowledge, no such dictionary exists that is able
to cope with Italian medical text documents, so we
built a resource to answer to this need.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Source Dictionaries</title>
        <p>The automatic development of a dictionary is not
a trivial task. We want to reach the highest
possible coverage for both general terms and
specific medical terminology, but at the same time we
cannot rely too much on unverified sources (e.g.,
crowd-sourced data) with the risk of introducing
misspellings and errors into the dictionary. We
selected different sources and arranged them into
four main classes:</p>
        <p>
          MED: a collection of medical terms built by
putting together five medical online
dictionaries
          <xref ref-type="bibr" rid="ref1 ref1 ref1 ref1 ref1 ref22 ref22 ref23 ref23 ref30 ref30 ref30 ref30 ref30 ref5 ref5 ref5 ref5 ref5">(torrinomedica.it, 2020; abcsalute.it,
2020; codifa.it, 2020; my-personaltrainer.it,
2020a; my-personaltrainer.it, 2020b)</xref>
          ,
containing medical specific terms and
medications names;
        </p>
        <sec id="sec-4-1-1">
          <title>ITA: a collection of Italian terms built by Table 3: Figures of the 5000 annotated tokens used to evaluate the dictionaries.</title>
          <p>
            merging three well-known Italian online
dictionaries
            <xref ref-type="bibr" rid="ref12 ref26 ref8">(Hoepli, 2020; Sabatini-Colletti,
2020; De Mauro, 2020)</xref>
            ;
          </p>
        </sec>
        <sec id="sec-4-1-2">
          <title>WMED: a collection of terms from Wikipedia</title>
          <p>
            pages pertaining the medical domain. The
list of Wikipedia medical pages has been
obtained by querying the SPARQL endpoint of
Wikidata
            <xref ref-type="bibr" rid="ref31">(Vrandecˇic´ and Kro¨tzsch, 2014)</xref>
            ,
while the pages have been taken from the 20
August 2020 Wikipedia dump;
WMOV: since medical records also contain a
brief narrative text of the events that led to
the (either violent or accidental) injuries, we
added terms associated to eventive and
narrative genres by collecting Wikipedia pages
pertaining to movies, television series and
literary work that are expected to contain
narrative terminology.
          </p>
          <p>The set of terms extracted from Wikipedia can
potentially contain misspellings and errors, and so
we also set a frequency minimum which allows for
the pruning of the tokens herein. We annotate this
parameter with a subscript next to the set name,
e.g., WMOV1 indicates that the threshold was set to
1 for the terms frequency.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation</title>
        <p>Building the dataset. In order to assess the
quality of the collected dictionaries we started
from the 49; 116 unique tokens in the dataset,
removed the stop words4 and randomly selected
5; 000 of them to be manually annotated. The
annotation was carried out by four of the authors
of this paper. The selection algorithm was
designed so to increase the probability of a token to
be selected in accordance to its frequency in the
dataset. These 5; 000 tokens were then annotated
4We used the set of stop words made available by Spacy
(https://spacy.io/) for the Italian language.
with one of the following four classes: correct
words (regardless of their domain specificity),
abbreviations, acronyms and misspellings. The first
three classes represent terms that should be found
in our resource, while the last category contains
words that should not be present in the dictionary.
Table 3 reports the statistics featuring the dataset
annotated for evaluation purposes.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluating the dictionary. In Table 4 we report</title>
        <p>the results of the dictionaries evaluation. Each
dictionary has been built by taking into consideration
one or more of the previously presented sources.
Multiple sources have been simply merged into a
unique set of terms, without repetitions. We
assess the quality of each dictionary via two
measures, coverage and correctness. The coverage is
the percentage of words that were found in the
dictionary (either correct words, abbreviations or
acronyms), while the correctness is the percentage
of misspellings that were not present in the
dictionary. We considered different combinations of the
sources, the tuning of the frequency-based filtering
parameter, and an additional lemmatization step.</p>
        <p>We observe that both the ITA and the MED sets
are fundamentally correct, even though they also
include words that in the common usage are
frequently misspelled, such as passeggiero in place
of the correct form passeggero. On the other side,
its :62 coverage is unsatisfactory (please refer to
the second row of Table 4, MED, ITA); it also
witnesses that medical jargon is only partially grasped
by dictionaries in the MED set. As expected, the
introduction of terms from Wikipedia improves the
coverage, but with detrimental effect on the
correctness. This also holds for the WMOV set, which
is rich but also pretty noisy. By fine tuning the
frequency thresholds of both WMED and WMOV we
were able to prune most of the noise and to
preserve the coverage at the same time, finally
obtaining a good dictionary with the combination MED,
ITA, WMED1, WMOV5.</p>
        <p>This setting was also tested by applying a
lemmatization step on both Wikipedia terms and
our dataset tokens. Interestingly, the
lemmatization introduces more mistakes than it solves: this
is due the the fact that unpredictably the
lemmatizer converts misspellings into legitimate words
that do not necessarily correspond to their
correct spelling. This fact shows also that
lemmatization, which is acknowledged as a task almost
completely solved from a scientific point of view, still
poses relevant issues for the medical jargon and
for domain-specific languages more in general.</p>
        <p>A lot of abbreviations are not yet covered in
the dictionary. Once again, these abbreviations
are dataset-specific (and perhaps also follow local
uses rather than widely accepted practices), and
thus these are very hard to find even on
specialized public medical resources. For instance, incid
(incidente – accident) appears very frequently and
its easily understandable by humans but its not a
common or medical abbreviation. The same
phenomenon can also be observed on acronyms, that
are less sparse and more adherent to widely
accepted practices and standards.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>
        In this work we tackled the issue of detecting
textual noise in Italian room emergency reports,
focusing specifically on misspellings. Firstly we
examined the reports and found out that the sorts
of issues reported in literature for other languages
can also be found in Italian text documents.
Secondly, we developed and evaluated an Italian
dictionary suited for the task of noise detection. In
future work we plan to expand the dictionary by
including the terms from the Italian ICD-9 and
ICD-10 (International Classification of Diseases),
that may be useful to interpret acronyms and
resolve abbreviations. Moreover, we plan to employ
this dictionary in a fully fledged spell-checking
system. Finally, the usage of semantic —sense
indexed— representations such as, e.g.,
        <xref ref-type="bibr" rid="ref20">(Mensa et
al., 2018)</xref>
        and
        <xref ref-type="bibr" rid="ref21 ref21 ref6 ref6 ref7 ref7">(Colla et al., 2020a; Colla et al.,
2020b)</xref>
        will be explored, in order to deal with real
word mistakes, and more in general contextual
information
        <xref ref-type="bibr" rid="ref18 ref3 ref32">(Basile et al., 2019)</xref>
        will be considered
as a main cue in order to uncover and correct this
sort of errors. For example, by leveraging the
terminology surrounding noisy tokens we plan to
distinguish the more scattered misspellings from the
other terms that are not present in our dictionary.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The first author was supported by a grant
provided by Universita` degli Studi di Torino. This
research is also supported by Fondazione CRT, RF
2019:2263.</p>
      <p>2020.</p>
      <p>Colletti.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[abcsalute.it2020] abcsalute.it</source>
          .
          <year>2020</year>
          . Abcsalute.it - Dizionario Medico. http://www.abcsalute. it/dizionario-medico.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Aoki et al.2004
          <string-name>
            <surname>] Kiyoko F Aoki</surname>
          </string-name>
          ,
          <string-name>
            <surname>Atsuko Yamaguchi</surname>
            , Nobuhisa Ueda, Tatsuya Akutsu, Hiroshi Mamitsuka, Susumu Goto, and
            <given-names>Minoru</given-names>
          </string-name>
          <string-name>
            <surname>Kanehisa</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Kcam (kegg carbohydrate matcher): a software tool for analyzing the structures of carbohydrate sugar chains</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>32</volume>
          (
          <issue>suppl 2</issue>
          ):
          <fpage>W267</fpage>
          -
          <lpage>W272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Basile et al.2019]
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Tommaso Caselli, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Meaning in Context: Ontologically and linguistically motivated representations of objects and events</article-title>
          .
          <source>Applied Ontology</source>
          ,
          <volume>14</volume>
          :
          <fpage>335</fpage>
          -
          <lpage>341</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Browne et al.2000
          <string-name>
            <surname>] Allen</surname>
            <given-names>C Browne</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexa T McCray</surname>
            ,
            <given-names>and Suresh</given-names>
          </string-name>
          <string-name>
            <surname>Srinivasan</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>The specialist lexicon</article-title>
          .
          <source>National Library of Medicine Technical Reports</source>
          , pages
          <fpage>18</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[codifa.it2020] codifa.it</source>
          .
          <year>2020</year>
          . codifa.it - Dizionario dei Farmaci. https://www.codifa.it/ farmaci.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Colla et al.2020a]
          <string-name>
            <surname>Davide</surname>
            <given-names>Colla</given-names>
          </string-name>
          , Enrico Mensa, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          . 2020a.
          <article-title>Lesslex: Linking multilingual embeddings to sense representations of lexical items</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>46</volume>
          (
          <issue>2</issue>
          ):
          <fpage>289</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Colla et al.2020b]
          <string-name>
            <surname>Davide</surname>
            <given-names>Colla</given-names>
          </string-name>
          , Enrico Mensa, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          . 2020b.
          <article-title>Novel metrics for computing semantic similarity with sense embeddings</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>206</volume>
          :
          <fpage>106346</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>[De Mauro2020] De Mauro</surname>
          </string-name>
          .
          <year>2020</year>
          . Dizionario Italiano Nuovo De Mauro. https://dizionario. internazionale.it/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Dziadek et al.2017]
          <string-name>
            <given-names>Juliusz</given-names>
            <surname>Dziadek</surname>
          </string-name>
          , Aron Henriksson, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Duneld</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Improving terminology mapping in clinical text with context-sensitive spelling correction</article-title>
          .
          <source>Informatics for Health: Connected Citizen-Led Wellness and Population Health</source>
          ,
          <volume>235</volume>
          :
          <fpage>241</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Gianfelice et al.2013]
          <string-name>
            <given-names>Davide</given-names>
            <surname>Gianfelice</surname>
          </string-name>
          , Leonardo Lesmo, Monica Palmirani, Daniele Perlo, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Modificatory Provisions Detection: a Hybrid NLP Approach</article-title>
          . In Bart Verheij, editor,
          <source>Proceedings of ICAIL 2013: XIV International Conference on Artificial Intelligence and Law</source>
          , pages
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Gildea and Jurafsky2002] Daniel Gildea and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Automatic labeling of semantic roles</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>28</volume>
          (
          <issue>3</issue>
          ):
          <fpage>245</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[Hoepli2020] Hoepli</source>
          .
          <year>2020</year>
          .
          <article-title>Dizionario Italiano Hoepli</article-title>
          . https://dizionari.repubblica. it/italiano.html.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Kilicoglu et al.2015]
          <string-name>
            <given-names>Halil</given-names>
            <surname>Kilicoglu</surname>
          </string-name>
          , Marcelo Fiszman, Kirk Roberts, and
          <string-name>
            <surname>Dina</surname>
          </string-name>
          Demner-Fushman.
          <year>2015</year>
          .
          <article-title>An ensemble method for spelling correction in consumer health questions</article-title>
          .
          <source>In AMIA Annual Symposium Proceedings</source>
          , volume
          <year>2015</year>
          , page 727. American Medical Informatics Association.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Kreuzthaler et al.2016]
          <string-name>
            <given-names>Markus</given-names>
            <surname>Kreuzthaler</surname>
          </string-name>
          , Michel Oleynik, Alexander Avian, and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Schulz</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Unsupervised abbreviation detection in clinical narratives</article-title>
          .
          <source>In Proceedings of the clinical natural language processing workshop (ClinicalNLP)</source>
          , pages
          <fpage>91</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Lai et al.2015]
          <string-name>
            <surname>Kenneth H Lai</surname>
          </string-name>
          , Maxim Topaz,
          <string-name>
            <surname>Foster R Goss</surname>
            ,
            <given-names>and Li</given-names>
          </string-name>
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Automated misspelling detection and correction in clinical free-text records</article-title>
          .
          <source>Journal of biomedical informatics</source>
          ,
          <volume>55</volume>
          :
          <fpage>188</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Lesmo et al.2009]
          <string-name>
            <given-names>Leonardo</given-names>
            <surname>Lesmo</surname>
          </string-name>
          , Alessandro Mazzei, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Extracting Semantic Annotations from Legal Texts</article-title>
          .
          <source>In Proceedings of the International Conference on Hypertext, HT09</source>
          , pages
          <fpage>167</fpage>
          -
          <lpage>172</lpage>
          , Turin, Italy, July. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Liu et al.2012] Fei Liu, Fuliang Weng, and
          <string-name>
            <given-names>Xiao</given-names>
            <surname>Jiang</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A broad-coverage normalization system for social media language</article-title>
          .
          <source>In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>1035</fpage>
          -
          <lpage>1044</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Lo´
          <fpage>pez</fpage>
          -Herna´ndez et al.2019]
          <article-title>Je´sica Lo´ pezHerna´ndez, A´ ngela Almela, and Rafael ValenciaGarc´ıa</article-title>
          .
          <year>2019</year>
          .
          <article-title>Automatic spelling detection and correction in the medical domain: A systematic literature review</article-title>
          .
          <source>In International Conference on Technologies and Innovation</source>
          , pages
          <fpage>95</fpage>
          -
          <lpage>108</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Lyons et al.2015]
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Lyons</surname>
          </string-name>
          , Rupert Kisse, and
          <string-name>
            <given-names>Wim</given-names>
            <surname>Rogmans</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Eu-injury database introduction to the functioning of the injury database (idb)</article-title>
          . https://bit.ly/37FAKaB.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Mensa et al.2018]
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Mensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          , and Antonio Lieto.
          <year>2018</year>
          .
          <article-title>Cover: a linguistic resource combining common sense and lexicographic information</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>52</volume>
          (
          <issue>4</issue>
          ):
          <fpage>921</fpage>
          -
          <lpage>948</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Mensa et al.2020]
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Mensa</surname>
          </string-name>
          , Davide Colla, Marco Dalmasso, Marco Giustini, Carlo Mamo, Alessio Pitidis, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Violence detection explanation via semantic roles embeddings</article-title>
          .
          <source>BMC Medical Informatics and Decision Making</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>263</fpage>
          -
          <lpage>275</lpage>
          , Oct.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>[my-personaltrainer.it2020a] my-personaltrainer.it</article-title>
          . 2020a.
          <article-title>Lista delle Malattie di My Personal Trainer</article-title>
          . https://www.my-personaltrainer.it/ malattie_a_z.php.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>[my-personaltrainer.it2020b] my-personaltrainer.it</article-title>
          . 2020b.
          <article-title>Lista di Sintomi di My Personal Trainer</article-title>
          . https://www.my-personaltrainer.it/ sintomi_a_z.php.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Patrick et al.2010]
          <string-name>
            <given-names>Jon</given-names>
            <surname>Patrick</surname>
          </string-name>
          , Mojtaba Sabbagh, Suvir Jain, and
          <string-name>
            <given-names>Haifeng</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Spelling correction in clinical notes with emphasis on first suggestion accuracy</article-title>
          .
          <source>In Proceedings of 2nd Workshop on Building and Evaluating Resources for Biomedical Text Mining (BioTxtM2010)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Pitidis et al.2014]
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Pitidis</surname>
          </string-name>
          , Gianni Fondi, Marco Giustini, Elo¨ıse Longo, Giuseppe Balducci,
          <article-title>Gruppo di lavoro SINIACA-IDB, and Dipartimento di Ambiente e Connessa Prevenzione Primaria</article-title>
          , ISS.
          <year>2014</year>
          .
          <article-title>Il Sistema SINIACA-IDB per la sorveglianza degli incidenti</article-title>
          .
          <source>Notiziario dell'Istituto Superiore di Sanit a`</source>
          ,
          <volume>27</volume>
          (
          <issue>2</issue>
          ):
          <fpage>11</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>[</surname>
          </string-name>
          Sabatini-Colletti2020]
          <article-title>Sabatini-Colletti</article-title>
          . Dizionario Italiano Sabatibi https://dizionari.corriere.it/ dizionario_italiano/.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Sayle et al.2012]
          <string-name>
            <given-names>Roger</given-names>
            <surname>Sayle</surname>
          </string-name>
          , Paul Hongxing Xie, and
          <string-name>
            <given-names>Sorel</given-names>
            <surname>Muresan</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Improved chemical text mining of patents with infinite dictionaries and automatic spelling correction</article-title>
          .
          <source>Journal of chemical information and modeling</source>
          ,
          <volume>52</volume>
          (
          <issue>1</issue>
          ):
          <fpage>51</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>[Sikl o´si et al</source>
          .2013]
          <article-title>Borba´la Siklo´ si, Attila Nova´k, and Ga´bor Pro´ sze´ky</article-title>
          .
          <year>2013</year>
          .
          <article-title>Context-aware correction of spelling errors in hungarian medical documents</article-title>
          .
          <source>In International Conference on Statistical Language and Speech Processing</source>
          , pages
          <fpage>248</fpage>
          -
          <lpage>259</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>[SNOMED-CT2020] SNOMED-CT</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>International Health Terminology Standards Development Organisation</article-title>
          . http://www.ihtsdo. org/snomed-ct/.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>[torrinomedica.it2020] torrinomedica</article-title>
          .it. torrinomedica.it - Dizionario dei https://www.torrinomedica.it/ schede-farmaci.
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Vrandecˇic´ and Kro¨ tzsch2014]
          <article-title>Denny Vrandecˇic´</article-title>
          and Markus Kro¨ tzsch.
          <year>2014</year>
          .
          <article-title>Wikidata: A free collaborative knowledgebase</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>57</volume>
          (
          <issue>10</issue>
          ):
          <fpage>78</fpage>
          -
          <lpage>85</lpage>
          , September.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Workman et al.2019]
          <string-name>
            <given-names>T</given-names>
            <surname>Elizabeth</surname>
          </string-name>
          <string-name>
            <given-names>Workman</given-names>
            , Yijun Shao, Guy Divita, and
            <surname>Qing</surname>
          </string-name>
          Zeng-Treitler.
          <year>2019</year>
          .
          <article-title>An efficient prototype method to identify and correct misspellings in clinical text</article-title>
          .
          <source>BMC research notes</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
            et al.2011]
            <given-names>Yonghui</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>S Trent</given-names>
          </string-name>
          <string-name>
            <surname>Rosenbloom</surname>
          </string-name>
          , Joshua C Denny,
          <article-title>Randolph A Miller, Subramani Mani, Dario A Giuse,</article-title>
          and
          <string-name>
            <given-names>Hua</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Detecting abbreviations in discharge summaries using machine learning methods</article-title>
          .
          <source>In AMIA Annual Symposium Proceedings</source>
          , volume
          <year>2011</year>
          , page 1541. American Medical Informatics Association.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [Zapirain et al.2013]
          <string-name>
            <given-names>Benat</given-names>
            <surname>Zapirain</surname>
          </string-name>
          , Eneko Agirre, Lluis Marquez, and
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Surdeanu</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Selectional preferences for semantic role classification</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>39</volume>
          (
          <issue>3</issue>
          ):
          <fpage>631</fpage>
          -
          <lpage>663</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>