<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>BioVec-Ita: Biomedical Word Embeddings for the Italian Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcello Bavaro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Dolci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Piantella</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano - Department of Electronics</institution>
          ,
          <addr-line>Information and Bioengineering</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>In the healthcare field, the information created by digital technologies that collect clinical care notes, health service reports and patients' records, is generating terabytes of data, a great part of which is in textual format. These datasets may become an incredibly valuable asset only if the knowledge they carry is extracted using the appropriate artificial intelligence techniques, and specifically natural language processing (NLP) ones. Unfortunately, most existing tools support NLP for the English language, while local administrations and hospitals typically work in their native language, and therefore it becomes very important to have NLP tools to process biomedical data written also in these languages. Word embeddings are a popular and powerful NLP technique to extract semantics from textual data that could be very useful to solve the problem, but unfortunately for the Italian language there are no such tools specialized in the biomedical field. In this paper we propose BioVec-Ita, a new word embedding model for Italian, specialized in the biomedical field and designed using Word2vec, a flexible model for semantic representation that can be easily integrated with other pipelines. We also evaluate the performance of our word embeddings model in capturing the semantic similarities of biomedical terms, using three very popular test datasets translated into Italian.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Word embeddings</kwd>
        <kwd>Word2vec</kwd>
        <kwd>natural language processing</kwd>
        <kwd>healthcare</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the field of biomedicine, the digitization of clinical care processes and health services is
producing more and more medical data, much of it in textual format: reports, nursing notes,
discharge letters and emergency room reports are just a few of the digital documents that are
generated every day in hospitals. Moreover, humans are being digitized through new medical
devices, apps, and monitoring technologies, which track, analyze and store a massive amount
of data. It has been estimated that by 2025 the annual growth rate of data for healthcare will
reach 36%, more than the general growth rate estimated at 27% [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Biomedical pieces of information are a huge asset for medical researchers and, since much
of them are in unstructured form, it is necessary to leverage artificial intelligence (AI) and
natural language processing (NLP) techniques to extract and create knowledge from them. This
knowledge can then be used to improve patients’ care and management, reduce costs, speed up
procedures, and more. Among the most popular NLP models for biomedicine are pre-trained
language representations such as BioBERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This model is able to perform text mining tasks
like named entity recognition, relation extraction, and question answering. Moreover, language
models can be used to classify documents based on diseases described in the text [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], or to
extract information from Electronic Health Records to predict future diseases [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Despite a widespread use of the English language, local administrations and hospitals still
work in their native language. For this reason, biomedical researchers need tools for data analysis
in other languages, such as Italian. For example, word embeddings are a powerful semantic
representations used for many diferent NLP tasks, but there exists few word embeddings for
the Italian language [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ], and none specifically trained for the biomedical field, except for
some preliminary work [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Word embeddings for the biomedical field are commonly used
as feature input to machine learning and deep learning models, enabling techniques for the
contextualization of textual data to be used for many tasks, such as readmission prediction [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In this work, we introduce a biomedical word embeddings model for the Italian language,
named BioVec-Ita. Our model is based on Word2vec [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and takes inspiration from previous
literature on English biomedical word embeddings [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. A key step of creating powerful
language models is to search for well-documented biomedical data in text format to be used
for the training phase. To train our model, we use a set of corpora ofered by OPUSnlp [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
an online repository of textual data, and additional data extracted from Wikipedia dumps of
thousands of pages in Italian belonging to biomedical topics. Moreover, we include material
from the International Classification of Diseases (ICD9) translated into Italian. 1 Finally, we test
the performance of BioVec-Ita on three popular test datasets for biomedical word embeddings
evaluation, manually translated into Italian.2
      </p>
      <p>The rest of the paper is organized as follows: Section 2 explores the state of the art, Section 3
introduces the BioVec-Ita embeddings and its training phase, Section 4 describes results of the
experimental evaluation, Section 5 concludes the the paper and outlines future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Over the years, there have been several studies on the creation of specialized word embeddings
for the biomedical field and on their usage in a variety of downstream medical tasks. For instance,
Xiao et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] combined word embeddings and recurrent neural networks to predict future
readmissions of patients after their discharge, while Liu et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] integrated word embeddings
in a drug name recognition system. Katikapalli et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] studied the use of diferent pre-trained
language models, including the famous ELMo, BERT, and sentenceBERT, in capturing the
semantic of biomedical terms. Similar results on intrinsic performance of biomedical word
embeddings were obtained by Chiu et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] using a much simpler Word2vec model trained
on a big text corpus of titles and abstracts of papers from PubMed3, totaling 2.7 billion tokens.
1The Italian version of ICD9 is available at https://www.salute.gov.it/portale/temi/manuale-icd9cm/
2BioVec-Ita embeddings and test datasets are available at https://github.com/MarcelloBavarof/BioVec-Ita
3https://pubmed.ncbi.nlm.nih.gov/
      </p>
      <p>
        Although there are some works about NLP for biomedical applications regarding the Italian
language (e.g., extraction of medical concepts [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or medical notes understanding [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) the
limited amount of freely available textual data and resources hinder the research in this area.
In fact, there is a small number of works even on general-purpose Italian word embeddings:
Berardi et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] first addressed the issue in 2015, and more recently Di Gennaro et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The only exception is represented by the recent preliminary work by Bondarenko et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
where the researchers tried to improve a fine-tuned version of BERT on Italian medical texts by
combining contrastive learning and knowledge graph embeddings. However, their work leaves
room for improvements considering the performance score obtained by equivalent English
biomedical word embeddings.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. BioVec-Ita</title>
      <p>
        Biomedical word embeddings are a popular and powerful tool for language understanding, but
there is a lack of research regarding the Italian language. For this reason, we take inspiration
from the much richer English literature on biomedical word embeddings, particularly regarding
Word2vec models for biomedicine. Word2vec [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is a flexible and powerful model to retrieve
accurate semantic representations of words, and it is easy to train and retrain even with a big
amount of data. It is part of the family of static vector models: contextualized vector models
provide instead representations based also on the context of the sentence. However, despite
contextualized models being generally considered more powerful for language understanding
tasks, biomedical terms are rarely synonymical, and for many medical tasks static embeddings
obtain comparably high results in intrinsic performance tests [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Our work starts from the experience of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], whose authors provide an in-depth overview of
parameters setup for biomedical Word2vec model training in English. In this section, we first
illustrate how we identified and pre-processed the training data used for the creation of
BioVecIta. Then, after describing the training parameters adopted, we proceed with a preliminary
evaluation on the intrinsic quality of the word embedding model produced.
      </p>
      <sec id="sec-3-1">
        <title>3.1. Training Data</title>
        <p>
          The main problem when it comes to biomedical word embeddings is that there are no resources
in Italian like those ofered in English by PubMed or MIMIC III [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], containing huge amounts
of textual and medical data. Therefore, our training data mainly comes from OPUSnlp, an
open-source collection of text corpora in diferent languages [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Specifically, we selected three
corpora on biomedical topics, namely ELRC-wikipedia_health, EMAv3 from European Medicines
Agency, and Tilde MODEL-multilingual open data for EU languages. Additionally, we include the
Italian version of ICD9.
        </p>
        <p>
          Finally, following the methodology described in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], we integrate a number of Wikipedia
pages on medical concepts in Italian. In particular, we download the Italian Wikipedia pages
belonging to the following categories: “Salute”, “Medicina”, “Procedure mediche”, “Diagnostica
medica”, “Specialità mediche”, “Farmacologia”, “Farmaci”, “Chirurgia”, and “Infermieristica”.
Moreover, we add all the respective subcategories at depth two: for instance, considering the
category “Salute”, its subcategories include “Alimentazione” at depth one, which in turn includes
the subcategory “Alimentazione animale” at depth two. Duplicated pages are properly removed.
In total, 14,395 pages are considered.
        </p>
        <p>The overall resulting training data is then split into sentences, with each sentence tokenized
to separate its words (also called tokens). In addition, all special characters (e.g., emojis and
other non-character symbols) are removed and the order of the sentences in the datasets is
randomized. The overall training dataset contains about 41M tokens.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Parameters and Setup</title>
        <p>Two important parameters to adjust are the number of training epochs and the minimum number
of occurrences (min_count) required for a word to appear in BioVec-Ita. If min_count is too
large, important words may be left out of the model, if it is too small, the size of the vocabulary
increases too much without accurate representations. Regarding the number of epochs, we
carry out several tests by varying the number of epochs between 10 and 100. Additionally, we
test the training process with both 4 and 5 as min_count values. Table 1 shows an overview of
all the parameters involved in the training of BioVec-Ita.</p>
        <p>We train BioVec-Ita on a machine equipped with an NVIDIA GPU RTX 360 (12Gb VRAM)
and an Intel I5-12400F, using the Gensim library in version 4.2.0 in an environment with Python
3.9.15. The resulting word embeddings contain 158,251 words with min_count equal to 4, and
134,697 words with min_count equal to 5.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and Results</title>
      <p>In this section, we illustrate the tests used to assess the performance of BioVec-Ita. We start
from a description of the test datasets and the evaluation metrics, followed by the presentation
and discussion of the experimental results.</p>
      <sec id="sec-4-1">
        <title>4.1. Test Datasets</title>
        <p>We evaluate the results BioVec-Ita on the translated version of the following three datasets:
MayoSRS [19], UMNSRS-similarity and UMNSRS-relatedness [20].</p>
        <p>
          MayoSRS and UMNSRS-similarity in Italian are provided in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and we make only few changes
to certain terms to improve the quality of the data. For instance, we decide not to translate
American drug names into the equivalent Italian active ingredient, but into the equivalent most
popular drug name in our country. For example, Coumadin (original drug) is not translated into
Warfarina (active ingredient), but instead into Warfarin (Italian equivalent drug). Concerning
UMNSRS-relatedness, we manually translated the medical terms from English to Italian. All
three datasets are composed of pairs of medical terms, associated with a score indicating the
similarity between the two. Medical terms may consist of a single word (e.g., “fever”), or
multi-word expressions (e.g., “portal hypertension”).
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Metrics</title>
        <p>To evaluate the quality of word embeddings in representing the semantics of biomedical words,
we calculate the cosine similarity between the word pairs in each tuple of the test datasets,
comparing it to their gold similarity score. Considering two embeddings ⃗ and ⃗, their cosine
similarity is defined as:
(, ) = () =  ·  ,
|||| · || ||
where ||⃗|| is the magnitude of ⃗. Once the similarity is calculated for all the tuples, the Spearman
function is used to assess whether the distribution of the obtained scores follows the distribution
of the gold scores indicated in the dataset. The final score is between 1 and 100. Since BioVec-Ita
represents only single words, for multi-word expressions we take the average vector of all the
word vectors that compose the expression. While averaging, semantically meaningless words
such as prepositions are excluded. Tests are carried out with the help of BioNLP-20164, a toolkit
for word embeddings evaluation.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Performance Results</title>
        <p>4https://github.com/cambridgeltl/BioNLP-2016</p>
        <p>48,42
44,3
45,96
44,83
44,02</p>
        <p>
          44,36
47,82
0. To ensure a fair comparison, we also calculate BioVec-Ita scores using the original versions of
MayoSRS and UMNSRS-similarity without any of our changes. In both cases, our model obtain
significantly higher results. Finally, Figure 3 compares the results of our best BioVec-Ita model
with the state-of-the-art English biomedical word embeddings obtained from replicating the
methodology described in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. BioVec-Ita obtains fairly comparable results, despite a diference
in the size of the training data of two orders of magnitude. Moreover, our model has a lower
miss rate on MayoSRS compared to the English embeddings.
45,96
8,91
        </p>
        <p>49,26
18,33
22,94
18,81
6,01</p>
        <p>6,98</p>
        <p>BioVec-Ita
MayoSRS score MayoSRS miss rate
UMNSRS-sim miss rate UMNSRS-rel score</p>
        <p>Chiu et al.</p>
        <p>UMNSRS-sim score
UMNSRS-rel miss rate</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The ever-growing amount of healthcare data demands advanced techniques to be properly
exploited. Word embeddings represent a powerful and flexible NLP tool for managing all sorts
of biomedical tasks involving language understanding, but there is a chronic lack of research
on biomedical word embeddings for the Italian language. In this paper, we introduce
BioVecIta, an Italian word embedding model trained specifically on biomedical data. BioVec-Ita is
based on Word2vec, thus providing a flexible and simple, yet powerful model for a variety
of NLP tasks. After describing the textual data used to train it, and the overall methodology
adopted, we tested our model on three translated datasets to assess the quality of its semantic
representations. Additionally, we compared BioVec-Ita both with previous Italian biomedical
word embeddings from literature and with state-of-the-art biomedical embeddings in English,
showing that our model achieves high-quality semantic representations comparable to those of
its English counterpart, given the smaller size of the training data.</p>
      <p>Future work includes retrieving more textual data in Italian regarding the biomedical field,
for instance by including the digitized version of medical books. Moreover, we plan to train and
test BioVec-Ita with diferent parameters. For example, increasing the dimensionality of the
vectors to 300 or even higher, to try improving the semantic granularity of the representations.
On the other hand, this would also increase the training time and the complexity of the model.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We are grateful to Letizia Tanca for helping us in the definition and realization of this work. We
also thank Carmela Giannelli for supervising the translation of the test datasets into Italian.</p>
      <p>Alma Mater Studiorum, Università di Bologna, 2018. URL: http://amslaurea.unibo.it/16714/.
[19] T. Pedersen, S. V. Pakhomov, S. Patwardhan, C. G. Chute, Measures of semantic similarity
and relatedness in the biomedical domain, JBI 40 (2007) 288–299.
[20] S. Pakhomov, B. McInnes, T. Adam, Y. Liu, T. Pedersen, G. B. Melton, Semantic
similarity and relatedness between clinical terms: an experimental study, in: AMIA Annual
Symposium Proceedings, 2010, p. 572.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>D. R.-J. G.-J. Rydning</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Reinsel</surname>
            ,
            <given-names>J. Gantz,</given-names>
          </string-name>
          <article-title>The digitization of the world from edge to core</article-title>
          ,
          <source>Framingham: International Data Corporation</source>
          <volume>16</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>So</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <article-title>Biobert: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>36</volume>
          (
          <year>2020</year>
          )
          <fpage>1234</fpage>
          -
          <lpage>1240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Luo,</surname>
          </string-name>
          <article-title>Clinical text classification with rule-based features and knowledgeguided convolutional neural networks</article-title>
          ,
          <source>BMC Medical Informatics and Decision Making</source>
          <volume>19</volume>
          (
          <year>2019</year>
          )
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , N. Razavian,
          <article-title>Deep ehr: Chronic disease prediction using medical notes</article-title>
          ,
          <source>in: Machine Learning for Healthcare Conference</source>
          , PMLR,
          <year>2018</year>
          , pp.
          <fpage>440</fpage>
          -
          <lpage>464</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Di Gennaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Buonanno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Di</given-names>
            <surname>Girolamo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ospedale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Palmieri</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Fedele,</surname>
          </string-name>
          <article-title>An analysis of word2vec for the italian language</article-title>
          ,
          <source>Progresses in Artificial Intelligence and Neural Systems</source>
          (
          <year>2021</year>
          )
          <fpage>137</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Berardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Marcheggiani</surname>
          </string-name>
          ,
          <article-title>Word embeddings go to italy: A comparison of models and training datasets</article-title>
          .,
          <source>in: Proceedings of the 6th Italian Information Retrieval Workshop</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Bondarenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferrod</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Di</given-names>
            <surname>Caro</surname>
          </string-name>
          ,
          <article-title>Combining contrastive learning and knowledge graph embeddings to develop medical word embeddings for the italian language</article-title>
          ,
          <source>arXiv preprint arXiv:2211.05035</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiao</surname>
          </string-name>
          , T. Ma,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Dieng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Readmission prediction via deep contextual embedding of clinical concepts</article-title>
          ,
          <source>PloS one 13</source>
          (
          <year>2018</year>
          )
          <article-title>e0195024</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Chiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Crichton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Korhonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <article-title>How to train good word embeddings for biomedical nlp</article-title>
          ,
          <source>in: Proceedings of the 15th Workshop on Biomedical Natural Language Processing</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>166</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          ,
          <article-title>Parallel data, tools and interfaces in opus</article-title>
          ,
          <source>in: Proceedings of the Eighth International Conference on Language Resources and Evaluation</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>2214</fpage>
          -
          <lpage>2218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Efects of semantic features on machine learning-based drug name recognition systems: Word embeddings vs</article-title>
          .
          <source>manually constructed dictionaries, Information</source>
          <volume>6</volume>
          (
          <year>2015</year>
          )
          <fpage>848</fpage>
          -
          <lpage>865</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Kalyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sangeetha</surname>
          </string-name>
          ,
          <article-title>A hybrid approach to measure semantic relatedness in biomedical concepts</article-title>
          ,
          <source>arXiv preprint arXiv:2101.10196</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Agnello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Ansaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Azzalini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Piantella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rabosio</surname>
          </string-name>
          , L. Tanca,
          <article-title>Extraction of medical concepts from italian natural language descriptions</article-title>
          ,
          <source>in: 29th Italian Symposium on Advanced Database Systems</source>
          , SEBD, volume
          <volume>2994</volume>
          ,
          <year>2021</year>
          , pp.
          <fpage>275</fpage>
          -
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferrod</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Brunetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Di</given-names>
            <surname>Caro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dragoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ghidini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Marinello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sulis</surname>
          </string-name>
          ,
          <article-title>A support for understanding medical notes: Correcting spelling errors in italian clinical records</article-title>
          .,
          <source>in: SMARTERCARE@ AI* IA</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F. K.</given-names>
            <surname>Khattak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jeblee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pou-Prom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdalla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Meaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rudzicz</surname>
          </string-name>
          ,
          <article-title>A survey of word embeddings for clinical text</article-title>
          ,
          <source>JBI</source>
          <volume>100</volume>
          (
          <year>2019</year>
          )
          <fpage>100057</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , T. J.
          <string-name>
            <surname>Pollard</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>L.-w. H.</given-names>
          </string-name>
          <string-name>
            <surname>Lehman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ghassemi</surname>
            , B. Moody, P. Szolovits,
            <given-names>L. Anthony</given-names>
          </string-name>
          <string-name>
            <surname>Celi</surname>
          </string-name>
          , R. G. Mark,
          <article-title>Mimic-iii, a freely accessible critical care database</article-title>
          ,
          <source>Scientific data 3</source>
          (
          <year>2016</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>G.</given-names>
            <surname>Calarota</surname>
          </string-name>
          ,
          <article-title>Domain-specific word embeddings for ICD-9-CM classification</article-title>
          ,
          <source>Ph.D. thesis,</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>