<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ICB-UMA at CLEF e-Health 2020 Task 1: Automatic ICD-10 coding in Spanish with BERT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guillermo Lopez-Garc a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose M. Jerez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francisco J. Veredas</string-name>
          <email>franveredasg@uma.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departamento de Lenguajes y Ciencias de la Computacion, ETSI Informatica, Universidad de Malaga</institution>
          ,
          <addr-line>Malaga</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This working notes paper presents our contribution to the CLEF eHealth 2020 Task 1. Our team has participated in the CodiEsp-D subtask, the rst shared task consisted in the automatic clinical coding of medical cases in Spanish, annotated with ICD-10-CM codes. We tackled the task as a multi-label classi cation problem using BERT model [4]. With the aim of leveraging all the language modeling capacities of the deep bidirectional encoder architecture of BERT, we developed a tailored approach to annotate short fragments of text extracted from the long clinical cases present in the CodiEsp corpus and use them as input to the model. Two publicly available Spanish versions of BERT, namely BETO [3] and BERT-SciELO [1], were ne-tuned on the CodiEsp-D corpus extended by a set of abstracts annotated with ICD-10 codes, following our fragment-based classi cation approach. BERT-SciELO, a BERT-Base model pre-trained from scratch on an unlabeled corpus of biomedical articles in Spanish, achieved the best results among our three submitted systems, obtaining a nal Mean Average Precision (MAP) metric score of 0.482 on the evaluation set.</p>
      </abstract>
      <kwd-group>
        <kwd>Clinical coding</kwd>
        <kwd>Spanish clinical cases</kwd>
        <kwd>BERT si cation</kwd>
        <kwd>Transfer learning</kwd>
        <kwd>Clinical NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The increasingly adoption of electronic health records (EHRs) as a key
component in many hospital information systems across the globe has posed a series
of questions to the scienti c community that remain partially unresolved. One
of the main issues is how to e ectively leverage the information stored in the
system to improve patient care. EHRs store heterogeneous data in a wide variety
of formats, including free-text documents like clinical notes [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. These medical
textual representations contain crucial patient information related to diagnosis,
treatments, or procedures. However, their unstructured nature makes it specially
challenging to extract the relevant medical concepts from the data.
      </p>
      <p>
        Automatic clinical coding is the task of transforming unstructured clinical
text into a structured format, following standard coding terminologies and using
computational methods [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The produced structured data can be subsequently
used for medical billing, conducting epidemiological studies, exchanging
information between medical institutions, performing statistical analysis, and many
other purposes, with the additional advantage of not requiring any human
intervention throughout the coding process. Given the importance of storing natural
language descriptions of medical cases in modern EHR systems, automatic
clinical coding constitutes an essential task in the process of extracting valuable
information from EHR data, improving many aspects of clinical care.
      </p>
      <p>
        Historically, natural language processing (NLP) techniques have been applied
to the problem of clinical coding [
        <xref ref-type="bibr" rid="ref15 ref17 ref18 ref8">15, 18, 17, 8</xref>
        ]. However, most of the previous
works only focus on English text, as the availability of corpora annotated with
clinical coding information and additional linguistic resources in languages other
than English is scarce. With the intention of overcoming this issue, over the past
four years, the CLEF eHealth Lab has organised a series of clinical coding shared
tasks on non-English or multilingual corpora. Concretely, CLEF eHealth 2016
Task 2 focused on the assignment of International Statistical Classi cation of
Diseases and Related Health Problems (ICD-10) codes to French free-text death
certi cates, using CepiDC corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In 2017, the CLEF eHealth Task 1
continued with automatic coding of death certi cates, but turning the challenge into a
bilingual task, using both CepiDC (French) and CDC (English) annotated
corpora [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. One year later, the CLEF eHealth 2018 Task 1 explored a multilingual
clinical coding challenge, with the assignment of ICD-10 codes to death reports
in French, Hungarian and Italian [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Finally, last year, the CLEF eHealth 2019
Task 1 again focused on the assignment of ICD-10 codes, but instead of using
death reports, non-technical summaries of animal experiments in German were
employed [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        This year, CLEF eHealth 2020 Task 1 corresponds to the CodiEsp track [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
the rst shared task consisted in the automatic coding of clinical cases in
Spanish, using the Spanish version of ICD-10 (CIE-10) and the CodiEsp corpus, a
synthetic corpus of 1K clinical cases in Spanish manually curated by the
organisers of the task. The CodiEsp track is composed of three di erent subtasks:
CodiEsp Diagnosis (CodiEsp-D) subtask, CodiEsp Procedure (CodiEsp-P)
subtask and Explainable AI (CodiEsp-X) subtask. Given a free-text clinical case in
Spanish, CodiEsp-D subtask requires assigning a set of diagnosis
codes|ICD10-CM or CIE-10 Diagnostico in Spanish|to the medical document, whereas in
CodiEsp-P subtask procedures codes|ICD-10-PCS or CIE-10 Procedimiento|
are predicted for each clinical case. On the other hand, systems participating in
CodiEsp-X subtask are required to predict both diagnosis and procedures codes,
but also to provide references in the text justifying the coding predictions.
      </p>
      <p>
        In this work, we present our contribution to the CLEF eHealth 2020 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
where our team has participated in the CodiEsp-D subtask. We have tackled the
problem as a multi-label text classi cation task using BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a contextualized
neural language model that achieved state-of-the-art results on eleven distinct
NLP tasks, and was employed by the best performing team in the CLEF eHealth
2019 Task 1 [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Since BERT was specially designed to process short fragments
of text|in contrast with the long clinical notes present in the CodiEsp corpus|,
we have developed a tailored approach to turn the text classi cation task into a
short segments text classi cation problem, in order to leverage all the predictive
capabilities of the BERT model when applied to the CodiEsp corpus. In this way,
each medical document from the corpus was split into short fragments of text.
Then, using the information available for the CodiEsp-X subtask, we annotated
each short fragment with ICD-10-CM codes information, and used the annotated
segments as input to the model. Once the probabilities for individual codes were
predicted by the model on every fragment of a document, a maximum probability
criterion was used to obtain the probability of each ICD-10 diagnosis code at the
document level. Finally, for every document, a list of codes ordered by con dence
or relevance was produced, which was used by the organisers to evaluate the
performance of the participating systems. For reproducibility purposes, all the
code generated to implement our approach is publicly available at https://
github.com/guilopgar/CLEF-2020-CodiEsp.
      </p>
      <p>The rest of the paper is organised as follows. In Section 2, a short description
of the CodiEsp corpus is given, as well as the details of our proposed strategy
to tackle the CodiEsp-D subtask are described. We analyze the obtained results
in Section 3, whereas Section 4 presents some conclusions and perspectives for
future work.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Materials and Methods</title>
      <p>Corpora
According to the task organisers1, the CodiEsp corpus comprises 1K clinical
case studies covering a broad diversity of medical subjects, such as pneumology,
urology, cardiology or oncology. The entire corpus was randomly split into three
di erent subsets, the training (500 documents), development (250 documents)
and test (250 documents) sets. Also, jointly with the test set, a background set
of 2751 clinical cases was released. The latter had the intention of guaranteeing
that teams participating in the CodiEsp track could not do manual corrections,
since submissions to the task had to include predictions for the 3001 medical
documents present in both the background and test sets, but they were only
evaluated on the test set.</p>
      <p>For the CodiEsp-D subtask, annotation tables containing the assignment of
ICD-10-CM codes to the medical documents in the corpus were also available.
Furthermore, along with the CodiEsp corpus, a set of additional documents
was provided2. The supplementary documents correspond to Spanish abstracts
obtained from LILACS3 and IBECS4 biomedical literature databases, which were
1 https://temu.bsc.es/codiesp/index.php/category/data/
2 https://zenodo.org/record/3606662#.XvxBT59fg8o
3 https://lilacs.bvsalud.org/es/
4 http://ibecs.isciii.es/
annotated with ICD-10 codes information. From the whole set of abstracts, we
only selected those texts annotated with ICD-10-CM codes present in the list of
valid codes for the CodiEsp track supplied by the organisers5.</p>
      <p>
        Table 1 contains a basic description of the CodiEsp-D corpus (training,
development and test subsets) as well as the additional abstracts annotated with
diagnosis codes information. As it is shown in the table, CodiEsp-D is a
considerably challenging task, given the limited number of clinical cases (1000) and the
large number of unique codes (2557) present in the corpus, including 363 codes
that are only present in the test subset. Additionally, the number of codes
annotations is also scarce, resulting in a highly imbalanced multi-label classi cation
problem, where for each code, the number of negative samples clearly surpasses
the number of positive cases, i.e. number of documents annotated with a certain
code. For these reasons, we decided to expand the CodiEsp-D corpus using the
additional set of abstracts. Considering that most of the diagnosis codes (2153
out of 2984) included in the abstracts corpus are not present in the CodiEsp-D
annotations, we experimented with two distinct ways of expanding the
training and development CodiEsp-D corpora: either using all available abstracts, or,
alternatively, solely using the abstracts annotated with ICD-10-CM codes
contained in the training and development sets, leading to a reduced version of the
abstracts corpus comprising 115457 documents and 160652 codes annotations,
from which 733 are unique ICD-10 codes.
We have tackled the CodiEsp-D challenge using BERT model [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. BERT is a
contextual language representation model, based on the encoder part of the
Transformer architecture [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], designed to extract deep bidirectional contextual
representations both at the token and sentence level. The model can be
pretrained on an unlabeled corpus in an unsupervised manner using two language
modeling objectives, namely next sentence prediction and masked language
modeling. Unlike other contextual language models, BERT can be transferred to be
      </p>
      <sec id="sec-2-1">
        <title>5 https://zenodo.org/record/3706838#.XvxD3J9fg8o</title>
        <p>used in a downstream task, by ne-tuning the whole architecture in a supervised
way, then adapting all its pretrained weights to solve a speci c task.</p>
        <p>
          BERT has gained a lot of attention in the NLP community, including in
biomedical and clinical domains, where BERT-based approaches have obtained
state-of-the-art results in a wide range of tasks [
          <xref ref-type="bibr" rid="ref21 ref7">7, 21</xref>
          ]. In this work, we have
experimented with two Spanish versions of the BERT-Base architecture: BETO [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
and BERT-SciELO [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] models. While both models were pretrained from scratch
on unlabeled Spanish corpora, BETO was trained on a compilation of general
domain texts6, whereas BERT-SciELO was pre-trained on a corpus of biomedical
articles retrieved from SciELO7.
        </p>
        <p>
          In contrast to sequential models such as recurrent neural networks (RNNs),
for computational feasibility reasons, Transformer architectures cannot deal with
long input sequences of variable length, since for the self-attention layers the
complexity is quadratic on the length of the input sequence [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. For instance, in
the original implementation of BERT|which uses WordPiece [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] tokenization to
subdivide each input token into further sub-token units|, the maximum input
sub-tokens sequence length is 512. This constitutes an important limitation when
dealing with document or long-text classi cation tasks like CodiEsp-D, where
the sub-tokens sequence size of many clinical cases is clearly above the maximum
length supported by BERT.
        </p>
        <p>
          Over the last year, a few works have already explored di erent strategies to
overcome this limitation. The most straightforward approach is to use a text
truncation method, like the one adopted by the CLEF eHealth 2019 Task 1 best
performing team [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], which simply consisted in using the rst 5108 sub-tokens
of each document as input to the model. In [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], additional truncation strategies
were considered, and the best performing method was to select the rst 128 and
the last 382 sub-tokens from each input text. Authors hypothesized that the
most relevant information in a document appears at the beginning and at the
end of it, which may be the case for the texts present in the corpora analyzed
in that work, speci cally the movie review IMDb corpus and the Chinese Sogou
news articles dataset. However, in the CodiEsp-D corpus, clinical information
relevant to solve the task may be spread anywhere within the clinical cases,
hence eliminating parts of the documents may not be the most appropriate
strategy for our needs.
        </p>
        <p>
          Another recent work explored a di erent approach that does not make use of
any truncation method to adapt BERT to solve a text classi cation task. In this
way, authors in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] proposed to segment the input texts into smaller parts, and
then ne-tune the BERT model on a segment-level supervised task, assigning to
each segment the labels associated with the entire document where the segment
comes from. If we applied the same strategy to solve the CodiEsp-D multi-label
classi cation task, we would annotate all the segments obtained from a single
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>6 https://github.com/josecannete/spanish-corpora 7 https://www.scielo.org/es/</title>
        <p>8 Since BERT always adds two special tokens ([CLS] and [SEP]) at the rst and last
positions, respectively, of an input sequence.
clinical case with the same ICD-10-CM codes present in the complete document.
This constitutes a problematic situation, as many fragments would be annotated
with diagnosis codes which are not represented within the fragment, but appear
in other segments from the same clinical document.</p>
        <p>
          With the aim of adopting a strategy to adjust BERT to the distinctive
features of the CodiEsp-D subtask, we have developed a three-phases custom
approach that transforms the multi-label long-text classi cation task into a
multilabel short-fragment classi cation problem. Unlike [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], using the annotations
available for CodiEsp-X named-entity recognition (NER) subtask, for each
fragment we only assign the labels occurring within the speci c fragment, avoiding
misleading the model by assigning codes to a fragment that are present in other
parts of the document. In the next paragraphs, the three stages of the developed
approach are described.
        </p>
        <p>Splitting each clinical case into fragments. As indicated before, BERT
supports input sequence lengths up to a maximum value of N (N = 512 in
the original implementation). For this reason, for every clinical case in the
CodiEsp-D corpus, after performing WordPiece tokenization, we split the
resulting sub-tokens sequence w = (w1; w2; :::; wk), of length k, into a sequence
of m = dk=(N 2)e contiguous sub-token fragments f = (f1; f2; :::; fm) =
((w1; :::; wN 2); (w(N 2)+1; :::; w2 (N 2)); :::; (w(m 1) (N 2)+1; :::; wk)), in which
the m 1 rst fragments f1; ::; fm 1 have a length of N 2 sub-tokens while
the last fragment fm contains the remaining nal sub-tokens of the document
(considering that the tokens [CLS] and [SEP] have to be later added at the
beginning and the end of each fragment, respectively, to get sequences of size N
sub-tokens that will constitute the input to BERT).</p>
        <p>Annotating text fragments with ICD-10-CM codes. Annotations
available for the CodiEsp-D subtask contain the assignment of a set of diagnosis
codes to each clinical document (see Fig. 1A). On the other hand, in
CodiEspX subtask|which uses the same CodiEsp corpus as the CodiEsp-D subtask|,
codes annotations include an extra eld indicating the reference in the text
that explains the coding assignment (see Fig. 1B). In this way, using the
information provided for the CodiEsp-X subtask, we managed to annotate each of
the fragments|resulting from the splitting procedure|exclusively with those
ICD-10-CM codes annotations whose text references were contained within the
fragment. Thus, given a certain fragment fl and the set of its associated
diagnosis codes Cfl , as well as a CodiEsp-X annotation ai consisted in a code ci and
a text reference spanning a group of continuous or discontinuous sub-tokens9
ti = (wi1; wi2; :::; wi ), we annotated fl with ci when any of the sub-tokens of
ti was contained inside the limits of fl, i.e. ci 2 Cfl () 9wij 2 ti j wijs
fls ^ wije fle , where wijs and wije stand for the starting and ending character
positions of the sub-token wij, respectively, whereas fls and fle represent the
9 Considering that annotations text references were also tokenized using WordPiece.
starting and ending character positions of the fragment fl. To illustrate the
annotation process, in Fig. 2, we show the annotations generated for the fragments
extracted from the S1139-76322012000400011-1 CodiEsp document using the
information available for the same document in the CodiEsp-X subtask (see
Fig. 1B).</p>
        <p>Obtaining codes probabilities at document level. Using the sub-tokens
fragments from the clinical cases in the CodiEsp corpus annotated with
ICD10-CM codes, we trained the BERT model on a fragment-level classi cation
problem. Since many distinct codes may be assigned to the same medical
document, we treated the task as a multi-label classi cation problem. To ne-tune the
whole architecture of BERT on a supervised learning task at text-level, the
representation produced by the model for the initial [CLS] token was fed into a nal
output layer for classi cation. As we were dealing with a multi-label task|which
is equivalent to multiple independent binary classi cation tasks|, we used the
sigmoid activation function and D output units, with D representing the number
of unique diagnosis codes occurring in the texts used to train the model. Hence,
given a text fragment as input to the model, the produced output vector could
be interpreted as the probability of each code to occur within the input
fragment. However, CodiEsp-D task was a document classi cation problem, and the
evaluation of the participating systems was performed at document level.
Accordingly, using a maximum probability criterion, we post-processed the model
predicted probabilities to obtain codes probabilities at document level. Thus,
given a sequence f containing all fragments generated from a single document d
as input, the model produces the output probability matrix M 2 Rjfj D. Then,
selecting the maximum probability value across every column of M , a nal
vector p 2 RD of codes probabilities at document level is generated, which contains
the probability of each code to appear in d. Using this method, we were able to
produce, for each clinical case, a list of D distinct codes sorted in descending
order according to their probability values predicted by the model, which was
nally used by the organisers to evaluate the performance of the classi cation
system.</p>
        <p>It should be noted that, in the case of the additional abstracts corpus (see
Section 2.1), the available codes annotations did not contain any extra eld
indicating the text reference that supports the coding assignment. Therefore, our
fragment-based custom approach could not be applied to the abstracts corpus,
but only to the CodiEsp-D corpus enriched with the information available for
the CodiEsp-X subtask. For this reason, exclusively the abstracts that, after
performing WordPiece tokenization, contain a maximum number of N 2
subtokens were used to expand the CodiEsp-D corpus.
2.3</p>
        <p>Experiments
We experimented with two Spanish versions of the BERT-Base model, namely
BETO and BERT-SciELO. Since the vocabulary of the BERT-SciELO model
R30.0</p>
        <p>R31.0</p>
        <p>R31.9</p>
        <p>R50.9</p>
        <p>B65.0</p>
        <p>B65.9
Presentamos el caso de un varón de 11 años original de Gambia
que consulta por hematuria macroscópica de predominio al final
de la micción y disuria de un año de evolución; sin antecedentes
de fiebre. Al realizar la anamnesis refieren un viaje reciente a su
país de origen y en el transcurso del mismo varios baños en
lagos de la región. La exploración física es anodina.</p>
        <p>Ante la sospecha clínica de bilharzhiasis, se contacta con el
Servicio de Microbiología del hospital de referencia, donde
indican recogida de orina de tres días consecutivos,
preferentemente del mediodía y del final de la micción (momento
en que es máxima la excreción de huevos) y se solicita estudio
mediante ecografía renovesical. El estudio microbiológico
demostró huevos de Schistosoma haematobium.</p>
        <p>La ecografía vesical practicada puso de manifiesto un
engrosamiento parietal que llegaba a alcanzar un grosor máximo
de 9 mm en un radio de 20 mm, lo que sugiere
esquistosomiasis, por lo que se prescribió tratamiento con
prazicuantel.</p>
        <p>B Presentamos el caso de un varón de 11 años original de Gambia que</p>
        <p>R31.9</p>
        <p>R31.0
consulta por hematuria macroscópica de predominio al final de la</p>
        <p>R50.9
R30.0
B65.0</p>
        <p>B65.9
micción y disuria de un año de evolución; sin antecedentes de fiebre. Al realizar la
anamnesis refieren un viaje reciente a su país de origen y en el transcurso del
mismo varios baños en lagos de la región. La exploración física es anodina.</p>
        <p>B65.9
Ante la sospecha clínica de bilharzhiasis, se contacta con el Servicio de
Microbiología del hospital de referencia, donde indican recogida de orina de tres
días consecutivos, preferentemente del mediodía y del final de la micción
(momento en que es máxima la excreción de huevos) y se solicita estudio
mediante ecografía renovesical. El estudio microbiológico demostró huevos de
Schistosoma haematobium.</p>
        <p>La ecografía vesical practicada puso de manifiesto un engrosamiento parietal que
llegaba a alcanzar un grosor máximo de 9 mm en un radio de 20 mm, lo que
sugiere esquistosomiasis, por lo que se prescribió tratamiento con prazicuantel.</p>
        <p>
          Fig. 2. Illustration of the text fragments ICD-10-CM codes annotations
obtained after applying the rst two phases of our developed approach to the
S1139-76322012000400011-1 clinical case from the CodiEsp training corpus. To
generate the sub-tokens fragments, we used the WorPiece tokenizer of the BETO model,
setting a maximum fragment length of N 2 = 78 sub-tokens.
did not include any punctuation character, we used a pre-processed version of the
expanded CodiEsp-D corpus (together with the additional abstracts) in which
punctuation marks were substituted by spaces. In the case of BETO, the raw
text from the documents was employed, as punctuation marks were contained
in the vocabulary of the model. Regarding the hardware resources employed,
all experiments were executed on a single GeForce GTX 1080 Ti 11 GB GPU.
Given the limitations imposed by the hardware, we used a maximum input
sequence length of N = 230 for the BERT-SciELO and a value of N = 275 for
the BETO model, with both models having 184M trainable weights. Finally,
with respect to the hyperparameters of the models, for ne-tuning, we used
RAdam [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] with learning rate of 3 10 5, a batch size of 16 and the number
of epochs were experimentally determined on the CodiEsp-D development set
using early-stopping, with an upper limit of 40 epochs.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In this section, we present the results obtained by our team, ICB-UMA, at the
CodiEsp-D subtask. Task organisers allowed each participating team to submit
up to 5 runs of their classi cation systems. We submitted three distinct runs.
The rst two submissions (ICB-UMA-run1 and ICB-UMA-run2) corresponded
to the BERT-SciELO model ne-tuned on the CodiEsp-D training and
development corpora expanded using additional abstracts annotated with ICD-10-CM
codes present in the training and development sets. Thus, the extended
corpus contained 2194 distinct diagnosis codes|the number of unique ICD-10-CM
codes present in the training and development corpora (see Table 1)|, having a
nal BERT-SciELO model with an output classi cation layer of D = 2194 units.
Since exclusively the abstracts containing a maximum number of 230 2
WordPiece sub-tokens were considered for the BERT-SciELO model (see Section 2.2),
a nal set of 66620 abstracts with 92239 codes annotations were employed to
expand the training and development corpora. Submission ICB-UMA-run2 used a
dropout probability of 0:1 on the output layer of BERT-SciELO model, whereas
in submission ICB-UMA-run1 no dropout was used on the output layer. For
its part, in submission ICB-UMA-run3 the BETO model was ne-tuned on the
CodiEsp-D training and development sets extended using all available abstracts
with a maximum number of 275 2 WordPiece sub-tokens, which comprised a
total of 87871 documents and 208076 diagnosis codes annotation, from which
2204 were unique ICD-10-CM codes, i.e. codes that were not present in the
training and development corpora. Finally, for submission ICB-UMA-run3 the
output layer of the BETO model had D = 4398 units, with no dropout on the
nal layer. Apart from these three runs, we also experimented with the
BERTSciELO model ne-tuned on the CodiEsp-D corpus expanded using all available
abstracts, as well as with the BETO model ne-tuned on the CodiEsp-D corpus
extended using additional abstracts annotated with training and development
diagnosis codes. However, these two approaches obtained worse results on the
development set, and we submitted the three strategies that best performed on
the development corpus.</p>
      <p>
        Table 2 and Table 3 show the predictive performance of our three di erent
submitted runs on the CodiEsp test corpus, as a result of the evaluation
performed by the CodiEsp track organisers. Namely, Table 2 presents the results
obtained according to the o cial evaluation metric of the CodiEsp-D subtask,
i.e. the Mean Average Precision (MAP). The second column of the table
contains the MAP values calculated considering all codes present in the CodiEsp
test subset, whereas the results shown in the third column (MAP codes ) were
computed taking into account only the predictions for the test codes that were
also present in the training and development subsets. Finally, the fourth
column (MAP30 ) measures MAP-at-k metric (MAP@k), with k equals 30, while
the last column (MAP30 codes) measures MAP@30 considering only the test
codes that were present in the training and development corpora (as in the third
column of the table). According to the results observed in Table 2, the
BERTSciELO-based model outperformed the BETO-based classi er on the CodiEsp-D
subtask, since both the ICB-UMA-run1 and ICB-UMA-run2 systems achieved
higher values for all analyzed metrics than the ICB-UMA-run3 submitted
system. The best performance is obtained by the BERT-SciELO model when no
dropout is used on the output layer, as the ICB-UMA-run1 results slightly
surpassed the performance of the ICB-UMA-run2 system across the four examined
evaluation metrics. To summarize, we can say that, according to our obtained
results, the BERT-SciELO model outperformed the BETO on the CodiEsp-D
predictive task. Therefore, pre-training the BERT-Base architecture from scratch
on an unlabeled corpus of Spanish biomedical articles, rather than using a general
domain Spanish corpus, leads to a better performance of the model on a clinical
coding task in the context of Spanish medical narrative. The results obtained in
this work for the CodiEsp-D subtask support the hypothesis already explored
in previous works [
        <xref ref-type="bibr" rid="ref2 ref21">2, 21</xref>
        ], claiming that a clinical domain-speci c BERT yields
superior performance on medical classi cation problems than a general domain
version of the model.
      </p>
      <p>
        On the other hand, to perform a larger analysis of the results, the task
organisers evaluated the performance of the systems according to a set of additional
metrics, other than the o cial ones. In this way, in Table 3, the second, third
and fourth columns show the calculated values using precision (P ), recall (R)
and the F1 score (F1 ) metrics, respectively, considering all codes present in the
CodiEsp-D test set, while the fth (P codes), sixth (R codes) and seventh (F1
codes) columns contain the results computed using the same three metrics but
taking into consideration only the codes present in the training and development
subsets. Finally, in the last three columns (P cat, R cat and F1 cat ), the
previous metrics are used to evaluate the submitted predictions at the hierarchical
category level of the ICD-10-CM codes contained in the test set. As it can be
observed from Table 3, for the precision and consequently for the F1 score, our
three submitted systems obtained extremely poor values, though for the recall
the computed values are abnormally high. The reason for this is that, with the
aim of maximizing the score obtained for the o cial evaluation metric, i.e. MAP,
for each test document we submitted all codes considered by the model|2194
codes in the case of ICB-UMA-run1 and ICB-UMA-run2 and 4398 codes for
the ICB-UMA-run3 submission|ordered by their predicted probability of
occurrence (see Section 2.2). If we had optimized precision, recall and F1 score
metrics instead of MAP, in place of submitting all codes, a decision threshold
would have been de ned to select solely a subset of the codes according to their
predicted probabilities.
In this paper, we present our contribution to the CodiEsp-D subtask from the
CodiEsp track [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] of CLEF eHealth 2020 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The shared task proposes the
automatic assignment of ICD-10-CM codes to Spanish clinical cases. The scarce
number of medical cases present in the CodiEsp corpus in combination with the
large quantity of unique diagnosis codes used to annotate the documents, make
CodiEsp-D a considerably challenging task.
      </p>
      <p>
        We have tackled the challenge as a multi-label text classi cation task using
BERT model [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A fragment-based classi cation approach was developed in
order to take advantage of all the predictive capacity of BERT when receiving the
long-text clinical cases contained in the CodiEsp corpus as input to the model.
Our strategy consisted in using the available annotations for the CodiEsp-X
NER subtask to turn the CodiEsp-D multi-label document classi cation task
into a multi-label short-fragment classi cation problem. We experimented with
two publicly available Spanish versions of the BERT model, ne-tuned on the
CodiEsp-D corpus expanded using a set of available abstracts annotated with
ICD-10-CM codes. BERT-SciELO [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a BERT-Base architecture pre-trained on
a corpus of biomedical articles in Spanish, yielded the best performance among
our three submitted systems, obtaining a MAP score of 0:482 on the evaluation
set. The obtained results in this work reinforced the idea that a medical domain
version of BERT achieves higher performance on clinical classi cation tasks than
nonspeci c domain versions of the model.
      </p>
      <p>In future works, we will try to enhance the developed fragment-based
classi cation strategy to further improve the obtained results on the CodiEsp-D
subtask. For instance, when splitting each clinical case into fragments, we could
perform the text segmentation at the sentence level, producing fragments
comprising a sequence of sentences with a complete semantic meaning. On the other
hand, given the superior performance observed from the BERT-SciELO model,
it is worth investigating whether alternative clinical-speci c versions of BERT
pre-trained on Spanish medical corpora more similar to the CodiEsp corpus
could increase the results further. Additionally, due to the widespread
adoption of BERT in multi-lingual setups across domains, we could also explore the
pre-training of the model on a multi-lingual medical corpus. This would permit
the creation of enormous medical corpora comprising clinical documents written
in many di erent languages. Because of the sub-word vocabulary employed by
BERT and the common etymology of numerous medical terms across distinct
languages, multi-lingual clinical corpora could serve as a valuable source of data
to pre-train BERT models. The resulting BERT's deep bidirectional architecture
could leverage its language modeling capabilities to produce e ective contextual
representations that could be used in applications within a vast number of
medical NLP information-extraction problems.</p>
      <p>Acknowledgments. This work was partially supported by the project
TIN201788728-C2-1-R, MINECO, Plan Nacional de I+D+I, and I Plan Propio de
Investigacion y Transferencia of the Universidad de Malaga.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akhtyamova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mart</surname>
            <given-names>nez</given-names>
          </string-name>
          , P.,
          <string-name>
            <surname>Verspoor</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cardi</surname>
          </string-name>
          , J.:
          <article-title>Testing Contextualized Word Embeddings to Improve NER in Spanish Clinical Case Narratives</article-title>
          .
          <source>Preprint (Version</source>
          <volume>1</volume>
          ) available at Research Square (
          <year>2020</year>
          ). https://doi.org/10.21203/rs.2.22697/v1
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Alsentzer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boag</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>W.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jindi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McDermott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Publicly available clinical BERT embeddings</article-title>
          .
          <source>In: Proceedings of the 2nd Clinical Natural Language Processing Workshop</source>
          . pp.
          <volume>72</volume>
          {
          <fpage>78</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>W19</fpage>
          -1909
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Can~ete, J.,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>J.: Spanish</given-names>
          </string-name>
          <string-name>
            <surname>Pre-Trained BERT</surname>
          </string-name>
          Model and
          <article-title>Evaluation Data</article-title>
          . In: to appear
          <source>in PML4DC at ICLR</source>
          <year>2020</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          . pp.
          <volume>4171</volume>
          {
          <fpage>4186</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota (
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>N19</fpage>
          -1423
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miranda-Escalada</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krallinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Saez Gonzales,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2020</article-title>
          . In: Arampatzis,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kanoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Tsikrika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Vrochidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Joho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Lioma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Eickho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Neveol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            , andNicola Ferro, L.C. (eds.)
            <surname>Experimental IR Meets Multilinguality</surname>
          </string-name>
          , Multimodality, and
          <source>Interaction: Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ).
          <source>Lecture Notes in Computer Science</source>
          , vol.
          <volume>12260</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krikun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thorat</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viegas</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wattenberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Google's multilingual neural machine translation system: Enabling zero-shot translation</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>5</volume>
          ,
          <issue>339</issue>
          {
          <fpage>351</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoon</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>So</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>BioBERT: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          .
          <source>Bioinformatics</source>
          <volume>36</volume>
          (
          <issue>4</issue>
          ),
          <volume>1234</volume>
          {
          <fpage>1240</fpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1093/bioinformatics/btz682
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <source>Automated ICD-9 Coding via A Deep Learning Approach. IEEE/ACM Transactions on Computational Biology and Bioinformatics</source>
          <volume>16</volume>
          (
          <issue>4</issue>
          ),
          <volume>1193</volume>
          {
          <fpage>1202</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>J</given-names>
            ., Han, J
          </string-name>
          .:
          <article-title>On the Variance of the Adaptive Learning Rate and Beyond (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Miranda-Escalada</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez-Agirre</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Armengol-Estape</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krallinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020</article-title>
          . In: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          . CEUR Workshop Proceedings (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>R.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavergne</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.: CLEF</given-names>
          </string-name>
          <article-title>eHealth 2017 multilingual information extraction task overview: ICD10 coding of death certi cates in English and French</article-title>
          . In:
          <article-title>Proc of CLEF eHealth Evaluation lab</article-title>
          . Dublin, Ireland (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavergne</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Clinical information extraction at the CLEF eHealth evaluation lab 2016</article-title>
          .
          <article-title>In: Proc of CLEF eHealth Evaluation lab</article-title>
          . Evora,
          <string-name>
            <surname>Portugal</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grippo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgand</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orsi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelikan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramadier</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.: CLEF</given-names>
          </string-name>
          <article-title>eHealth 2018 Multilingual Information Extraction Task Overview: ICD10 Coding of Death Certi cates in French, Hungarian and Italian</article-title>
          . In:
          <article-title>Proc of CLEF eHealth Evaluation lab</article-title>
          . Avignon, France (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Neves</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Butzke</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Dorendahl,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Leich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Hummel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            , Schonfelder, G.,
            <surname>Grune</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2019</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.E.</surname>
          </string-name>
          , Muller, H. (eds.)
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>CLEF 2019. Lecture Notes in Computer Science</source>
          , vol.
          <volume>11696</volume>
          . Springer, Cham (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Pakhomov</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buntrock</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          :
          <article-title>Automating the Assignment of Diagnosis Codes to Patient Encounters Using Example-based and Machine Learning Techniques</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>13</volume>
          (
          <issue>5</issue>
          ),
          <volume>516</volume>
          {
          <fpage>525</fpage>
          (
          <year>2006</year>
          ). https://doi.org/10.1197/jamia.M2077
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Pappagari</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zelasko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villalba</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carmiel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dehak</surname>
          </string-name>
          , N.:
          <article-title>Hierarchical Transformers for Long Document Classi cation</article-title>
          .
          <source>In: 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)</source>
          . pp.
          <volume>838</volume>
          {
          <issue>844</issue>
          (
          <year>2019</year>
          ). https://doi.org/10.1109/ASRU46091.
          <year>2019</year>
          .9003958
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Perotte</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pivovarov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natarajan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiskopf</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
          </string-name>
          , N.:
          <article-title>Diagnosis code assignment: models and evaluation metrics</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>21</volume>
          (
          <issue>2</issue>
          ),
          <volume>231</volume>
          {
          <fpage>237</fpage>
          (
          <year>2013</year>
          ). https://doi.org/10.1136/amiajnl-2013
          <source>-002159</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pestian</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brew</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matykiewicz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovermale</surname>
            ,
            <given-names>D.J.</given-names>
          </string-name>
          , Johnson, N.,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duch</surname>
            ,
            <given-names>W.: A Shared</given-names>
          </string-name>
          <string-name>
            <surname>Task</surname>
          </string-name>
          <article-title>Involving Multi-Label Classi cation of Clinical Free Text</article-title>
          .
          <source>In: Proceedings of the Workshop on BioNLP 2007: Biological, Translational, and Clinical Language Processing</source>
          . p.
          <volume>97</volume>
          {
          <fpage>104</fpage>
          . BioNLP '07,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Sanger,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Kittner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Leser</surname>
          </string-name>
          ,
          <string-name>
            <surname>U.</surname>
          </string-name>
          :
          <article-title>Classifying German Animal Experiment Summaries with Multi-lingual BERT at CLEF eHealth 2019 Task 1</article-title>
          . In: CLEF (Working Notes) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Shickel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tighe</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bihorac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashidi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Deep</surname>
            <given-names>EHR</given-names>
          </string-name>
          :
          <article-title>A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis</article-title>
          .
          <source>IEEE Journal of Biomedical and Health Informatics</source>
          <volume>22</volume>
          (
          <issue>5</issue>
          ),
          <volume>1589</volume>
          {
          <fpage>1604</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Si</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Enhancing clinical concept extraction with contextual embeddings</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>26</volume>
          (
          <issue>11</issue>
          ),
          <volume>1297</volume>
          {
          <fpage>1304</fpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1093/jamia/ocz096
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. Stan ll, M.H.,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fenton</surname>
            ,
            <given-names>S.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenders</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hersh</surname>
            ,
            <given-names>W.R.:</given-names>
          </string-name>
          <article-title>A systematic literature review of automated clinical coding and classi cation systems</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>17</volume>
          (
          <issue>6</issue>
          ),
          <volume>646</volume>
          {
          <fpage>651</fpage>
          (
          <year>2010</year>
          ). https://doi.org/10.1136/jamia.
          <year>2009</year>
          .001024
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qiu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>How to Fine-Tune BERT for Text Classi cation? In: Chinese Computational Linguistics</article-title>
          .
          <source>CCL 2019. Lecture Notes in Computer Science</source>
          , vol.
          <volume>11856</volume>
          . Springer, Cham (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          . In: Guyon,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.V.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Vishwanathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Garnett</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pp.
          <volume>5998</volume>
          {
          <fpage>6008</fpage>
          . Curran Associates, Inc. (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>