<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Team NOVA LINCS @ BIOASQ12 MultiCardioNER Track: Entity Recognition with Additional Entity Types</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rodrigo Gonçalves</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>André Lamúrias</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NOVA LINCS, NOVA School of Science and Technology</institution>
          ,
          <addr-line>Lisbon</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the contribution of the NOVA LINCS team to the task MultiCardioNER at BioASQ12. BioASQ is a long-running challenge that focuses mostly on biomedical semantic indexing and question answering (QA). The specific task of MultiCardioNER focuses on the multilingual adaptation of clinical NER systems to the cardiology domain. We leverage a state-of-the-art spanish pre-trained model to perform NER on the DisTEMIST and DrugTEMIST datasets provided by this task, aided by the additional clinical cases corpus, CardioCCC, for validation. Experiments were done both with just the entity type to which each of those datasets refers to, and with the additional entity types from other datasets that use the same documents, in order to determine if any advantage could be obtained by leveraging the knowledge of the additional entities. The models trained on the combined dataset achieved a very slight and not significant boost in F1-score when compared to their one-entity counterparts, with all of them lacking heavily in recall, due to pre- and post-processing errors. However, one run achieved the highest precision of the task (0.9242). Code to reproduce our submission is available at https://github.com/Rodrigo1771/BioASQ12-MultiCardioNER-NOVALINCS.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Cardiology</kwd>
        <kwd>Clinical Cases</kwd>
        <kwd>Named Entity Recognition</kwd>
        <kwd>Language Models</kwd>
        <kwd>Transfer Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The MultiCardioNER [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] task, part of the BioASQ 2024 challenge [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], focuses on the automatic
recognition of two key clinical concept types: diseases and medications. In this task, the extraction of those
types of entities is specifically performed over cardiology clinical case documents. This represents a
worthwhile efort because, with cardiovascular diseases being the world’s leading cause of death, it’s
imperative that better automatic semantic annotation resources and systems for high impact clinical
domains such as cardiology are developed. Furthermore, this task also represents an efort to aid in the
development of systems that can perform in multiple languages, not just English, which is much needed
as the prevalence and impact of cardiovascular diseases are global, requiring robust and adaptable tools
that can cater to diverse linguistic and clinical environments.
      </p>
      <p>
        The datasets provided by each of the two MultiCardioNER subtasks, namely DisTEMIST and
DrugTEMIST, both make use of and annotate over the same set of documents, the SPACCC corpus [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Additionally, this corpus is also used by both MedProcNER [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and SympTEMIST [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. What this means
in practice is that the SPACCC corpus is annotated with four diferent types of entities: diseases for
DisTEMIST, medications for DrugTEMIST, medical procedures for MedProcNER, and symptoms for
SympTEMIST. Thus, we focused on training two models per each subtask: one solely with the entity
type of that same subtask, and another one with all four entity types. That way, we could evaluate the
benefit, or lack thereof, of having annotations regarding additional entity types. For the second subtask,
we considered only the documents in Spanish.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data</title>
      <p>This shared task utilizes two main training datasets, as previously mentioned: the DisTEMIST and
DrugTEMIST datasets. Each of these corpora is composed by the same 1000 documents that belong to</p>
      <sec id="sec-2-1">
        <title>DisTEMIST</title>
      </sec>
      <sec id="sec-2-2">
        <title>DrugTEMIST</title>
      </sec>
      <sec id="sec-2-3">
        <title>SympTEMIST</title>
      </sec>
      <sec id="sec-2-4">
        <title>MedProcNER</title>
      </sec>
      <sec id="sec-2-5">
        <title>CardioCCC-DisTEMIST</title>
      </sec>
      <sec id="sec-2-6">
        <title>CardioCCC-DrugTEMIST</title>
      </sec>
      <sec id="sec-2-7">
        <title>CombinedDataset-Train</title>
      </sec>
      <sec id="sec-2-8">
        <title>DisTEMIST-Train</title>
      </sec>
      <sec id="sec-2-9">
        <title>DrugTEMIST-Train</title>
      </sec>
      <sec id="sec-2-10">
        <title>DisTEMIST-Dev</title>
        <p>DrugTEMIST-Dev
the Spanish Clinical Case Corpus (SPACCC), a manually classified collection of clinical case reports
written in the Spanish language. While DisTEMIST is focused on extracting disease mentions, thus
having ENFERMEDAD as its sole and primary label, DrugTEMIST focuses on drug and medication
extraction, with FARMACO being its entity type. Also previously mentioned is the fact that two former
subtasks, MedProcNER and SympTEMIST, also make use of the SPACCC corpus to extract medical
procedures and symptoms, having PROCEDIMIENTO and SINTOMA as their entity types, respectively.</p>
        <p>Additionally, the organization also provided a separate cardiology clinical case reports dataset,
CardioCCC, to be used for the domain adaptation part of the task. It contains a total of 508 documents,
split into 258 documents initially intended for validation, and 250 for testing. As expected, these
documents only include annotations for disease and medication mentions, the entity types considered
for this competition. Nevertheless, we used this dataset both for training and validation of the model.
Some statistics on all of these datasets are shown in Table 1.</p>
        <p>After combining the four mentioned datasets (apart from CardioCCC) into a single dataset that
includes the annotations of all four entity types, and given the task objective of adapting clinical models
to the cardiology domain specifically, we joined this new combined dataset with CardioCCC, and split
it so that 80% of examples (sentences) were used for training and 20% for validation. This way, we can
include cardiology clinical case reports in the training of the model. Table 2 shows some statistics on
every training and validation set.</p>
        <p>Note that the three training sets are all equal between themselves in terms of the example sentences
that they contain. Likewise, the two validation sets have the same exact sentences between them. The
only thing that changes is the annotated entity types in those examples.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology and Experiments</title>
      <p>
        For this task, we leveraged the capabilities of the pre-trained model bsc-bio-ehr-eb1 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a
RoBERTabased model [7] trained on a biomedical-clinical corpus in Spanish collected from several sources
totalling more than 1 billion tokens. This model has previously been fine-tuned on similar shared tasks
like CANTEMIST [8] and PharmaCoNER [9], achieving remarkable results.
      </p>
      <p>We conducted two experiments. The first experiment compares the model when fine-tuned on the
1https://huggingface.co/PlanTL-GOB-ES/bsc-bio-ehr-es
combined dataset with the four entities against the model when fine-tuned on just the DisTEMIST
dataset. The second experiment also compares the model when fine-tuned on the combined dataset, but
this time against the model when fine-tuned on just the DrugTEMIST dataset. This brings the total of
runs submitted to four:
• Track 1 (DisTEMIST):
• Track 2 (DrugTEMIST), Spanish subset:
– 1_bsc-bio-ehr-es_distemist_4: bsc-bio-ehr-es trained on the combined dataset
with the 4 entity types (ENFERMEDAD, FARMACO, PROCEDIMIENTO, SINTOMA), and
validated on the validation set for DisTEMIST (only with ENFERMEDAD).
– 2_bsc-bio-ehr-es_distemist_1: trained and validated on the training and validation
sets, respectively, for DisTEMIST (only with ENFERMEDAD).
– 3_bsc-bio-ehr-es_drugtemist_4: bsc-bio-ehr-es trained on the combined dataset
with the 4 entity types (ENFERMEDAD, FARMACO, PROCEDIMIENTO, SINTOMA), and
validated on the validation set for DrugTEMIST (only with FARMACO).
– 4_bsc-bio-ehr-es_drugtemist_1: trained and validated on the training and validation
sets, respectively, for DrugTEMIST (only with FARMACO).</p>
      <p>During training, the best checkpoint was kept, when evaluated on the validation set after each epoch.
The model’s hyperparameters for all four runs were the following:
• Learning rate: 5e-05
• Total train batch size: 16
• Epochs: 10</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>The results for each run obtained on our validation set, both during the development of our
approach and after submitting it (the latter obtained using the oficial evaluation library 2, after it
was released) are presented in Tables 3 and 4, respectively. Additionally, the oficial gold
standard test set results are shown in Table 5. The naming convention for each run is as follows:
{run_id}\_{model_name}\_{dataset}\_{number_of_entity_types_in_training}.</p>
      <p>First of all it is important to point out the way in which we obtained our results on the validation
set (Table 3), as well as the final submitted predictions, for the runs where the model was trained on
the combined dataset (specifically runs 1 and 3), as those models naturally predict more entities than
the entity relevant for the track. Thus, on those particular runs, after obtaining the predictions (be
it the final test set predictions or the predictions used to further calculate the validation results), the
ones that corresponded to entity types that did not match the entity type of the track were ignored.
This way, only predictions of the entity type related to the track were considered. For example, in
2https://github.com/nlp4bia-bsc/multicardioner_evaluation_library</p>
      <sec id="sec-4-1">
        <title>Precision</title>
      </sec>
      <sec id="sec-4-2">
        <title>Recall F1-Score</title>
      </sec>
      <sec id="sec-4-3">
        <title>Precision</title>
      </sec>
      <sec id="sec-4-4">
        <title>Recall F1-Score</title>
        <p>run 1, the predictions related to the entity types FARMACO, PROCEDIMIENTO and SINTOMA were
ignored and only ENFERMEDAD was included in the final metric assessment seen on Table 3 and in
the submitted predictions. Likewise, in run 3, the predictions related to the entity types ENFERMEDAD,
PROCEDIMIENTO and SINTOMA were ignored and only FARMACO was considered.</p>
        <p>The significant diference between the recall scores on our internal evaluation and on the oficial
evaluation of the runs is evident. Usually, in a situation of high precision and low recall, it means
that the system is only retrieving a small portion of the relevant entities (low recall), but when it does
retrieve one, it identifies it correctly the vast majority of the time (high precision). Our first hypothesis
to try to explain what happened was that this was due to our use of the BIO tagging schema. In this
tagging schema, a sequence containing an entity, for example, of type ENFERMEDAD like "hipertensión
pulmonar severa", should be classified as "B-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD". Our
reasoning was that the model could not generalize for entities that were not present in the training
data, specially since we picked the checkpoint that obtained the best token classification and observed
some overfitting when comparing the training and validation losses, shown in Figure 1. This could lead
to the model under or over-classifying entities:
• Under classification of "hipertensión pulmonar severa" - "B-ENFERMEDAD I-ENFERMEDAD O"
(missing the "severa").
• Over classification of "estenosis píloro - duodenal de carácter extrínseco" - should be correctly
classified as "B-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD O O O"
but is overclassified as "B-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD
I-ENFERMEDAD I-ENFERMEDAD I-ENFERMEDAD" (including the additional information "de
carácter extrínseco").</p>
        <p>This under and over-classification could then lead to a low recall, where the correct entities are not
extracted and the number of false negatives grows. However, this would also lead to low precision, as
both under and over-classification of entities would also lead to the extraction of incomplete entities
which in turn leads to more false positives, and our models achieved good and even great precision. To
achieve great precision and bad recall, our models would have to extract a low number of entities as to
not increase the number of false positives, and they do as demonstrated by Tables 6 and 7.</p>
        <p>Then, through further error analysis, we realized that this might be happening due to our data parsing
approach. Because of how we parsed the datasets (same approach as PlanTL-GOB-ES’s when parsing
6
10</p>
        <p>2</p>
        <p>Epoch
3_bsc-bio-ehr-es_drugtemist_4
6</p>
        <p>Epoch
4_bsc-bio-ehr-es_drugtemist_1
Epoch
6
8
8</p>
        <p>Training Loss
Validation Loss
Best Model Epoch</p>
        <p>
          10
CANTEMIST 3 or PharmaCoNER 4 for NER in order to fine-tune the same model that we used [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]), the
model was not only trained, but also locally evaluated on individual sentences, not entire documents.
This approach difers from the way that the oficial evaluation library evaluates runs, as the models
3https://huggingface.co/PlanTL-GOB-ES/bsc-bio-ehr-es-cantemist
4https://huggingface.co/PlanTL-GOB-ES/bsc-bio-ehr-es-pharmaconer
are expected to predict on the full content of each test document, making the spans for each extracted
entity relative to the whole document, not just to individual sentences.
        </p>
        <p>This conflict of approaches might explain the discrepancy between the quality of the results,
particularly the recall, obtained on our validation set with our evaluation method, and the quality of the
results obtained with the oficial evaluation library, both on our validation set and on the test set: as
just mentioned, the examples that were fed to the model during training were individual sentences,
meaning that they were overall much shorter (average length of 21 tokens, 124 chars) than the ones
fed to the model when evaluating it on the test set (entire documents, average length of 1013 tokens,
5754 chars). Thus, the small size of the training examples might have induced the model to only extract
entities up until a certain span when fed longer examples (i.e. documents). Furthermore, and possibly
more relevant, the maximum input sequence length of this model is 512 tokens, which means that when
obtaining the test set predictions to submit, with the entire documents as input, the tokens that went
over this limit were not classified by our models.</p>
        <p>These two reasons explain the low number of entities extracted, as the majority of the misses occur
near the end of the documents. On the other hand, the results using our evaluation method, during
the model’s training, did not show this low recall because the model was also evaluated on individual
sentences. The information presented in Tables 8 and 9 supports our claims, as the average start span
of (true and false) positive, i.e retrieved entities, is much lower than the average span of false negatives,
i.e. relevant but not retrieved entities.</p>
        <p>We believe that this is the reason why our approach showed consistent results across the board on
our internal validation, but failed to replicate them on the oficial testing: basically the input length
of the model when evaluating it on a document basis being smaller than many documents, with the
additional nuance that the diference in length between the training and the test examples (the first
much shorter than the second) may have induced the model to only predict entities until a certain span.
We plan to re-classify the dev and tests sets using sentences instead of the full document in order to
verify if the recall would improve and become closer to our internal evaluation.</p>
        <p>Nevertheless, and focusing on the objective of experiment itself, we can observe that training with
the combined dataset did not show any substantial improvements regarding the F1-Score. In fact, both
tasks showed a barely significant disadvantage on the validation set: for task 1, a drop of 0.38pp on our
evaluation method and of 0.85pp using the oficial evaluation library, and on task 2 a drop of 0.92pp
for our evaluation method and of 0.76pp using the oficial library. On the test set, while the precision
for all runs is quite good, even achieving the best precision score for any team on task 2 with run 3
(0.9242), the recall achieved by all runs is very low, which results in poor F1-scores, more specifically
20.64pp and 16.62pp below the mean for tasks 1 and 2 respectively, for our best run from each task.
Furthermore, when checking for the main objective of the experiments, we observe again no substantial
advantage for training with the combined dataset, only increasing the F1-Score by 0.95pp for task 1 and
0.80pp for task 2.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>We employ the state-of-the-art transformer model bsc-bio-ehr-es and explore the hypothesis of if
training with more entity types within biomedical domain would be beneficial for extracting a specific
type, with the results indicating it does not make a significant diference. Furthermore, with the scores
obtained on the test set, we show that this model can achieve results with high precision and with
minimal fine-tuning. However, the model’s poor recall on the test set is noteworthy and likely caused
by the the way in which the data is parsed for training and evaluation.</p>
      <p>In the future, we intend to perform hyperparameter optimization on this model, to try to improve
the scores that we obtained for the MultiCardioNER task (both on the DisTEMIST and DrugTEMIST
subtracks), and apply English and Italian biomedical pre-trained language models to the corresponding
subsets of the DrugTEMIST dataset.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgements</title>
      <p>This work is supported by NOVA LINCS ref.
(https://doi.org/10.54499/UIDB/04516/2020) and ref.
(https://doi.org/10.54499/UIDP/04516/2020) with the financial support of FCT.IP.</p>
      <p>UIDB/04516/2020
UIDP/04516/2020
cessing, Association for Computational Linguistics, Dublin, Ireland, 2022, pp. 193–199. URL:
https://aclanthology.org/2022.bionlp-1.19. doi:10.18653/v1/2022.bionlp-1.19.
[7] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov,
RoBERTa: A Robustly Optimized BERT Pretraining Approach, 2019. URL: http://arxiv.org/abs/1907.
11692, arXiv:1907.11692 [cs].
[8] A. Miranda-Escalada, E. Farré-Maduell, M. Krallinger, Named Entity Recognition, Concept
Normalization and Clinical Coding: Overview of the Cantemist Track for Cancer Text Mining in Spanish,
Corpus, Guidelines, Methods and Results, 2020. doi:10.5281/zenodo.3773228.
[9] A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, PharmaCoNER:
Pharmacological Substances, Compounds and proteins Named Entity Recognition track, in:
Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, Association for Computational
Linguistics, Hong Kong, China, 2019, pp. 1–10. URL: https://www.aclweb.org/anthology/D19-5701.
doi:10.18653/v1/D19-5701.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rodríguez-Miret</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rodríguez-Ortega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lenkowicz</surname>
          </string-name>
          , G. Ceroni,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kossof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          , G. Katsimpras, G. Paliouras,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Krallinger, Overview of MultiCardioNER task at BioASQ 2024 on Medical Speciality and Language Adaptation of Clinical NER Systems for Spanish, English and Italian</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . García Seco de Herrera (Eds.),
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Davydova</surname>
          </string-name>
          , E. Tutubalina, G. Paliouras,
          <source>Overview of BioASQ</source>
          <year>2024</year>
          :
          <article-title>The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Maria Di Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ),
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Intxaurrondo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <surname>SPACCC</surname>
          </string-name>
          ,
          <year>2018</year>
          . URL: https://zenodo.org/doi/10.5281/zenodo.1563762. doi:
          <volume>10</volume>
          .5281/ZENODO.1563762.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gascó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          , G. Katsimpras, G. Paliouras,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Krallinger, Overview of medprocner task on medical procedure detection and entity linking at bioasq 2023</article-title>
          , Working Notes of CLEF (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gasco-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rodríguez-Miret</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Krallinger, Overview of SympTEMIST at BioCreative VIII: corpus, guidelines and evaluation of systems for the detection and normalization of symptoms, signs and findings from text</article-title>
          ,
          <year>2023</year>
          . URL: https://zenodo.org/doi/10. 5281/zenodo.10104547. doi:
          <volume>10</volume>
          .5281/ZENODO.10104547.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Llop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gutiérrez-Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Armengol-Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silveira-Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Valencia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gonzalez-Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Pretrained Biomedical Language Models for Clinical NLP in Spanish</article-title>
          ,
          <source>in: Proceedings of the 21st Workshop on Biomedical Language</source>
          Pro-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>