<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>S. Piron);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Named Entity Recognition with GLiNER and Relation Extraction with LLMs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Samuel Piron</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio Maria Di Nunzio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <addr-line>Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Biomedical Information Extraction from Natural Language Processing (NLP) is one of the newest challenges driving innovation in the biomedical scientific field. In this work, we present our implementation pipeline for the GutBrain shared task covering both Named Entity Recognition and Relation Extraction. For Subtask 6.1 (NER), we fine-tuned the GLiNER framework on expert-annotated GutBrain datasets, achieving robust entity recognition between the predefined categories. For the RE Subtasks (6.2.1-6.2.3), we injected entity markers into text and employed fine-tuned BiomedBERT and pubmed-bert classifiers to predict relations between entities. By exploring Precision-oriented, Recall-oriented, and balanced configurations, we identified the best setups for maximizing Precision, Recall, and F1 for each task. Finally, we show our results with scatter plots and discuss the trade-of each run ofers.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Information Extraction</kwd>
        <kwd>Named Entity Recognition</kwd>
        <kwd>GLiNER</kwd>
        <kwd>Relation Extraction</kwd>
        <kwd>Large Language Models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the biomedical field, large volumes of textual data are generated daily, such as electronic health
records and biomedical literature. Extracting and structuring this information in an eficient way is
crucial for improving healthcare quality, supporting clinical decision-making, and advancing medical
research [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Natural Language Processing (NLP) techniques have therefore become fundamental tools
in medical text mining and Information Extraction (IE).
      </p>
      <p>
        A key subtask of information extraction is Named Entity Recognition (NER), which involves
identifying and categorizing spans of text into predefined categories such as diseases, treatments, and
anatomical entities [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Based on NER, Relation Extraction (RE) identifies and extracts relationships between named entities
from the underlying content [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. RE is crucial for facilitating the extraction of information from large
datasets, particularly when the data is unstructured. It supports many downstream applications, such
as transforming unstructured corpora into knowledge graphs, question answering, and automated
document processing [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ]. Figure 1 illustrates an example of NER and RE annotation in biomedical
text.
      </p>
      <p>
        The CLEF 2025 conference is the 16th edition of the Conference and Labs of the Evaluation Forum
(CLEF), continuing the popular CLEF campaigns that have been running since 2000 and contributing to
the systematic evaluation of information access systems through experimentation with shared tasks1.
In particular, BioASQ 2025 Lab Task 6, known as GutBrainIE, promotes the development of Information
Extraction systems by extracting named entities and the relations between them2 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The challenge is divided into 4 Subtasks:</p>
      <p>dietary
supplement
human
Does ⏞Fibre⏟-fix provided to ⏞peo⏟ple, with i⏞rritable bow⏟el syndrome, who are consuming a low
target microbiome</p>
      <p>FODMAP diet, improve their gut health, ⏞gut micr⏟obiome sleep and mental health?
• Subtask 6.1, the participants are provided with PubMed abstracts about the gut-brain axis focusing
on the Parkinson’s disease and mental health. Their task is to classify specific entity mentions
into one of the 13 predefined categories: anatomical location, animal, biomedical technique,
bacteria, chemical, dietary supplement, disease disorder or finding (DDF), drug, food, gene,
human, microbiome, and statistical technique. An example is shown in Table 1.
• The next challenges concerns RE over the entities identified by the NER systems:
– Subtask 6.2.1 is the Binary Tag-Based RE, where participants are asked to identify which
entities are in relation within a document. An example is shown in Table 2.
– Subtask 6.2.2 is the Ternary Tag-Based RE: participants are required to identify the actual
entities involved in a relation and predict the type of relation. An example is shown in Table
3.
– Finally, Subtask 6.2.3, the Ternary Mention-Based RE, where participants are required to
identify the actual entities involved in a relation and predict the type of relation. An example
is shown in Table 4.</p>
      <p>
        For more detailed information about the GutBrainIE task and the proposed subtasks, please refer to the
overview paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Early approaches to NER and RE relied on hand-crafted rules, which were later replaced by probabilistic
models such as Hidden Markov Models (HMMs) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Conditional Random Fields (CRFs) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
Mikolov et al. [12] introduced distributed word representations, marking one of the first significant
advances in capturing word similarity, which remains foundational in modern NLP. Lample et al. [13]
proposed a BiLSTM-CRF architecture for NER, where a Bidirectional Long Short-Term Memory network
is inserted between the input words and the CRF output layer.
      </p>
      <p>RE first appeared prominently in SemEval-2010 Task 8 [ 14], which focused on the “Relation
Classification Subtask" and assigned a single label to a marked entity pair. The deep learning era began
with convolutional neural networks (CNNs), which enabled mapping entire sentences to relation labels
without manual feature engineering [15]. Subsequently, the advent of pre-trained language models
(PLMs) marked a major advancement in RE, as highlighted by Li et al. [16]. These fine-tuned models,
such as BERT and RoBERTa, trained on diverse datasets, have become standard for many modern NLP
tasks.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>NER models are widely used in data mining, textual analysis, and text processing. However, they often
lack flexibility, and training them can be a challenging task. To address these limitations, Zaratiana
et al. [17] proposed GLiNER, a recent and efective alternative to traditional NER models, which are
typically restricted to predefined entity types and rely on expensive Large Language Models (LLMs).
GLiNER is a NER model capable of recognizing a wide range of entity types using a Bidirectional
Transformer architecture.</p>
      <p>The model employs a Bidirectional Language Model (BiLM) and takes as input a set of entity type
prompts and a sentence or text, with each entity separated by a learned token [ENT]. The BiLM outputs
representations for each token. Entity embeddings are passed into a FeedForward Network, where
input word representations are passed into a span representation layer to compute embeddings for each
span. Finally, it computes a matching score between entity representations and span representations
(using dot product and sigmoid activation) [17]. Figure 2 shows the overall architecture.</p>
      <p>In our implementation to perform NER on the GutBrain datasets, we used NuNER Zero, a zero-shot
NER model3. NuNER Zero is based on the GLiNER architecture and expects input as a concatenation of
entity types and the target text. It was trained on the NuNER-v2.0 dataset4, which combines annotated
subsets of the Pile5 and C46 corpora. Annotations were generated using large language models following
the NuNER procedure, which employs GPT-3.5-turbo to label entity mentions in a large-scale English
corpus (C4) [18] with semantically meaningful concepts. The LLM was prompted to extract as many
relevant entities as possible from each sentence and assign them to one of approximately 200k unique
concepts (e.g., “wellness”). This annotation process resulted in over 4.3 million labeled entities, providing
high conceptual diversity but also showing class imbalance and ambiguities, which were addressed
through filtering and training procedures [19].</p>
      <p>Our implementation pipeline follows these steps:
• Data analysis: we begin by examining the composition of the GutBrain corpus (titles, abstracts,
and entity annotations), dropping any annotations whose location is not "title" or "abstract".
• Data loading: load the JSON files containing metadata and character-level entity labels.</p>
      <sec id="sec-3-1">
        <title>3https://huggingface.co/numind/NuNER_Zero 4https://huggingface.co/numind/NuNER-v2.0 5https://huggingface.co/datasets/EleutherAI/pile 6https://huggingface.co/datasets/allenai/c4</title>
        <p>• Tokenizer initialization: initialize the microsoft/deberta-v3-large tokenizer7, which
splits the text into tokens and provides ofset mappings for model input.
• Preprocessing: combine the title and abstract for each document; convert character-level entity
spans into token-level spans using the tokenizer’s ofset mappings; store the tokenized text and
corresponding entities.
• Hyperparameter configuration : define the number of training steps based on dataset size,
along with the maximum number of epochs, batch size, and maximum token length.
• Model setup: initialize the token classification model numind/NuNER_Zero based on GLiNER
with pre-trained weights, and configure sampling parameters to manage training complexity.
• Training loop: iterate over batches, perform forward passes, compute loss, and apply
backpropagation. Update model weights using the AdamW optimizer8.
• Evaluation: periodically evaluate model performance on dev.json, reporting Micro and Macro</p>
        <p>Precision, Recall, and F1 scores.</p>
        <p>To predict NER entities with our trained models, we structured our implementation pipeline in this
way: for each article, we ran the model separately on its title and abstract, requesting predictions for
all thirteen entity types (e.g., animal, anatomical location, DDF). The raw output consisted of a list of
spans with start and end ofsets, labels, and confidence scores. We then merged any adjacent spans
with the same label and converted the character-level ofsets into the required format. Finally, we saved
the predictions in a JSON file, including the (start_idx) and (end_idx) fields.</p>
        <p>For the other RE Subtasks of the GutBrainIE challenge, we employed two diferent
pretrained models: NeuML/pubmedbert-base-embeddings9, a fine-tuned model based on
sentencetransformers trained on a dataset of randomly sampled PubMed10 title–abstract pairs and similar titles;
and microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext11, a biomedical
7https://huggingface.co/microsoft/deberta-v3-large
8https://pytorch.org/docs/stable/generated/torch.optim.AdamW.html
9https://huggingface.co/NeuML/pubmedbert-base-embeddings
10https://pubmed.ncbi.nlm.nih.gov/
11https://huggingface.co/microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224
model pre-trained from scratch using PubMed abstracts and full-text articles from PubMed Central12
[20]. Our pipeline for these Subtasks is as follows:
• Data loading: we first load the JSON files containing the GutBrain datasets (platinum, gold, and
development collections) and extract article texts, entity annotations, and relation annotations.
We then initialize the microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fu
lltext tokenizer to prepare the input for the neural network, inserting special entity
markers—[E1]. . . [/E1] for subjects and [E2]. . . [/E2] for objects—into the raw text.
• Preprocessing: split the entities between the title and the abstract, and filter for relations
involving entities from both.
• Model configuration : extend the BERT embedding layer to recognize the newly introduced
entity markers. Instantiate the BiomedBERT model for sequence classification, where each input
sequence (text with entity markers) is processed to determine the presence of a binary relation.
• Training setup: configure the AdamW optimizer and learning rate scheduler, and define the
hyperparameters used to train the model, including number of epochs, batch size, learning rate,
and sequence length.
• Model training: in each training step, perform a forward pass, compute the loss, backpropagate
gradients, and update model weights via the optimizer and scheduler.
• Validation: after each epoch, evaluate model performance on the test split, reporting Micro and</p>
        <p>Macro Precision, Recall, and F1-score.</p>
        <p>To predict relations between entities in the test dataset, we implemented the following pipeline.
Starting from the named entities identified in Subtask 6.1, we generate all ordered pairs of distinct entities
(1, 2). We insert special markers ([E1], [E2]) into the raw text to highlight the subject and object,
respectively. The marked text is then tokenized into a fixed-length sequence and converted into BERT
input tensors. We feed the encoded sequence into a fine-tuned BertSequenceClassification
model, whose output logits corresponds to the relation labels defined in the label2id.json file.
Finally, we apply the softmax function, select the top-scoring label with a confidence above 0.7, and
save the predictions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setup</title>
      <sec id="sec-4-1">
        <title>The experimental setup for this project includes the following components:</title>
        <p>• The project source code is available in the ataupd2425-gainer repository on GitHub:
https://github.com/Vezzero/ataupd2425-gainer.
• The dataset collection was provided by the CLEF BioASQ 2025 organizers and is available at:
https://hereditary.dei.unipd.it/challenges/gutbrainie/2025/.
• The evaluation script used is evaluation.py, made available by the organizers via their oficial
GitHub repository:
https://github.com/MMartinelli-hub/GutBrainIE_2025_Baseline/blob/main/Eval/evaluate.py.</p>
        <p>
          Full details about provided data, baselines, and evaluation can be found in the overview paper [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
• Model training and prediction for both NER and RE tasks were performed using the following
hardware:
– Tesla T4 GPU on Google Colab: https://colab.research.google.com/
– Dual NVIDIA T4 GPUs on Kaggle: https://www.kaggle.com/
– 8× NVIDIA A40 GPUs on DEI (Dipartimento di Ingegneria dell’Informazione) cluster:
https://docs.dei.unipd.it/
        </p>
        <p>A README file is provided in the GitHub repository with complete reproducibility instructions.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In this section, we report and describe the results obtained from the evaluation on the test data.</p>
      <p>The runs submitted have been evaluated by the organizers on a held-out test set of 40 expert-annotated
articles. These annotations are used as ground truth to compute the performance metrics.</p>
      <sec id="sec-5-1">
        <title>5.1. Named Entity Recognition Results</title>
        <p>For the NER Subtask, we trained three variants of the numind/NuNER_Zero model: PironA, PironD,
and PironS. All models share the same architecture but difer in the training dataset, number of epochs,
and training steps as reported in Table 5.</p>
        <p>The NER test results are reported in Table 6. The PironA model with the ma run achieves the highest
Micro-F1 score with high Micro-Precision and Micro-Recall. In particular, its Micro-Precision indicates
that almost all entities are labeled correctly, resulting in the fewest overall errors. PironD with the md
run obtains the highest Macro-Recall with lower Precision, suggesting it captures more true instances
while generating more false positives. In contrast, PironS with the last run, ms, performs worst on both
Micro and Macro scores, likely a consequence of the noisier annotations in the silver dataset.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Relation Extraction Results</title>
        <p>For the RE Subtasks, we trained four diferent models, finetuning the BiomedNLP-BiomedBERT-base
-uncased-abstract-fulltex13 and NeuML/pubmedbert-base-embeddings14. BiomedBERT
was fine-tuned under three diferent configurations, each varying only in training data and
hyperparameters. Table 7 reports the models setup.</p>
        <p>As explained in Section 3, we base our RE models on the entity spans produced by our three NER
variants (ma, md, and ms). For each test document, we first apply one of these variants to identify and
save entities and then feed them into the following RE model to generate the final relation labels.
13https://huggingface.co/microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext
14https://huggingface.co/NeuML/pubmedbert-base-embeddings
5.2.1. Binary Tag-Based RE Results
Table 8 reports both macro- and micro-averaged metrics for all submitted runs. The bp run, trained
on the combined platinum, gold, and dev relation sets, with the BinaryTGDNeuml classifier and the
NER predictions of the ma run, achieves the highest Macro F1 and Micro F1. We attribute this strong
performance to its underlying BiomedBERT encoder: initially pre-trained on millions of PubMed
abstracts and then fine-tuned in a sentence-transformers framework, it produces 768 dimensional
embeddings that encode biomedical semantics. These more accurate representations help the model
distinguish valid relations and invalid ones better and assign higher confidence scores to its predictions.
The bd runs group, based on the NER prediction in the md run, shows a high Macro and Micro Recall
but low Macro and Micro Precision, suggesting it is a Recall-oriented model. The ba group shows
the opposite: higher Macro and Micro Precision but lower Macro and Micro Recall, suggesting it is a
Precision-oriented model. The bs runs group had the lower Micro and Macro Precision (even if the
Macro and Micro Recall are similar to the bd group). This shows again that the silver dataset’s noise
negatively influences the Precision value.</p>
        <p>In the scatter plot, Figure 3, each point represents a run, with Micro Precision on the x-axis and Micro
Recall on the y-axis. The Recall-oriented bd/bs runs group cluster in the top-left with high Recall and
low Precision. The Precision-oriented ba group is located in the mid-left. The balanced ones, bp group,
appears toward the bottom-right, corresponding to their highest Micro F1.
5.2.2. Ternary Tag-Based RE Results
For the Subtask 6.2.2, we leveraged our PironBinary models, updating the inference pipeline to predict
and return relations between entities. Table 9 reports the performance of each run.</p>
        <p>The td run, obtained using the PironBinaryTGD model and the ma run, achieves the highest Micro
Precision, resulting in the highest Micro F1. This suggests that the run correctly predicted nearly
all relations between entities. In contrast, the ts run shows the best score in Macro Precision, as
both rare and common relations contribute equally to the final average, leading to the best Macro
F1. This comparison highlights that PironBinaryTGD provides the highest overall accuracy in the td
configuration, while PironBinaryTGS, in the ts setup with fewer positive examples, excels in achieving
balanced performance across all relation classes.
5.2.3. Ternary Mention Based RE Results
Finally, Table 10 reports the performance of all 6.2.3 runs. The ts configuration, using the
PironBinaryTGS model with the entities of the ma run, achieves the highest Micro and Macro F1, making it the
best run overall.</p>
        <p>To explore how each run balances false positives and false negatives, Figure 4 plots Micro Precision
against Micro Recall. In this scatter plot, the tma runs cluster in the upper left, with high Precision
but low Recall, indicating they efectively avoid false positives but miss many true relations. The tmd
group shows a well-balanced trade-of between Precision and Recall. The tms runs overlap with the
tmd seeds, but the aggregated tms shifts into the top right quadrant, achieving the best combination of
Micro Precision and Micro Recall—and consequently, the highest Micro F1.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and Future Perspectives</title>
      <p>In this paper, we presented the implementation pipelines for biomedical information extraction in the
GutBrain corpus, covering both NER and RE. For Subtask 6.1, we fine-tuned the NuNer_Zero model under
three hyperparameter configurations, achieving an almost 80% Micro F1 score. The chosen datasets used
to train the model impacted the run performance, enabling us to see that the silver collections contained
noisy annotations and leading to the worst results. In Subtasks 6.2.1, 6.2.2, and 6.2.3, we carried out
RE as a marker-based sequence classification problem, using PironBinary models trained using the</p>
      <p>BiomedBERT base model. We were able to produce Precision-oriented, Recall-oriented, and balanced
variants, each performing the best in diferent Subtasks, but without a dominating configuration.</p>
      <p>Looking ahead, we plan to investigate and test new LLMs’ models and more noise-robust training
techniques, such as confidence-aware loss functions, to improve both NER and RE performance. In
addition, we aim to integrate a semantic perspective grounded in linguistic analysis to enrich the
linguistic and conceptual interpretation of extracted terms and relations. Specifically, we would like to
apply semic analysis, which decomposes terms into minimal semantic units, as a structured approach
to uncovering the internal organization of meaning in medical terminology [21, 22]. Incorporating
this technique may enhance our ability to align terminological outputs with underlying conceptual
structures, improving not only model interpretability but also the precision of information retrieval in
domain-specific biomedical contexts.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work is partially supported by the HEREDITARY Project, as a part of the European Union’s Horizon
Europe research and innovation programme under grant agreement No GA 101137074.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <sec id="sec-8-1">
        <title>The authors have not employed any Generative AI tools.</title>
        <p>[12] T. Mikolov, I. Sutskever, K. Chen, G. Corrado, J. Dean, Distributed representations of
words and phrases and their compositionality, 2013. URL: https://arxiv.org/abs/1310.4546.
arXiv:1310.4546.
[13] G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, C. Dyer, Neural architectures for named
entity recognition, 2016. URL: https://arxiv.org/abs/1603.01360. arXiv:1603.01360.
[14] I. Hendrickx, S. N. Kim, Z. Kozareva, P. Nakov, D. Ó Séaghdha, S. Padó, M. Pennacchiotti, L. Romano,
S. Szpakowicz, SemEval-2010 task 8: Multi-way classification of semantic relations between pairs
of nominals, in: K. Erk, C. Strapparava (Eds.), Proceedings of the 5th International Workshop
on Semantic Evaluation, Association for Computational Linguistics, Uppsala, Sweden, 2010, pp.
33–38. URL: https://aclanthology.org/S10-1006/.
[15] D. Zeng, K. Liu, S. Lai, G. Zhou, J. Zhao, Relation classification via convolutional deep neural
network, in: J. Tsujii, J. Hajic (Eds.), Proceedings of COLING 2014, the 25th International Conference
on Computational Linguistics: Technical Papers, Dublin City University and Association for
Computational Linguistics, Dublin, Ireland, 2014, pp. 2335–2344. URL: https://aclanthology.org/
C14-1220/.
[16] J. Li, T. Tang, W. X. Zhao, J.-Y. Nie, J.-R. Wen, Pretrained language models for text generation: A
survey, 2022. URL: https://arxiv.org/abs/2201.05273. arXiv:2201.05273.
[17] U. Zaratiana, N. Tomeh, P. Holat, T. Charnois, Gliner: Generalist model for named entity recognition
using bidirectional transformer, 2023. URL: https://arxiv.org/abs/2311.08526. arXiv:2311.08526.
[18] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring
the limits of transfer learning with a unified text-to-text transformer, Journal of Machine Learning
Research 21 (2020) 1–67. URL: http://jmlr.org/papers/v21/20-074.html.
[19] S. Bogdanov, A. Constantin, T. Bernard, B. Crabbé, E. Bernard, Nuner: Entity recognition encoder
pre-training via llm-annotated data, 2024. arXiv:2402.15343.
[20] Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao, H. Poon,
Domain-specific language model pretraining for biomedical natural language processing, 2020.
arXiv:arXiv:2007.15779.
[21] V. Bonato, G. M. Di Nunzio, F. Vezzani, A Novel Approach to Semic Analysis: Extraction of
Atoms of Meaning to Study Polysemy and Polyreferentiality, Languages 9 (2024) 121. URL: https:
//www.mdpi.com/2226-471X/9/4/121. doi:10.3390/languages9040121, number: 4 Publisher:
Multidisciplinary Digital Publishing Institute.
[22] V. Bonato, G. M. Di Nunzio, F. Vezzani, Preliminary Considerations on a Systematic Approach to
Semic Analysis: The Case Study of Medical Terminology, Umanistica Digitale (2021) 211–234. URL:
https://umanisticadigitale.unibo.it/article/view/12621. doi:10.6092/issn.2532-8816/12621,
number: 10.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dunn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dagdelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Rosen</surname>
          </string-name>
          , G. Ceder,
          <string-name>
            <given-names>K.</given-names>
            <surname>Persson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>Structured information extraction from complex scientific text with fine-tuned large language models</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2212.05238. arXiv:
          <volume>2212</volume>
          .
          <fpage>05238</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>K. M. S. Islam</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          <string-name>
            <surname>Nipu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Madiraju</surname>
          </string-name>
          ,
          <article-title>Llm-based prompt ensemble for reliable medical entity recognition from ehrs</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2505.08704. arXiv:
          <volume>2505</volume>
          .
          <fpage>08704</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , H. Cheng, W. Lam,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey on relation extraction: Recent advances</article-title>
          and new frontiers,
          <year>2024</year>
          . URL: https://arxiv.org/ abs/2306.
          <year>02051</year>
          . arXiv:
          <fpage>2306</fpage>
          .
          <year>02051</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Nayak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Poria</surname>
          </string-name>
          ,
          <article-title>Deep neural approaches to relation triplets extraction: a comprehensive survey</article-title>
          ,
          <source>Cognitive Computation 13</source>
          (
          <year>2021</year>
          )
          <fpage>1215</fpage>
          -
          <lpage>1232</lpage>
          . URL: http://dx.doi.org/10. 1007/s12559-021-09917-7. doi:
          <volume>10</volume>
          .1007/s12559-021-09917-7.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Si</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Hybrid transformer with multi-level fusion for multimodal knowledge graph completion</article-title>
          ,
          <source>in: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2022</year>
          , pp.
          <fpage>904</fpage>
          -
          <lpage>915</lpage>
          . URL: http://dx.doi.org/10.1145/3477495.3531992. doi:
          <volume>10</volume>
          .1145/ 3477495.3531992.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Knowledge base question answering via encoding of complex query graphs</article-title>
          , in: E.
          <string-name>
            <surname>Rilof</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Chiang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hockenmaier</surname>
          </string-name>
          , J. Tsujii (Eds.),
          <source>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>2185</fpage>
          -
          <lpage>2194</lpage>
          . URL: https://aclanthology.org/D18-1242/. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D18</fpage>
          -1242.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Diaz-Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A. D.</given-names>
            <surname>Lopez</surname>
          </string-name>
          ,
          <article-title>A survey on cutting-edge relation extraction techniques based on language models</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2411.18157. arXiv:
          <volume>2411</volume>
          .
          <fpage>18157</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bonato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Irrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Menotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Vezzani</surname>
          </string-name>
          , Overview of GutBrainIE@CLEF 2025:
          <article-title>Gut-Brain Interplay Information Extraction</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , D. Spina (Eds.),
          <source>CLEF 2025 Working Notes</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rodríguez-Ortega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rodriguez-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sakhovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tutubalina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dimitriadis</surname>
          </string-name>
          , G. Tsoumakas,
          <string-name>
            <given-names>G.</given-names>
            <surname>Giannakoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bekiaridou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Samaras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinelli</surname>
          </string-name>
          , G. Silvello, G. Paliouras,
          <source>Overview of BioASQ</source>
          <year>2025</year>
          :
          <article-title>The thirteenth BioASQ challenge on large-scale biomedical semantic indexing and question answering</article-title>
          ,
          <source>volume TBA of Lecture Notes in Computer Science</source>
          , Springer,
          <year>2025</year>
          , p.
          <source>TBA.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gader</surname>
          </string-name>
          ,
          <article-title>Generalized hidden markov models. i. theoretical frameworks</article-title>
          ,
          <source>Fuzzy Systems, IEEE Transactions on 8</source>
          (
          <year>2000</year>
          )
          <fpage>67</fpage>
          -
          <lpage>81</lpage>
          . doi:
          <volume>10</volume>
          .1109/91.824772.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <article-title>An introduction to conditional random fields</article-title>
          ,
          <year>2010</year>
          . URL: https://arxiv. org/abs/1011.4088. arXiv:
          <volume>1011</volume>
          .
          <fpage>4088</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>