<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kunal Suri</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prakhar Mishra</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saumajit Saha</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Atul Singh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Optum</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Finetuning Large Language Models helps improve the results for domain-specific use cases. End-to-end ifnetuning of large language models is time and resource intensive and has high storage requirements to store the finetuned version of the large language model. Parameter Eficient Fine Tuning (PEFT) methods address the time and resource challenges by keeping the large language model as a fixed base and add additional layers, which the PEFT methods finetune. This paper demonstrates the evaluation results for one such PEFT method Low Rank Adaptation (LoRA), for Clinical Dialogue Summarization. The evaluation results show that LoRA works at par with end-to-end finetuning for a large language model. The paper presents the evaluations done for solving both the Subtask A and B from ImageCLEFmedical It is important to record conversations between medical personnel and patients for compliance, training, and evaluation purposes. To that end, summaries of such conversations serve as valuable tools for medical personnel and patients to refer back to and comprehend their prior interactions. Therefore, a concise summary must be produced to facilitate the next medical consultation and provide a source for future reference. Currently, such summaries are created manually; this summarization process is costly and labour-intensive. AI-based summarization techniques can help here by reducing the time and cost associated with manual summarization and facilitating the generation of more accurate representations of doctor-patient conversations by human scribes in less time. Sequence-to-Sequence (Seq2Seq) Architectures [1] have been at the forefront of creating summaries. Transformers [2] further improved the performance of this architecture. Over time, we have seen that the performance of these models have improved significantly 1 but it comes at the cost of increased model size which made it very dificult to fit such models on consumer grade hardware such as K80 or T4. Recently a couple of techniques such as</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dialogue Summarization</kwd>
        <kwd>Parameter Eficient Fine Tuning</kwd>
        <kwd>Clinical Dialogue Summarization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        LoRA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Prefix Tuning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], P-Tuning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Prompt Tuning [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] have been introduced which are
collectively referred to as Parameter Eficient Fine Tuning (PEFT) techniques. These techniques
are used for eficiently adapting pre-trained language models (PLMs) to various downstream
applications without fine-tuning all the model’s parameters. PEFT methods only trains a small
number of (extra) model parameters, significantly decreasing computational and storage costs
because fine-tuning large-scale PLMs is prohibitively costly. For this paper, we use the PEFT
implementation from Huggingface 2.
      </p>
      <p>
        This paper presents the experimental results of our explorations with LoRA on Clinical
Dialogs to accomplish both Subtask A and B of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] Shared Tasks from [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The solution of
SubTask B presented in this paper was ranked first among all the submissions for SubTask B.
The paper uses LoRA based models for both assigning conversations to a pre-defined set of
clinical notes sections and summarization of conversations. Through this work, the paper also
compares the performance of fine-tuned Transformer based models with LoRA based models for
classification and summarization tasks. In addition to this comparison, we also evaluate impact
of ensembling outputs from multiple Seq2Seq models using [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Our simulations show that
LoRA works as well as finetuning of Transformer-based models. This is very important because
it shows that we can get the equivalent performance as we get after fine tuning Transformer
models while using only a fraction of parameters which means that such models could be fine
tuned on consumer grade hardware such as K80 and T4.
      </p>
      <p>This paper is organized as follows. Section 3 presents a brief overview of SubTask A and
B - including available labeled data and evaluation metrics. Then the paper describes current
state-of-the-art for dialog classification and summarization in Section 2 that this paper builds
upon. This is followed by the description of the approach used to solve SubTask A in Section
4 and SubTask B in Section 5. Then the results of our solutions for both of these subtasks
are presented. Finally, the paper ends with a conclusion of the work. The paper includes an
appendix containing exploratory data analysis and material that will help to better understand
the solution presented in the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Finetuning large language models enables better performance for domain-specific use cases.
In-context finetuning performs well in few-shot scenarios enabling the end users to provide
examples with the prompt to enable LLMs to learn for the use case at hand. This approach does
not scales as it restricts sending multiple examples with the prompt. End-to-end finetuning of
LLMs is resource and time intensive and has the additional drawback of storing and managing
multiple copies of large-size models.</p>
      <p>
        Parameter Eficient Fine Tuning (PEFT) Methods attempt to solve the problems mentioned
above by finetuning a smaller number of existing or newly introduced parameters of the large
language model while keeping the rest of the parameters frozen. In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Lilian et al. divide
PEFT methods into the following four categories: additive, selective, reparameterization-based,
and hybrid methods. Additive methods such as adapters [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] introduce and train only a new
set of parameters or layers. Selective methods finetune only a few top layers of the network.
2https://huggingface.co/docs/peft/index
Reparametrization-based methods use a low-dimensional representation of the network to
reduce the number of parameters to be trained during finetuning. This paper evaluates
LowRank Adaptation (LoRA) a prominent example of this category of methods.
      </p>
      <p>
        Parameter Eficient Fine Tuning (PEFT) methods reduce the need to host a large-sized model
for each use case. They enable users to use a frozen base model with a small layer of model
weights that vary with the use case. In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the authors compare the performance of four
diferent PEFT techniques for scenarios where low, medium and high counts of samples are
available for fine-tuning. The evaluation results show that LoRA gives near-best performance
when low to medium data samples are available for summarization tasks. In another similar
related study in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the evaluations demonstrate that the best summarization for radiology
reports is achieved using a model pre-trained on the clinical text and then fine-tuned using
LoRA. In this paper, the authors have used LoRA and ensembling for summarization.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task Description</title>
      <p>
        This Section provides a high-level overview of the MEDIQA-Sum 2023 Task (including both
SubTask A and B) from ImageCLEFmed MEDIQA[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The Section starts with a description of
diferent SubTask goals followed by basic counts of available labeled data. The metric used to
evaluate this task is arithmetic mean of ROUGE-1 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], Bertscore F1 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and BLEURT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>3.1. Task Definition</title>
        <p>Given a short conversation between a Doctor and a patient or another Doctor (Dialogue), the
goal of SubTask A is to create a system that automatically predicts the Section to which the
conversation belongs to which is denoted by Section Header. There are twenty Sections Headers in
this dataset. Some examples of Section Headers are FAM/SOCHX, GENHX, PASTMEDICALHX,
CC. All of these Section Headers and their descriptions (Section Description) can be found
in Table A2. The goal of SubTask B is to create a system that generates a summary which
matches the human generated summary (Section Text) as closely as possible while optimizing
the metric for evaluation.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Labeled Data</title>
        <p>In this paper we have used the labeled data provided by MEDIQA-Sum 2023 organizers for
training the models. A sample data point from the labeled data set for SubTask A and B can
be found in Table A1. The oficial data consists of a training and validation split. For SubTask
A and B, training data contains 1201 and validation data contains 180 &lt;dialogue, section-text,
section-header&gt; triplets.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. SubTask A Methodology</title>
      <p>Given a short conversation between a doctor and a patient, the goal of SubTask A is to predict
its Section Header. This Section starts with a description of the approach used to predict the
Section Header.</p>
      <sec id="sec-4-1">
        <title>SubTaskA - Train Split SubTaskA - Test Split Best Multi Class Classification</title>
      </sec>
      <sec id="sec-4-2">
        <title>Section Header</title>
      </sec>
      <sec id="sec-4-3">
        <title>Multi Class</title>
      </sec>
      <sec id="sec-4-4">
        <title>Classification</title>
        <p>(a) SubTask A -- Training</p>
        <p>
          We have achieved success using Bio-ClinicalBERT [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] for classification in the healthcare
domain. Hence we choose it as the backbone and initialize LoRA layer on top of it. We use
this architecture for classification of Dialogue to a Section Header in SubTask A. We limit the
number of input tokens to 300 tokens because that is the length of majority of dialogues, as
shown in Figure A1. We use a 3 Fold Cross Validation approach for modeling purposes. This
is to ensure that we capture all information in the data. For every fold, we split its test part
into validation and test. We do this so that we can use validation split to select best model
using Early Stopping and test split to calculate its performance. The hyper-parameters used for
training and performance for all folds can be found in Table A3. During inference, we pass a
given Dialogue through all three models, take an average of the logits for all the classes and
output the class with the highest logit score.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. SubTask B Methodology</title>
      <p>Given a short conversation between a doctor and a patient, the goal of SubTask B is to summarize
it while ensuring that the generated summary is as fluent and as close to Section Text as possible.
This Section starts with a description of the methodology used to summarize the conversation.
For Dialogue Summarization, we have trained a LoRA layer on top of Seq2Seq models. This
Section also describes the processed labeled data used for training these models, followed by the
actual training steps. Then this Section looks at the steps used to generate the summary from
the decoder. Finally, we discuss the approach used for ensembling the outputs of these models.</p>
      <p>We train LoRA based Seq2Seq models using labeled data (Dialogue + Section Header, Section
Text) as (Input, Output) pair. Section Text is a part of the labeled data and is a human subject
matter expert-created summary of Dialogue. As a preprocessing step, we replace all new line
SubTask A - Train</p>
      <p>Split
Get Best Decoding Strategy
On BioBART-V2-Large
On FLAN-T5-Large</p>
      <p>Ensemble summaries
from all models</p>
      <p>Final Summary
(c) SubTask B - Get the best
decoding strategy and
ensemble
characters with whitespaces. The Dialogue is concatenated with the section description of its
Section Header by the SEP token of the Seq2Seq architecture. During training and inference,
we use the actual section description for the actual Section Header. No changes are made to
Section Text.</p>
      <p>
        We use a 3-fold cross validation scheme as described in 4 and train LoRA on two Seq2Seq
architectures - BioBart-V2-Large [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and Flan-T5-Large [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Here we need to select the number
of input tokens for encoder and decoder. For encoder, we have selected token length of 512
tokens and for decoder, we have selected token length of 400 tokens. All the hyper-parameters
used to train each of the above architecture can be found in Table A4. To select the best model,
we use early-stopping [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] based on Validation Negative Log Loss. Results on the test part of
each of these models can be found in Table A5. The distribution of tokens for Dialogue and
Section Text can be found in Figure A2 and Figure A3 respectively.
      </p>
      <p>
        To generate summaries that match the human generated summaries, we need a way to
control the text generated by the decoder component of a Seq2Seq model. This can be done by
using decoding strategies such as Beam Search [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], Top-k Sampling [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], Top-p Sampling [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
Contrastive Search [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] etc. In this module, we use Beam Search with TPESampler Algorithm
from Optuna3 to search for the optimal decoding strategy trying to maximize ROUGE-1,
ROUGE2, and BertScore rather than relying on manual tweaking of these metrics. We use TPESampler
here because it supports multivariate optimization and also it handles Float, Integer, and
Categorical values better than other algorithms present in Optuna4. We use Optuna here due to
ease of implementing Hyper-parameter optimization algorithms. We did not use BLEURT during
search because it is extremely time consuming. For this module, we use four hyper-parameters
3https://optuna.readthedocs.io/en/stable/reference/samplers/generated/optuna.samplers.TPESampler.html
4https://optuna.readthedocs.io/en/stable/reference/samplers/index.html
for Beam Search - Early Stopping, Number of Beams, No Repeat N-gram Size, Length Penalty.
The search space of each of these variables can be found in the Table 1.
      </p>
      <p>
        The results from the diferent models are ensembled using Generating Best Summary by
semantic similarity - a post-ensemble method [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to identify the summary which is closest to
all the generated summaries. The paper uses this output summary as the final summary for the
given Dialogue.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. SubTask A Results and Analysis</title>
      <p>This Section presents the results for SubTask A using the approach described in Section 4. We
have made only one submission for predicting Section Header whose Multi Class Accuracy was
73.5% on the test set given by the organizers, obtaining a rank of 8 among 23 submissions. In
this submission, we pass Dialogues through all three LoRA based Bio-ClinicalBERT models, take
an average of the logits for all the classes and output the class with the highest logit score. The
table containing our team’s standing can be found in the Tables A6. Standings of all the teams
have been calculated using multi class accuracy. We compared performance of Bio-ClinicalBERT
when it is fine-tuned end-to-end and when it is used as a backbone for LoRA. We observe that
Bio-ClinicalBERT with LoRA score 73.3% on validation data whereas end-to-end fine-tuned
Bio-ClinicalBERT score 72% on the same validation data.</p>
    </sec>
    <sec id="sec-7">
      <title>7. SubTask B Results</title>
      <p>This Section presents the results for SubTask B using the approach described in Section 5. We
have made three submissions (mentioned as runs in the result tables) for generating summaries
from Dialogues. For the summarization task, we have submitted results from three runs. In run
1 and run 2, we train LoRA on BioBart-V2-Large and Flan-T5-Large respectively while run 3
presents the results of ensembling summaries from both of these models. The details for each
run are as follows:
1. Run 1 - We generate summary from BioBart-V2-Large model trained on each fold and
ensemble output of all the models using 5
2. Run 2 - We generate summary from Flan-T5-Large model trained on each fold and
ensemble output of all the models using Generating Best Summary by semantic similarity.
3. Run 3 - We generate summary from BioBart-V2-Large and Flan-T5-Large model trained
on each fold and ensemble output of all the models using Generating Best Summary by
semantic similarity.</p>
      <p>The table containing our team’s standing can be found in Table A7. Standings of all the teams
have been calculated by calculating arithmetic mean of Rouge-1, Bertscore, BLEURT for the
Dialogue summary.</p>
      <p>The experiments show that Run3 performs the best scoring rank 1 out of 13 submissions.
This is also intuitive since it contains summaries from 3 models of BioBART-V2-Large and 3
models of Flan-T5-Large. Run2 scored 5th rank and Run1 scored 6th rank. This is an interesting
observation since Flan-T5-Large is an enhanced version of T5 that has been finetuned in a
mixture of tasks whereas BioBart-V2-Large has been trained solely on medical corpus so ideally
Run1 should have scored better than Run2 but it seems that bigger models work better than
domain specific models although this hypothesis needs to be validated.</p>
      <sec id="sec-7-1">
        <title>7.1. Analysis of diferent Transformer Architectures on SubTask B</title>
        <p>We compare performance of BioBart-V2-Large and Flan-T5-Large when they are fine-tuned
end-to-end and they are treated as backbone for LoRA. We observe that the models trained with
LoRA perform better than the models which were fine-tuned end-to-end. The performance
was evaluated by calculating arithmetic mean of ROUGE-1, ROUGE-2, and BertScore-F1. We
do not use BLEURT here as it is extremely time consuming and based on our observations,
ROUGE-2 and BLEURT have a very strong correlation. The average score across all folds for
each architecture can be found in the Table A5.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>The paper presents the solution and the results for SubTask A and B of ImageCLEFmed
MEDIQASum task. The solution uses LoRA to finetune Transformer based models to classify and
summarise Clinical Dialogues, and our simulation results show that the performance of Transformer
based models finetuned using LoRA is equivalent to the performance of Transformer based
models finetuned using resource and time-intensive end-to-end finetuning. The success of
Transformer based model finetunes using LoRA implies organizations can easily finetune and
deploy domain-based models.</p>
      <p>The authors observe that metrics such as ROUGE are inefective for evaluating the
performance of models like OpenAI GPT3 as they focus on syntactic similarity. Metrics such
as Bertscore and BLEURT seem more suitable for such models since they focus on semantic
similarity. Finally, the paper also evaluates two diferent ensemble techniques, and the
results demonstrate that the Post Ensemble technique performs the best while giving minimum
hallucinations.</p>
    </sec>
    <sec id="sec-9">
      <title>A. Appendix</title>
      <sec id="sec-9-1">
        <title>A.1. Data Exploration and Explanation</title>
        <p>This section discusses data exploration and explanation so that audience can understand why
we made the decisions that we made. A sample data point from dataset for SubTask A and B
can be seen in Table A1.</p>
        <p>The description of each of the Section Headers present in the data can be found in Table A2
The Class distribution of Section Headers for SubTask A is give by Figure A1
The Dialogue Token Distribution for SubTask A and B is give by Figure A2
The Clinical Note Token Distribution for SubTask B is give by Figure A3</p>
        <p>The hyper-parameters and performance metrics for Predicting Section Header i.e SubTask A
can be found in the Table A3.</p>
        <p>The hyperparameters used to fine tune Seq2Seq Models and LoRA i.e. SubTask B can be
found in Table A4. Each of these models were trained on 150 epochs, Gradient Accumulation of
16, Learning rate of 1e-3, AdamW optimizer, and Linear Learning Scheduler.</p>
        <p>The performance of diferent Seq2Seq Models using LoRA and Fine-tuning can be found in
Table A5</p>
        <p>Section Header Section Header Description</p>
        <p>FAM/SOCHX FAMILY HISTORY/SOCIAL HISTORY</p>
        <p>GENHX HISTORY OF PRESENT ILLNESS
PASTMEDICALHX PAST MEDICAL HISTORY</p>
        <p>CC CHIEF COMPLAINT
PASTSURGICAL PAST SURGICAL HISTORY</p>
        <p>ALLERGY ALLERGY</p>
        <p>ROS REVIEW OF SYSTEMS
MEDICATIONS MEDICATIONS
ASSESSMENT ASSESSMENT</p>
        <p>EXAM EXAM
DIAGNOSIS DIAGNOSIS
DISPOSITION DISPOSITION</p>
        <p>PLAN PLAN</p>
        <p>EDCOURSE EMERGENCY DEPARTMENT COURSE
IMMUNIZATIONS IMMUNIZATIONS</p>
        <p>IMAGING IMAGING</p>
        <p>GYNHX GYNECOLOGIC HISTORY</p>
        <p>PROCEDURES PROCEDURES
OTHER_HISTORY OTHER_HISTORY</p>
        <p>LABS LABS
Table A4
SubTask B - Hyperparameter Tuning for Diferent Architectures. Base Arch: Base Architecture, BS:
Batch Size, LR : Learning Rate, LoRA-A: LoRA-Alpha, LoRA-D: LoRA-Dropout, MaxSL : Maximum Source
Length, MaxTL : Maximum Target Length, MinTL : Minimum Target Length</p>
        <p>Base Arch</p>
        <p>Flan-T5-Large
Biobart-V2-Large</p>
        <p>BS
1
1</p>
        <p>LR
1e-3
1e-3</p>
        <p>LoRA-R
8
8</p>
        <p>LoRA-A
32
32</p>
        <p>LoRA-D
1e-3
1e-3</p>
        <p>MaxSL
512
512</p>
        <p>MaxTL
400
400</p>
        <p>MinTL
8
8
Section_Header Count</p>
        <p>300
t
u 200
n
o
C
100
0
/FSAXCHOM EXNHG ILPTESAACHDM CC ILPTSSAARCUG SRO LLYEARG IITESACNDOM TEESSSSANM EAXM IISSANDGO IIIPTSSNDOO LPAN EESRCUDO IIITAZUNNOMM IIANGGM YXNHG PEESRRCUDO I_TTESRHHOO LSAB</p>
        <p>X S YR</p>
        <p>section_header
FigureA1:ClassdistributionofSectionHeaders</p>
        <p>Token Length distribution for Dialogue</p>
        <p>Token Length distribution for Section Text
300
250
s 200
D
I
f
o
re 150
b
m
u
N 100
50
0
0
200</p>
      </sec>
      <sec id="sec-9-2">
        <title>A.2. Standing of our team</title>
        <p>Our standings (in bold) for SubTask A - Section Header Classification is in Table A6. We omitted
several teams from these standings and represent them by Ellipsis (...). This is done only to
conserve space.</p>
        <p>Table A6
SubTask A - Section Header Classification Standings</p>
        <p>Team Run
Cadence run1
...</p>
        <p>SuryaKiran
...</p>
        <p>SSNSheerinKavitha run1</p>
        <p>Our standings (in bold) for SubTask B - Summarization is in Table A7
Table A7
SubTask B - Section Text Summarization Standings</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Sequence to sequence learning with neural networks</article-title>
          , in: Z.
          <string-name>
            <surname>Ghahramani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cortes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Lawrence</surname>
          </string-name>
          , K. Weinberger (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>27</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2014</year>
          . URL: https://proceedings.neurips.cc/paper_files/paper/2014/file/ a14ac55a4f27472c5d894ec1c3c743d2-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , L. u. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          , in: I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>30</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2017</year>
          . URL: https://proceedings.neurips.cc/ paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Allen-Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Lora:
          <article-title>Low-rank adaptation of large language models</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2106</volume>
          .
          <fpage>09685</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Prefix-tuning: Optimizing continuous prompts for generation</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2101</volume>
          .
          <fpage>00190</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          , Gpt understands, too,
          <year>2021</year>
          . arXiv:
          <volume>2103</volume>
          .
          <fpage>10385</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Constant</surname>
          </string-name>
          ,
          <article-title>The power of scale for parameter-eficient prompt tuning</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>3045</fpage>
          -
          <lpage>3059</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>243</volume>
          . doi:
          <volume>10</volume>
          . 18653/v1/
          <year>2021</year>
          .emnlp-main.
          <volume>243</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Snider</surname>
          </string-name>
          , G. Adams,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Yetisgen, Overview of the mediqa-sum task at imageclef 2023: Summarization and classification of doctor-patient conversations</article-title>
          ,
          <source>in: CLEF 2023 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-M. Drăgulinescu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Yim</surname>
            ,
            <given-names>A. Ben</given-names>
          </string-name>
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Snider</surname>
            , G. Adams,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Yetisgen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bloch</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Brüngel</surname>
            , A. IdrissiYaghir, H. Schäfer,
            <given-names>S. A.</given-names>
          </string-name>
          <string-name>
            <surname>Hicks</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Thambawita</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Storås</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Papachrysos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Schöler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>A.-G.</given-names>
          </string-name>
          <string-name>
            <surname>Andrei</surname>
            ,
            <given-names>I. Filipovich</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Coman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stan</surname>
          </string-name>
          , G. Ioannidis,
          <string-name>
            <given-names>H.</given-names>
            <surname>Manguinhas</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
          </string-name>
          , Overview of ImageCLEF 2023:
          <article-title>Multimedia retrieval in medical, social media and recommender systems applications</article-title>
          , in: Experimental IR Meets Multilinguality, Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 14th International Conference of the CLEF Association (CLEF</source>
          <year>2023</year>
          ), Springer Lecture Notes in Computer Science LNCS, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kobayashi</surname>
          </string-name>
          ,
          <article-title>Frustratingly easy model ensemble for abstractive summarization</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>4165</fpage>
          -
          <lpage>4176</lpage>
          . URL: https://aclanthology.org/D18-1449. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D18</fpage>
          -1449.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lialin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Deshpande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rumshisky</surname>
          </string-name>
          ,
          <article-title>Scaling down to scale up: A guide to parametereficient fine-tuning</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>15647</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giurgiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jastrzebski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Morrone</surname>
          </string-name>
          , Q. de Laroussilhe,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gesmundo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Attariyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          ,
          <article-title>Parameter-eficient transfer learning for NLP</article-title>
          , CoRR abs/
          <year>1902</year>
          .00751 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1902</year>
          .00751. arXiv:
          <year>1902</year>
          .00751.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <article-title>Empirical analysis of the strengths and weaknesses of peft techniques for llms</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2304</volume>
          .
          <fpage>14999</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Veen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. V.</given-names>
            <surname>Uden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Attias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pareek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bluethgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polacin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chiu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-B. Delbrouck</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. M. Z. Chaves</surname>
            ,
            <given-names>C. P.</given-names>
          </string-name>
          <string-name>
            <surname>Langlotz</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          <string-name>
            <surname>Chaudhari</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Pauly</surname>
          </string-name>
          ,
          <source>Radadapt: Radiology report summarization via lightweight domain adaptation of large language models</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2305</volume>
          .
          <fpage>01146</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>C.-Y. Lin</surname>
            ,
            <given-names>ROUGE:</given-names>
          </string-name>
          <article-title>A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics</article-title>
          , Barcelona, Spain,
          <year>2004</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          . URL: https://aclanthology.org/W04-1013.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kishore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Artzi</surname>
          </string-name>
          , Bertscore:
          <article-title>Evaluating text generation with BERT</article-title>
          , CoRR abs/
          <year>1904</year>
          .09675 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1904</year>
          .09675. arXiv:
          <year>1904</year>
          .09675.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sellam</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Parikh</surname>
          </string-name>
          ,
          <article-title>Bleurt: Learning robust metrics for text generation</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2004</year>
          .04696.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Alsentzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Boag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. B. A. McDermott</surname>
          </string-name>
          ,
          <article-title>Publicly available clinical BERT embeddings</article-title>
          , CoRR abs/
          <year>1904</year>
          .03323 (
          <year>2019</year>
          ). URL: http: //arxiv.org/abs/
          <year>1904</year>
          .03323. arXiv:
          <year>1904</year>
          .03323.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Biobart: Pretraining and evaluation of a biomedical generative language model</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2204</volume>
          .
          <fpage>03905</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Longpre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fedus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Brahma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Webson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Suzgun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Castro-Ros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pellat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Valter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <surname>Scaling</surname>
          </string-name>
          instruction-finetuned
          <source>language models</source>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2210</volume>
          .
          <fpage>11416</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rosasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caponnetto</surname>
          </string-name>
          ,
          <article-title>On early stopping in gradient descent learning</article-title>
          ,
          <source>Constructive Approximation</source>
          <volume>26</volume>
          (
          <year>2007</year>
          )
          <fpage>289</fpage>
          -
          <lpage>315</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <article-title>Sequence transduction with recurrent neural networks</article-title>
          ,
          <source>CoRR abs/1211</source>
          .3711 (
          <year>2012</year>
          ). URL: http://arxiv.org/abs/1211.3711. arXiv:
          <volume>1211</volume>
          .
          <fpage>3711</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dauphin</surname>
          </string-name>
          ,
          <article-title>Hierarchical neural story generation, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>889</fpage>
          -
          <lpage>898</lpage>
          . URL: https://aclanthology.org/P18-1082. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P18</fpage>
          -1082.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Holtzman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Buys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Forbes</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Choi,</surname>
          </string-name>
          <article-title>The curious case of neural text degeneration</article-title>
          , CoRR abs/
          <year>1904</year>
          .09751 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1904</year>
          .09751. arXiv:
          <year>1904</year>
          .09751.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Collier</surname>
          </string-name>
          ,
          <article-title>Contrastive search is what you need for neural text generation</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2210</volume>
          .
          <fpage>14140</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>