<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DeepR3: Reducing, Reusing and Recycling Large Models for Developing Responsible and Green Language Technologies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aitor Soroa</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>German Rigau</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose M. Alonso-Moral</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcos Garcia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maite Melero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marta Villegas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Barcelona Supercomputing Center</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS)</institution>
          ,
          <addr-line>Universidade de Santiago de Compostela</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>HiTZ Basque Center for Language Technology - Ixa NLP Group, University of the Basque Country UPV/EHU</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the DeepR3 project, a coordinated project composed of three local projects at the Hitz Centre (University of the Basque Country), CiTIUS (University of Santiago de Compostela) and Barcelona Supercomputing Center, respectively. The main objective of DeepR3 is to research on parameter eficient ways to extend existing pre-trained language models for Spanish, Catalan, Basque, Galician plus English, and adapt them to new domains, genres and languages. In this project, we will apply the newly developed techniques to improve the state of the art on text generation tasks in the mentioned languages, reusing and recycling pre-trained models for machine translation, developing advanced content-based domain applications in sectors such as Meteorology, and developing new benchmarks and datasets for evaluating and assessing progress towards responsible natural language understanding and generation. DeepR3 is funded by MCIN/AEI/10.13039/501100011033 and by the European Union NextGenerationEU/PRTR.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computational Linguistics</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Ethical Guidelines</kwd>
        <kwd>Trustworthy AI</kwd>
        <kwd>Sustainable AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        pre-trained language models. Despite their impressive
capabilities, these models do come with severe drawbacks
The Natural Language Processing (NLP) community is from a research advancement, environmental, and
ethcurrently engaged in a paradigm shift with the produc- ical perspective. These models are extremely costly to
tion and exploitation of large pre-trained transformer- train and develop, both financially, due to the cost of
based language models. Compared to the previous state hardware and electricity or cloud computing time, and
of the art, the results are so good that systems are claimed environmentally, due to the carbon footprint required to
to obtain human-level performance when evaluated in fuel modern servers with multiple Graphics Processing
dificult language understanding tasks. Some authors call Unit (GPU) hardware. This also means that only a
limthese models “foundation models” to underscore their ited number of organisations with abundant resources
critically central yet incomplete character. This paradigm in terms of funding, computing capabilities, NLP experts
shift means that we have only just started to discover and data can currently aford to develop and deploy such
the new possibilities and concerns raised by these large models. A growing concern is that due to unequal
access to computing power, only certain firms and elite
SEPLN-CEDI-PD 2024: Seminar of the Spanish Society for Natural research groups can access modern Artificial Intelligence
Language Processing: Projects and System Demonstrations, June (AI) research [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We have no clear understanding of how
19-20, 2024, A Coruña, Spain large language models work, when they fail, or what
$ a.soroa@ehu.eus (A. Soroa); german.rigau@ehu.eus (G. Rigau); emergent properties they may present, as well as novel
jmosaercmoasr.giaa.racloian.gsoo.nmzaolreazl@@uusscc..gesal(J(.MM. .GAalrocnias)o;-mMaoirtea.lm); elero@bsc.es ways of exploiting these models eficiently. There are
(M. Melero); marta.villegas@bsc.es (M. Villegas) also worrying shortcomings in the text corpora used to
 https://ixa2.si.ehu.eus/asoroa/ (A. Soroa); train these anglo-centric models, ranging from a lack of
https://adimen.si.ehu.es/~rigau/ (G. Rigau); representation of less-resource languages such as
Catahttps://citius.gal/es/team/jose-maria-alonso-moral lan, Basque or Galician, to a predominance of harmful
(hJt.tMps.:/A/cloitniusos.-gMalo/treaal)m;/marcos-garcia-gonzalez (M. Garcia); stereotypes, and to the inclusion of personal information.
https://www.bsc.es/melero-nogues-maite (M. Melero); To tackle these questions, much critical interdisciplinary
https://www.bsc.es/villegas-marta (M. Villegas) collaboration and research are needed.
      </p>
      <p>
        0000-0001-8573-2654 (A. Soroa); 0000-0003-1119-0930 (G. Rigau); The DeepR3 project aims to address some of the
impor0000-0003-3673-421X (J. M. Alonso-Moral); 0000-0002-6557-0210 tant concerns that large language models raise from a
re((MM.. GVialrlecigaa)s;)0000-0001-9933-3224 (M. Melero); 0000-0003-0711-0029 search, innovation, but specifically from an
environmen© 2024 Copyright for this paper by its authors. Use permitted under Creative tal perspective. Instead of creating very expensive
monoCPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org) lingual and multilingual language models from scratch,
the aim of the project is to investigate parameter efi- electricity, and environmentally, due to the carbon
footcient methods to reuse and extend existing pre-trained print required to fuel multiple modern GPU hardware.
language models for Spanish, Catalan, Basque, Galician Nowadays, model training and development of large
lanplus English, and adapt them to new domains, genres guage models are likely to make up a substantial portion
and languages. The adapted models will be applied to of the greenhouse gas emissions attributed to the NLP
diferent use cases and tasks such as machine translation, area. In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] the authors benchmarked model training
applications in the biomedical domain and text genera- and development costs in financial terms and estimated
tion from meteorological data. The DeepR3 project may carbon dioxide emissions. While the average human is
therefore have a direct impact on both the Ecological responsible for an estimated five tons of carbon dioxide
Transition by avoiding unnecessary waste of energy and per year2, the authors trained a big neural architecture
the Digital Transition by creating better language models and estimated that the training process emitted 284 tons
for more languages that can be applied to multiple NLP of carbon dioxide.
tasks. There have been several proposals to address these
      </p>
      <p>
        DeepR3 follows the Findable, Accessible, Interoperable issues and develop methods that efectively train models
and Re-usable (FAIR) principles. First, the results will be without updating every parameter at every training
itused by the project partners which already provide NLP eration. Many of these techniques have been developed
products and services to third party companies. Second, within the framework of distributed or federated learning,
all project results are and will be distributed under open where individual workers train the models locally and
source licenses to many more additional end-users of communicate their changes to a centralized server [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
these technologies (e.g., academia, industry and public In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] the authors propose a method that chooses a small
administrations). In order to maximize impact and en- subset of the model parameters to update, according to
sure uptake of the methods developed in DeepR3 by the the relative importance of the parameter in the final
outscientific community, the advance in the state of the art put of the model. By training only on a small fraction of
is demonstrated by applying a carefully designed eval- the parameters (as few as 0.5%), they attain similar
peruation methodology, with results (software, data and formance to training all parameters. Certain models can
extended language models) made publicly available to have individual submodels added, removed, or updated
ensure reproducibility (by hosting datasets, code and data while the remainder of the parameters remain fixed. For
in public repositories such as Huggingface). We are cur- instance, in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] authors proposes the mixture-of-experts
rently involved in the organization of a workshop that is (MoE) layer, which consists of small feed sub-networks
directly aligned with the topic of green language models, and a trainable gating system that learns how to combine
collocated with the COLING-LREC2024 conference. A the outputs of the networks for each particular example.
second workshop entitled “Multimodal, interactive and At modest training budgets, MoEs can match the
perforafective eXxplainable AI (XAI)” has been accepted for mance of dense models using four times less computing
the European Conference on Artificial Intelligence (ECAI) efort [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
2024. It will include a special track on “Building, Reusing, In the transfer learning stage, standard fine-tuning
upRecycling, and Reducing Resources for Multi-modal XAI”. dates the whole parameter set of the model, and therefore
a language model fine-tuned for a specific task needs to
change and store all the parameters of the original model.
2. Related Work While not as expensive as the pre-training stage, this
parameter ineficiency still incurs high computational and
Language models that are pre-trained following a self- environmental costs, especially in the presence of many
supervision approach are usually very big. For in- downstream tasks. Recently, there has been a growing
instance, GPT-3 contains 175 billions of parameters and was terest in developing more parameter-eficient fine-tuning.
trained on 570 gigabytes of text [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], with a cost estimated Adapter modules are a light-weight alternative to full
at more than four million USD1. There are even bigger model fine-tuning, consisting of only a tiny set of newly
models such as Megatron-530B from NVIDIA, which con- introduced parameters at every transformer layer [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In
tains 530 billion trainable parameters which are also the [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the authors propose a new transfer learning method
number of activated model parameters per input token based on merging multiple models into one. For this, they
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. All these models are trained using some variation of develop a merging method that takes into account the
imthe gradient descent method, which requires updating all portance of each parameter when computing the average
the parameters of the model at every training iteration. of diferent models, and find that this form of merging can
As a consequence, building the models requires weeks eficiently transfer knowledge across models fine-tuned
and months of processing power. This incurs enormous from the same pre-trained checkpoint. In [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] authors
costs, both financially, due to the cost of hardware and
1https://lambdalabs.com/blog/demystifying-gpt-3/
2https://ourworldindata.org/co2-emissions
also show that it is possible to reach similar performances
on many downstream-tasks using much smaller language
models pre-trained with knowledge distillation, resulting
in models that are lighter and faster at inference time,
while also requiring a smaller computational training
budget [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ].
leader who coordinates the tasks in the WP, the
relationships and dependencies with other WPs, and contributes
to checking the overall quality of the project outcomes.
      </p>
      <p>Moreover, each task is coordinated by a researcher from
the research team.</p>
      <p>The technical WPs are briefly introduced below:</p>
    </sec>
    <sec id="sec-2">
      <title>3. Objectives</title>
      <p>The main goal of DeepR3 is to obtain substantial scientific
and technical contributions in many NLP tasks. In order
to achieve this goal, the project proposes:
• Developing eficient methods for extending to
new domains, genres and languages, existing
language models for the oficial languages of Spain
(Spanish, Catalan, Basque and Galician) plus
English.
• Exploring novel ways to pre-train and finetune
language models in a parameter-eficient way,
therefore lowering the carbon footprint required
to train such models.
• Addressing language understanding tasks by text
generation.
• Addressing explainability of Deep Learning
(DL)based language models for Natural Language
Generation (NLG) tasks.
• Developing new eficient techniques to reuse and
recycle pre-trained models for MT, especially for
settings with few or non-existing parallel data.
• Applying the newly developed techniques to
improve the state of the art in Natural Language
Understanding (NLU).
• Developing new benchmarks and datasets for
evaluating and assessing progress towards
responsible NLU and NLG.
• Developing a number of advanced content-based
domain applications for the oficial languages in
Spain (Spanish, Catalan, Basque, and Galician)
and English, in multiple sectors and domains
(Meteorology, Health, Tourism, Public
Administrations, etc).</p>
    </sec>
    <sec id="sec-3">
      <title>4. Methodology</title>
      <sec id="sec-3-1">
        <title>The duration of DeepR3 is 24 months. DeepR3 is a coordi</title>
        <p>nated project organized in three sub-projects with shared
methodology, objectives and techniques. The general
coordination is carried out by the HiTZ-UPV/EHU group.</p>
        <p>To achieve the objectives explained earlier, the work has
been organized in 5 technical Work Packages, plus a
WP0 Project Management and a WP6 for Dissemination
and Exploitation. Each work package has a responsible
WP1 Methodology and Design. The purpose of this</p>
        <p>WP is to define the overall methodology of the
project. This includes the specific defining
standard protocols, information flow and
architectures, adapting and integrating the modules,
resources, data structures, data formats, computing
facilities and module APIs in the DeepR3 project.</p>
        <p>The work package will also provide technical
coordination and support for the successful
achievement of the methodological and technological
objectives. In addition, we will consider
transversally through the entire project issues related to
Responsible AI as well as Ethical, Legal, Social,</p>
        <p>Economical, and Cultural (ELSEC) perspectives.</p>
        <p>WP2 Data Management. We target the collection and
curation of the corpora and data needed for
extending existing monolingual and multilingual
language models to new domains, genres and
languages, as well as the datasets required to
evaluate them. In this WP, we will also address
publishing, legal and ethical issues.</p>
        <p>WP3 Adapting compact language models. We
focus on providing eficient ways of adapting
pretrained language models to new domains and
languages. An important requirement is providing
these models in the most compact manner, in
terms of model size, to make their use as cost
eficient as possible. The goal is to recycle large
preexisting models instead of training from scratch
and to be environmentally responsible. Attention
will be paid to explainability and fairness of the
models. The recycled models obtained will be the
basis of a new generation of language models for</p>
        <p>Spanish, Catalan, Basque and Galician.</p>
        <p>WP4 Deployment of language models. We will
deploy the language models obtained in WP3
and adapt them to specific tasks. We will
research modular and parameter-eficient
finetuning methods as an eficient way to adapt
language models to new tasks and languages. The
language models developed in WP3 will be the
basis to build a new generation of high-level
semantic processing tasks in the biomedical domain.</p>
        <p>Another task that requires new and innovative
ways to reuse pre-trained language models is MT
in low-resource scenarios. Finally, many other
interesting tasks such as summarisation or question
answering require generative models.</p>
        <p>WP5 Evaluation and Assessment. The objective of shows that MedMT5 outperforms both encoders and
simthis WP is to measure the research progress via ilarly sized text-to-text models for the Spanish, French,
objective evaluation metrics and relevant open and Italian benchmarks while being competitive with
evaluation campaigns. In addition to standard current state-of-the-art LLMs in English.
metrics and comparative analysis for the
technologies developed in each WP, we plan to de- 5.3. Resources for Linguistic Evaluation
velop datasets and metrics that help to assess
the linguistic generalisation capabilities of the of Language Models
language models and resources developed in the
project (WP3 and WP4).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Current results</title>
      <sec id="sec-4-1">
        <title>This section provides details about work-in-progress which is currently undertaken within the project.</title>
        <sec id="sec-4-1-1">
          <title>5.1. Scaling Laws in Low-Resource</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Settings</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>There have been many works proposing formulas that</title>
        <p>
          relate the size of the model, the size of the dataset and the
computing budget, the so called “scaling laws”. However,
these laws are focused on rich languages and a large scale,
where the size of the available data is virtually infinite.
We have analyzed the efect those variables have on the
performance of language models in constrained settings,
by building lightweight models trained over a set of small
corpora [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Our conclusions conclude that the power
laws for parameters, data and compute for low-resource
settings difer from the optimal scaling laws previously
inferred, and data requirements should be higher. Our
insights are consistent across all the languages we study, as
well as across the MLM and downstream tasks.
Furthermore, we experimentally establish when the cost of using
a Transformer-based approach is worth taking, instead
to favouring other computationally lighter solutions.
        </p>
        <sec id="sec-4-2-1">
          <title>5.2. The medical MedMT5</title>
          <p>We have built MedMT5, the first open-source text-to-text
multilingual model for the medical domain. While there
already exist Large Language Models (LLMs) that have
been adapted to the medical domain, they have been
pretrained and evaluated with a focus on a single language
(English mostly). This is particularly true of text-to-text
models, which typically require large amounts of
domainspecific pre-training data, often not easily accessible for
many languages. MedMT5 has been trained on the largest
multilingual corpus for the medical domain in four
languages, namely English, French, Italian and Spanish. We
also developed two new evaluation benchmarks for all
four languages with the aim of facilitating multilingual
research in this domain. A comprehensive evaluation</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Regarding resources for the linguistic evaluation of lan</title>
        <p>
          guage models, we have made progress in the area of
Targeted Syntactic Evaluation, designing new datasets
for Spanish and Galician focused not only on
syntactic dependencies but also on the semantic properties
of control verbs [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Our results show that, although
transformer-based models have a high performance in
resolving syntactic dependencies, their performance drops
dramatically in cases in which syntax and semantics are
in interaction. Furthermore, we have created the Galician
Parallel Universal Dependencies (PUD) treebank, a new
manually annotated corpus for Galician which follows
the latest Universal Dependencies guidelines [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The
fact that the treebank is a parallel corpus translated by
professionals makes it an interesting resource for
evaluating machine translation models between Galician and
other of the 23 languages that have a PUD treebank.
        </p>
        <sec id="sec-4-3-1">
          <title>5.4. Reusing, Recycling and Reducing</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>Pre-trained Models for Developing</title>
          <p>and Evaluating Green Data-to-text</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>Systems: A Use Case with</title>
        </sec>
        <sec id="sec-4-3-4">
          <title>Meteorological Data</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>We have analyzed empirically how the reuse, recycling,</title>
        <p>
          and reduction of pre-trained language models can
enhance environmental sustainability in NLP. More
precisely, we paid attention to pre-trained models
performing sequence-to-sequence text generation in
Spanish and Galician with the METEOGALICIA-ES and
METEOGALICIA-GL datasets [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. We developed an
experimental pipeline for systematic experimentation,
facilitating the definition of baselines, creation of alternative
models, text generation, and comprehensive evaluation,
which includes assessing values related to energy
consumption and the quality of generated text to determine
the optimal model. In light of the reported results, we
validated the following research hypothesis: “Employing
knowledge transfer techniques enables the creation of
low-cost language models that yield results equivalent to
or superior to those produced by baseline models with a
higher computational cost.”
        </p>
        <sec id="sec-4-4-1">
          <title>5.5. An Empirical Study on the Number of</title>
        </sec>
        <sec id="sec-4-4-2">
          <title>Items in Human Evaluation of</title>
        </sec>
        <sec id="sec-4-4-3">
          <title>Automatically Generated Texts</title>
          <p>
            We have analyzed empirically the efect of the number
of items, i.e., texts to be evaluated, for the sake of
reproducibility in human evaluation of NLG systems. In
light of the reported results, we validated the following
research hypothesis: “well-known resampling statistical
methods can contribute to getting significant results even
with a small number of items to be evaluated by each
evaluator." [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]
          </p>
        </sec>
        <sec id="sec-4-4-4">
          <title>5.6. Enriching Interactive Explanations</title>
          <p>with Fuzzy Temporal Constraint</p>
        </sec>
        <sec id="sec-4-4-5">
          <title>Networks</title>
          <p>
            We have designed, implemented and validated a new
model for fuzzy temporal reasoning to overcome some
inconsistencies detected in pre-trained language
models [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ]. As a proof of concept, the new model is
integrated with TimeVersa, a conversational agent carefully
designed for acting as a virtual assistant for tourists. It
handles imprecise temporal constraints when providing
users with multilingual recommendations, and related
explanations, which are endowed with a good balance
between naturalness and fidelity. Taking as starting point
a knowledge graph that provides an intuitive
representation of the entities and relations in the application
domain, the temporal information is mapped onto a fuzzy
temporal constraint network. This way we represent
imprecise temporal information and are able to check
consistency in conversations.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Further work</title>
      <p>We plan to work on the main areas of the project with
main attention to extending and reusing existing
language models and adapting them to new new domains,
genres and languages, and with a special focus on the
languages spoken in the Iberian peninsula. We will continue
the work on the generation of data-to-text systems, and
extend it to new geographical areas such as the Basque
Country. We also plan to develop machine translation
systems for Iberian languages and English by reusing
existing pre-trained models in these languages. Finally,
we will continue developing benchmarks and datasets for
evaluating and assessing progress towards responsible
NLP, regarding both NLU and NLG tasks.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <sec id="sec-6-1">
        <title>This publication is TED2021-130295B-C31, supported by Grants TED2021-130295B</title>
        <p>C32, and TED2021-130295B-C33 funded by
MCIN/AEI/10.13039/501100011033 and by the “European
Union NextGenerationEU/PRTR”.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wahed</surname>
          </string-name>
          ,
          <article-title>The de-democratization of ai: Deep learning</article-title>
          and
          <source>the compute divide in artificial intelligence research</source>
          ,
          <year>2020</year>
          . arXiv:
          <year>2010</year>
          .15581.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          , N. Ryder, e. a. Subbiah,
          <article-title>Language models are few-shot learners</article-title>
          , in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>M. F.</given-names>
          </string-name>
          <string-name>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Patwary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Norick</surname>
          </string-name>
          , P. LeGresley, S. Rajbhandari,
          <string-name>
            <given-names>J.</given-names>
            <surname>Casper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Prabhumoye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zerveas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korthikanti</surname>
          </string-name>
          , et al.,
          <article-title>Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model</article-title>
          ,
          <source>arXiv preprint arXiv:2201.11990</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Strubell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <article-title>Energy and policy considerations for deep learning in NLP, in: Proceedings of the ACL, Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>3645</fpage>
          -
          <lpage>3650</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Aji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Heafield</surname>
          </string-name>
          ,
          <article-title>Sparse communication for distributed gradient descent</article-title>
          , in: M.
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hwa</surname>
          </string-name>
          , S. Riedel (Eds.),
          <source>Proceedings of EMNLO, Association for Computational Linguistics</source>
          , Copenhagen, Denmark,
          <year>2017</year>
          , pp.
          <fpage>440</fpage>
          -
          <lpage>445</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Sung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <article-title>Training neural networks with fixed sparse masks</article-title>
          , in: M.
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Vaughan</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>34</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2021</year>
          , pp.
          <fpage>24193</fpage>
          -
          <lpage>24205</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mirhoseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Maziarz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Outrageously large neural networks: The sparsely-gated mixture-of-experts layer</article-title>
          ,
          <source>in: Proceedings of ICLR</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhosale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mihaylov</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>e</year>
          . a. Ott,
          <article-title>Eficient large scale language modeling with mixtures of experts</article-title>
          , in: Y.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kozareva</surname>
          </string-name>
          , Y. Zhang (Eds.),
          <source>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>11699</fpage>
          -
          <lpage>11732</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pfeifer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Vulić</surname>
          </string-name>
          , I. Gurevych, S. Ruder,
          <string-name>
            <surname>MADX:</surname>
          </string-name>
          <article-title>An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer</article-title>
          ,
          <source>in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>7654</fpage>
          -
          <lpage>7673</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <article-title>Merging models with fisherweighted averaging</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2111</volume>
          .
          <fpage>09832</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Webson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bach</surname>
          </string-name>
          , L. Sutawika, e. a. Zaid Alyafeai,
          <article-title>Multitask prompted training enables zero-shot task generalization</article-title>
          ,
          <source>in: Proceedings of ICRL</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bucilâ</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Caruana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Niculescu-Mizil</surname>
          </string-name>
          ,
          <article-title>Model compression</article-title>
          ,
          <source>in: Proceedings of KDD, KDD '06</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2006</year>
          , p.
          <fpage>535</fpage>
          -
          <lpage>541</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Distilling the knowledge in a neural network</article-title>
          .,
          <source>CoRR abs/1503</source>
          .02531 (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Urbizu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Vicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Saralegi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agerri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soroa</surname>
          </string-name>
          ,
          <article-title>Scaling laws for bert in low-resource settings</article-title>
          ,
          <source>in: Findings of the ACL</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>I. de Dios-Flores</surname>
            ,
            <given-names>J. Garcia</given-names>
          </string-name>
          <string-name>
            <surname>Amboage</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia</surname>
          </string-name>
          ,
          <article-title>Dependency resolution at the syntax-semantics interface: psycholinguistic and computational insights on control dependencies</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
          </string-name>
          , J. BoydGraber, N. Okazaki (Eds.),
          <source>Proceedings of ACL, Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>203</fpage>
          -
          <lpage>222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>X.</given-names>
            <surname>Sánchez-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarymsakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <article-title>Increasing manually annotated resources for Galician: the Parallel Universal Dependencies Treebank</article-title>
          ,
          <source>in: Proceedings of PROPOR</source>
          , volume
          <volume>1</volume>
          , Association for Computational Linguistics,
          <year>2024</year>
          , pp.
          <fpage>694</fpage>
          -
          <lpage>699</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>González-Corbelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Alonso-Moral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Diz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Taboada</surname>
          </string-name>
          ,
          <article-title>Dealing with hallucination and omission in neural Natural Language Generation: A use case on meteorology</article-title>
          ,
          <source>in: Proceedings of NLG, ACL</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>121</fpage>
          -
          <lpage>130</lpage>
          . Https://aclanthology.org/
          <year>2022</year>
          .inlg-main.
          <volume>10</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalez-Corbelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Alonso-Moral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Crujeiras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bugarín-Diz</surname>
          </string-name>
          ,
          <article-title>An empirical study on the number of items in human evaluation of automatically generated texts</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Canabal-Juanatey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Alonso-Moral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Catala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bugarín-Diz</surname>
          </string-name>
          ,
          <article-title>Enriching interactive explanations with fuzzy temporal constraint networks</article-title>
          ,
          <source>International Journal of Approximate Reasoning</source>
          (
          <year>2024</year>
          )
          <fpage>109128</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>