<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Team GPLSI at AuTexTification Shared Task: Determining the Authorship of a Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Iván Martínez-Murillo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robiert Sepúlveda-Torres</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Estela Saquete</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena LLoret</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Palomar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GPLSI research group, Dept. of Software and Computing Systems, University of Alicante</institution>
          ,
          <addr-line>Ctra. San Vicente s/n, 03690, San Vicente del Raspeig, Alicante</addr-line>
          ,
          <country>España</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Spanish @user @user @user @user Si es así el spoiler, me va a emocionar ver esa localización en concreto @user por mensa Yo supuse que eras tú ¡Hola! ¿Cómo estás? ¿Estoy contento de que hayas adivinado qui</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>AuTexTification is a shared task within the IberLEF workshop which aims to determine whether a text has been generated by an Artificial Intelligence (AI) or a human. The objective of this paper is to report the participation and results of the GPLSI team from the University of Alicante (Spain) in subtask 1: Human or Generated of the AuTexTification challenge for English and Spanish languages. We propose and experiment with diferent approaches based on Transfer Learning; Ensemble Learning; fine-tuning existing language models, such as RoBERTa or RemBERT; or relying on linguistic features. Our best models for both languages were trained through Transfer Learning techniques, obtaining the 6th and 8th position in the English and Spanish versions of this subtask, respectively. Results obtained in the Spanish-version were close to the top-performing team.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Human Language Technologies</kwd>
        <kwd>Transformers</kwd>
        <kwd>Fine-tunning</kwd>
        <kwd>Multilinguality</kwd>
        <kwd>Ensemble classification</kwd>
        <kwd>Transfer Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, Natural Language Processing (NLP) has advanced exponentially, partly due to
the development of Transformers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and Large Language Models (LLMs) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Specifically, the
task of Natural Language Generation (NLG) has benefited from this development. Thanks to this,
some generative models such as GPT4 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], PALM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or BLOOM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] have been deployed with
the ability to produce texts that, in some cases, can be indistinguishable from human-generated
ones. Nevertheless, this rapid development also introduced some risks. On the one hand, even
though automatically generated texts can be semantically correct and written in a human-like
style, the meaning of the message, as well as the message itself, may be inaccurate or not
true, which leads to hallucinations [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Furthermore, bias can be introduced in some cases
during the training of these models, which could induce to produce unethical content [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. On
the other hand, it could be dificult to diferentiate an original work from one generated by a
machine in the academic context. In either of these cases, bad and unethical uses of Artificial
Intelligence (AI) and its related technology could be promoted, for instance, to potentially
generate misinformation.
      </p>
      <p>
        Given this context and taking these problems into account, in recent years, the task of
automatically detecting generated text has gained importance. Some proposals have been made,
such as the AI Text Classifier [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or GPTZero [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. However, these approaches are not completely
reliable yet. In order to fill this gap, AuTexTification shared task [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] has been proposed within
the 2023 IberLEF workshop [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for advancing the state of the art regarding the automatic
detection of generated text in English and Spanish through two diferent subtasks.
• Subtask 1: Human or Generated is a binary text classification task in which it has to
be determined whether a given text has been generated by an AI or not.
• Subtask 2: Model Attribution is a multiclass text classification task that consists in
determining which AI model has generated a given text.
      </p>
      <p>The objective of this paper is to report our participation (Team GPLSI) in the subtask 1 of
AuTexTification - IberLEF 2023 shared task for English and Spanish languages. We will present
the diferent models we have developed for automatically detecting AI-generated texts. We
make use of Transfer Learning models and Ensemble Learning techniques, fine-tuning existing
language models and linguistic feature extraction to address subtask 1 for both languages.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        AI Text detection is a task that has been studied for a while. Nevertheless, it has not been
until recently, when it has really gained great relevance due to the explosion of generative AI.
Particularly, in the context of NLG, despite its advantages, tools such as ChatGPT or Google
Bard can pose some threats to society if they are not used critically or ethically. These tools
generate texts very quickly, hardly indistinguishable from human-written texts, on top of a
wide variety of styles and languages. Considering this, in some sectors, such as academia [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
or medicine [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the need to detect if a text has been generated by an AI has arisen.
      </p>
      <p>
        In light of this, there have been eforts in the research community to address this issue with
the aim of advancing the state of the art. Some shared tasks related to automatic generated text
detection involving diferent scopes have been proposed (e.g., DAGPap22 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or RuATD-22
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]). The best model in DAGPap22 shared task was an ensemble of several fine-tuned models
based on BERT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], while in RuATD-22, the best models were those fine-tuned and based on
BERT with extra features such as tonality, reading ease, or lexical richness [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>In this context, AuTexTification shared task emerges as part of the IberLEF 2023 workshop.
The aim of IberLEF is to promote research in text processing, understanding, and generation
tasks in at least one of the Iberian languages. This 5th edition includes the AuTexTification
shared task, which as explained in Section 1, consists in automatically spotting if a text has been
generated by an AI or by a human (subtask 1) and detecting which AI model has generated a
given text (subtask 2).</p>
    </sec>
    <sec id="sec-3">
      <title>3. AuTexTification Subtask 1 Overview: Human or Generated text</title>
      <p>In this section, we provide a detailed description of subtask 1 of the AuTexTification shared
task. Given a text as input, the aim of this subtask is to classify whether it has been generated
by a human or a machine. This subtask can be addressed for two languages: English and/or
Spanish, having the restriction that participants are only allowed to use the data given by the
organisers to train the models.</p>
      <p>In the next section, an explanation of the train and test datasets is provided.</p>
      <sec id="sec-3-1">
        <title>3.1. Datasets</title>
        <p>In the first stage of the shared-task, only training data was provided, so the participants could
train and propose their approaches. This training data was provided in two distinct files, one
for each language. Unlabelled test data was also provided, which was then labelled at a second
stage, when the shared-task ended.</p>
        <p>The remaining of this section describes the training and test datasets in both languages.</p>
        <p>The first characteristic that is worth mentioning about the datasets is that, in some cases,
they are not complete. Therefore, this task involves an additional challenge because it is not
possible to make a distinction between AI-generated texts and human-written texts attending
to grammatical criteria. An example of this issue can be seen in Table 1.</p>
        <p>English
Let your friend know that you have noticed.</p>
        <p>This can encourage your friend to
Laying down by the West in, One World
Media was the first building to go up,
followed by the Westin, a</p>
        <p>Another fact to underline is the distribution between human and AI-generated texts. Table 2
shows the division of the dataset. It can be highlighted that the Spanish test dataset is not as
balanced as the training dataset. In contrast, the English datasets are more balanced.</p>
        <p>Moreover, analysing the length of these texts (number of characters) shown in Table 3 we
can see that AI-generated texts tend to be longer than human-generated texts for both datasets
in English and Spanish. As well, the test dataset contains longer texts than the training one,
regardless of the type (i.e., human or generated).</p>
        <p>Finally, another aspect to highlight is the domain of texts in both datasets in both languages.
While texts in the training dataset include several scopes such as legal documents, how-to
articles, and social media, texts in the test dataset change the domain to news and reviews.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Team GPLSI Strategy</title>
      <p>In this section the methodology followed to address subtask 1 is outlined.</p>
      <sec id="sec-4-1">
        <title>4.1. Preliminary Statistical analysis</title>
        <p>
          Based on existing literature [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], there are some patterns in a machine-generated text that could
help to distinguish them from a human-written one. Particularly, text generated by a machine
tends not to express sentiments and use more uncommon words. Relying on these findings, we
decided to first conduct an analysis of the presence of some readability and sentiment features
in the training dataset:
• To measure the readability and understandability, we used the textstat library (https:
//pypi.org/project/textstat/). The hypothesis is that AI-generated texts should be more
dificult to read and understand. We employed four diferent functions to calculate
statistics from the training dataset. Obtained results can be seen in Table 4.
1. The readability of a text can be assessed using the Flesch Reading Ease formula
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] (a higher score means that a text is easier to understand).
2. The understandability of a text can be rated by applying the Szigriszt-Pazos formula
[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] (a higher the score is indicates that a texts is easier to understand).
3. Fernandez-Huerta [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] formula measures the readability of the text, but this metric
is specifically used to evaluate texts in Spanish (a bigger value denotes that a text is
easy to read).
4. To calculate the readability of an English text by a foreign learner, we used the
McAlpine EFLAW Readability Score [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] (the lower the score of a text, the easier
it is to read for a foreigner person.). This metric is specifically used for the English
dataset.
        </p>
        <p>
          Based on the results obtained in Table 4, we can extract the conclusion that human-written
texts tend to be easier to read and understand than AI-generated ones. Based on this
ifnding, we consider that using these metrics will help to increase results for this task.
• We calculated sentiment statistics of the training dataset using the NLTK library [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ].
        </p>
        <p>The hypothesis to use this is that texts generated by an AI tend to be more neutral than
texts written by humans.</p>
        <p>Based on the findings of Table 5, we can see that in Spanish, most texts do not express
sentiment, being tagged as neutral. Analysing the other sentiment categories, there are
more texts expressing a positive sentiment in human-written texts than in AI-generated
texts.</p>
        <p>In contrast, the results obtained for the English dataset were not the expected ones. As we
can see in Table 5, most of the texts express a positive sentiment for both human and AI
generated texts, so texts generated by an AI can also be trained to express sentiments. On
the contrary, there is a greater number of human-generated texts expressing a negative
sentiment than texts generated by an AI in the English dataset.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Base Models</title>
        <p>
          In our research, Transfer Learning techniques have been applied. Transfer Learning is a subfield
of machine learning that applies knowledge learned by solving one task to approach a diferent
task [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. In this sense, we have used pre-trained models based on Transformers architecture
that obtains relevant results in NLP tasks. These models are trained in a general task and can
be fine-tuned on a more specific task. For this purpose, the model is set into a fine-tuning mode.
Afterwards, training examples are used to adjust the weights of a language model and the neural
network classifier, which makes the final prediction. Two types of pre-trained models have been
used to address this task, specific language models and multilingual models. The employed
models are described below:
• Specific language model: We use RoBERTa and BETO models trained in English and
Spanish. These are based on BERT, maintaining the same structure but optimising a few
parameters, such as larger training data or larger batch size. These optimisations make
RoBERTa performs better than BERT in the English language. Specifically, we use the
following models to address the English and Spanish tasks:
– English: RoBERTa-base-openai-detector was trained to detect GPT-generated
text. Further information can be found in [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. This model is published at
https://huggingface.co/roberta-base-openai-detector.
        </p>
        <p>
          RoBERTa-large has been pretrained with a collection of five diferent datasets. These
datasets have a total weight 160GB of text. More information, as well as the model
itself, is publicly available at https://huggingface.co/roberta-large [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ].
– Spanish: RoBERTa-base-bne was trained with a total of 570GB of clean Spanish
texts. Further information can be obtained in [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. This model is accesible in
https://huggingface.co/PlanTL-GOB-ES/roberta-base-bne.
        </p>
        <p>
          BETO was instructed with texts in Spanish from Wikipedia and the OPUS project
(https://opus.nlpl.eu/). More information is provided in [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. This model is published
at https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased.
• Multilingual model: We use XLM-RoBERTa and RemBERT models trained on a
multilingual dataset. These models can be trained for specific tasks in diferent languages.
XLM-RoBERTa is a multilingual model trained with 2.5TB of filtered CommonCrawl
data over 100 languages [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. This model can be found in https://huggingface.co/
xlm-roberta-base.
        </p>
        <p>
          As well as XLM-RoBERTa, RemBERT is also a multilingual model which can be seen as a
bigger version of mBERT [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]. This model is trained with 26B tokens of Wikipedia data
over 110 languages and is publicly available in https://huggingface.co/google/rembert.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Base models with features</title>
        <p>
          Some works include additional features to combine them with the output of the last layer of the
transfer learning models [
          <xref ref-type="bibr" rid="ref31 ref32 ref33">31, 32, 33</xref>
          ]. This strategy could improve the prediction of models based
on transfer learning. In this sense, we propose to include the features explained in section 4.1
because the statistical analysis shows diferences between human and generated texts. We relied
on the approach proposed in [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]. Figure 1 shows the internal architecture of this approach.
        </p>
        <p>
          The input text was encoded by using a transformer model. The encoded vector (model output)
is concatenated with the features in the first layer of the neural network. Afterwards, a dropout
layer is applied to prevent overfitting, and finally, an output layer classifies between human and
generated text. The output layer, in this case, is a dense layer with two output neurons using a
softmax activation function and cross-entropy as a loss function. In this case, the neural network
is a Multilayer Perceptron (MLP) with three layers. Our approach adopts the classification
architecture proposed by [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ], which combines Transformer model with external features.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Ensemble Models</title>
        <p>
          Based on the promising results obtained in other challenges [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], Ensemble models could
improve the performance of some tasks. Particularly, we have used an Ensemble Stacking
method. This approach consists in training some models to predict a given task and combining
its predictions with training a meta-learner to output a final prediction. The meta-learner inputs
the predictions of the stacked models as features, and the target label, and it learns how to best
combine the input predictions to make a better output prediction by using a traditional machine
learning algorithm [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental Setup</title>
      <p>
        In this section, the main experiments carried out to make the proposals for the workshop will
be explained. During the training phase, only the training dataset was available. The training
datasets for English and Spanish were, in turn, split into three subsets: one for training the
models, other for validation, and the last one for testing the models. The validation subset was
used to estimate model skills while performing hyperparameter tuning and to compare diferent
approaches. The test subset was used to measure the performance and generalisation of the
model with unseen data [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. Each validation and test subset approximately represent a 12% of
the training sets. Table 6 shows the distribution per class in each subset created.
As can be seen in Table 6, all the subsets are balanced between human and generated classes.
      </p>
      <p>
        To tackle the strategy defined in Section 4, four experiments have been proposed. These
experiments are described below:
1. Monolingual dataset fine-tuning: We fine-tuned some state-of-the-art Spanish, English,
and multilingual pre-trained models with an initial hyperparameter configuration. The
initial hyperparameter configuration was a maximum sequence length of 400, a batch
size of 8, a training rate of 1e-5, a manual seed of 1,509, and training epochs of 3. The
models with better performance to predict the validation subsets and with less overfitting
were used to perform a bayesian hyperparameter tuning. To automate hyperparameter
tuning, we used the Weight &amp; Biases library [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. Table 7 shows the search configuration.
2. Multilingual dataset fine-tuning: The second experiment concatenates English and
Spanish training datasets to train multilingual models with Transformer architecture. This
process can obtain more general models since more training data is available, and also
multilingual models tend to be very stable since they find patterns beyond the features of
each language.
3. Fine-tuning with features: Based on existing literature, adding features extracted from
the training data, such as sentiment or readability, while training a model could improve
its performance. Figure 1 shows the architecture proposed by Sepúlveda-Torres et al., that
was used in this experiment. Six features were employed for each language: a general
readability score (Flesch Reading Ease), a readability language specific score (McAlpine
EFLAW and Fernandez-Huerta, English and Spanish, respectively), an understandability
score, and three features that represent sentiments. Before concatenating the features
to the outputs of the transformer model, a normalisation of the readability and
understandability features were performed. We normalised these features since they were in a
diferent range of values from the sentiment features.
4. Ensemble: The last experiment for both languages is an ensemble classification system,
aiming to integrate the best obtained models and thus, improve prediction results. In this
case, we used a Logistic Regression (LR) algorithm as meta-learner. The logit outputs of
each model are stacked and used as predictors of LR algorithm. These logit outputs also
have been normalised to perform the training and prediction.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Results and discussion</title>
      <p>
        All of the experiments explained in Section 5 were implemented using Simple Transformer [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ]
and PyTorch [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ] libraries. The data employed to train the models were only derived from the
data provided by the challenge. The final hyperparameters configuration for each classifier and
the code implemented can be found at https://github.com/rsepulveda911112/Autextification.
      </p>
      <sec id="sec-6-1">
        <title>6.1. Training phase</title>
        <p>This section shows the results of all the experiments carried out with the diferent Transformers
models and the strategies addressed in section 4.
Each row is annotated according to the experiment it represents. To evaluate the performance
of each experiment, we have applied the oficial metrics of the AuTexTification shared task
(Macro-F1 in percentage mode).
IA vs. human prediction results for English on the validation and test subsets created by us.</p>
        <p>The first experiment in Tables 8 and 9 (rows 1, 2, and 3) obtains competitive results for the
RemBERT model in English and for RoBERTa-base-bne and RemBERT models in Spanish. The
results of this first experiment for these models are obtained after performing a hyperparameter
search.</p>
        <p>The second experiment (rows 4 and 5) achieves high results when using the RemBERT
language model. The other model gets very discrete results.</p>
        <p>In the case of the third experiment, results were not as good as expected. We obtained poor
results when adding features to an English model, and adding features to a Spanish model
performed similarly to without them.</p>
        <p>
          Finally, the best results in both languages were obtained after using the ensemble model,
in the same way as reported in published research works with similar approaches [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. This
approach significantly improved the results in the Spanish dataset. The English ensemble used
Run_1 and Run_2 models. In the case of Spanish, the ensemble has been constructed with the
following models:
• (Run_1) Roberta-base-bne with features.
• Roberta-base-bne with Hyperparameter tuning.
• (Run_2) RemBERT trained with datasets in both languages concatenated.
• RemBERT trained with the Spanish dataset.
        </p>
        <p>Through the training phase, we noticed that models did not overfit when splitting the training
dataset into three subsets, and results obtained when predicting both validation and test subsets
were similar. On account of this, we considered that training our models with more data would
increase our results. So, for making the final submission, we have only split the training data
into two subsets (one for training, and the other for test). Moreover, the Ensemble experiment
was performed after training the final models. Hence, we do not have validation results for the
Ensemble experiment, as it could only be tested with unseen data.</p>
        <p>These experiments carried out guided us to determine the three models to be finally submitted
for the AuTexTification subtask-1, either for English or Spanish. The submitted models will be
evaluated with the test set released by the subtask organisers.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Submission results: Subtask 1 - English</title>
        <p>The three submitted models for subtask 1 in English can be seen in Table 10. As the best results
while experimenting with the English dataset were making use of multilingual models, our
Run_1 and Run_2 are a RemBERT multilingual model. The main diference is that in Run_1
the model has been only trained with the English data, and in Run_2 the model has been trained
with both English and Spanish training data. Finally, Run_3 is an ensemble of both models.</p>
        <p>For the English subtask, our best model has been Run_1. It achieved the 6th position out of
76 proposals, with a score of 72.52 for the Macro-F1 metric. In contrast with the training phase,
Run_2 (the model trained with the multilingual dataset) did not perform as well as the model
trained with just the English dataset reaching 11th place with a score of 70.54 in the Macro-F1
metric. Finally, despite the fact that the ensemble model obtained the best results in the training
phase, Run_3 did not obtain the best results with the test dataset. Nonetheless, it has achieved
a 10th position in the ranking with similar results to our best approach (Macro-F1 of 71.39), so
this means that it is also a good strategy.</p>
        <p>Comparing these results with those obtained by other participants, the top-performing system
in this subtask obtained 80.91, followed by the second best one, which obtained 74.16. Our best
model was very close to the second one in the results.</p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Submission result: Subtask 1 - Spanish</title>
        <p>For the Spanish subtask 1, we have also submitted three diferent models. In this case, Run_1
has been a Spanish model, trained with the Spanish dataset concatenated with features. In the
same way as for English, Run_2 has been the multilingual RemBERT model trained with the
datasets in both languages concatenated. Finally, the proposal for Run_3 has been an ensemble
built with four diferent models.</p>
        <p>Results obtained in the final submission can be seen in Table 11. Our best approach at
predicting the Spanish task has been Run_2, achieving the 8th position in the ranking with
a Macro-F1 of 66.82. The multilingual model trained with the dataset concatenated in both
languages performed satisfactorily for this task. In contrast, extracted features seem not to be
relevant to predict whether a text has been generated by an AI or a human with the given test
dataset because Run_1 model has not achieved the expected results (22nd position obtaining
a Macro-F1 of 63.19). Finally, Ensemble did not perform as well as expected, considering that
Run_3 reached the 20th position with a Macro-F1 of 63.9, presumably because the ensemble
was built with models in diferent languages (multilingual and Spanish only), and it may not be
selecting the relevant features to predict just one language.</p>
        <p>When comparing the results with those achieved by other participants in Spanish, the leading
system in this particular subtask achieved a score of 70.77 in the Macro-F1 metric. In this case,
our best model was very proximate to the first position model.</p>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. General discussion</title>
        <p>The oficial results obtained after the submission of our approaches (Team GPLSI) show a
noticeable diference compared to the results obtained in the experimentation. The reason
for that could be that our models have learned to predict well some specific domains (legal
documents, how-to articles, and social media), but when changing the domain to predict unseen
domains (news and reviews), they do not behave in the same way. However, obtained results
indicate that all our models achieved a high score when predicting AI-generated text, being the
assignment of predicting human written the one which not achieves as good results. This could
happen because, as explained in Section 3.1, both human and AI generated texts are incomplete
in some cases. In turn, this may provoke that models not to learn as well some grammatical
patterns such as punctuation and consequently predict as an AI generated text one which is
actually written by a human.</p>
        <p>Another important issue that is worth discussing is that a multilingual model fine-tuned
with a dataset in a specific language performs better than fine-tuning the same model with
two diferent languages. Nonetheless, results of both models are close. Thereby, in some cases
would be interesting to construct a model that could make multilingual predictions, being more
eficient, than constructing a model that is only able to predict in one language.</p>
        <p>Moreover, looking at the competition results we can see that, generally speaking, models
predicting the Spanish dataset perform worse than the models trained to predict the English
dataset. Such results could arguably be a consequence of the structural diferences between
both languages, as the English language tends to show more pre-fixed grammatical patterns
with simpler structures, whereas in the case of Spanish, the grammatical constructions can be
built with a wider variety of linguistic structures. Consequently, it could be more dificult for
the models to cover every grammatical pattern concerning the Spanish language, which would
cause the worse performance shown in the results.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions and Future Work</title>
      <p>The task of automatically detecting texts generated by an AI is a growing task in the NLP field
as a consequence of the rapid development of NLG. Current generative models can generate
automatic texts that are hardly indistinguishable from human-written texts. This issue has
caused some social concerns.</p>
      <p>In this context, AuTexTification challenge was proposed as part of the IberLEF 2023 workshop.
We (Team GPLSI) participated in subtask 1 for both languages, Spanish and English. The results
achieved in the English subtask are considered good, ranking 6th of 76 participants with a
Macro-F1 of 72.52. In the Spanish subtask, we ranked 8th of 52 participants with a Macro-F1 of
66.82. The interesting fact about this model is that it is a multilingual model fine-tuned with
both datasets (English and Spanish) and achieved a high place in both tasks (8th with a F1 of
66.82 in Spanish and 11th with a F1 of 70.54 in English). This could confirm that fine-tuning a
multilingual pre-trained model could be a good option to automatically detect AI-generated
texts in several languages while being more eficient than building a model to predict just one
language. Another aspect to highlight is that results predicting AI-generated texts are high,
achieving almost a Macro-F1 of 80 in all of our models. However, these models do not predict
human written texts as well as when it comes to AI-generated texts.</p>
      <p>One future line of work is to analyse and extract other linguistic patterns in human texts
to train the model and help it to better understand the way humans write, and consequently,
increase the results obtained at predicting it. In addition to this, we will experiment with
other approaches, such as Active Learning or Generative Adversarial Networks to measure
how well these approaches perform in comparison with the proposed models during this task.
Furthermore, we will address these tasks with no restrictions on the data so we can train
our models with more text domains with the objective to make a better generalisation of the
knowledge and, consequently, predicting other domains better.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This research work is part of the R&amp;D projects “CORTEX: Conscious Text Generation”
(PID2021123956OB-I00) and “TRIVIAL: Technological Resources for Intelligent VIral AnaLysis through
NLP” (PID2021-122263OB-C22), both funded by MCIN/ AEI/10.13039/501100011033/ and by
“ERDF A way of making Europe”, and “CLEAR.TEXT:Enhancing the modernization
public sector organizations by deploying Natural Language Processing to make their digital
content CLEARER to those with cognitive disabilities” (TED2021-130707B-I00), funded by
MCIN/AEI/10.13039/501100011033 and “European Union NextGenerationEU/PRTR”. Moreover,
it has been also partially funded by the Generalitat Valenciana through the project “NL4DISMIS:
Natural Language Technologies for dealing with dis- and misinformation with grant
reference (CIPROM/2021/21)", and by the European Commission ICT COST Action “Multi-task,
Multilingual, Multi-modal Language Generation” (CA18231).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          , et al.,
          <article-title>A survey of large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2303.18223</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] OpenAI, Gpt-4
          <source>technical report</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>08774</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Anil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Firat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lepikhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shakeri</surname>
          </string-name>
          , E. Taropa,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          , et al.,
          <source>Palm 2 technical report, arXiv preprint arXiv:2305.10403</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Akiki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pavlick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ilić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hesslow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Castagné</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Luccioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yvon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gallé</surname>
          </string-name>
          , et al.,
          <article-title>Bloom: A 176b-parameter open-access multilingual language model</article-title>
          ,
          <source>arXiv preprint arXiv:2211.05100</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Frieske</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ishii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Bang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Madotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <article-title>Survey of hallucination in natural language generation</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Homolak</surname>
          </string-name>
          ,
          <article-title>Opportunities and risks of chatgpt in medicine, science, and academic publishing: a modern promethean dilemma</article-title>
          ,
          <source>Croatian Medical Journal</source>
          <volume>64</volume>
          (
          <year>2023</year>
          )
          <article-title>1</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>OpenAI</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://beta.openai.com/ai-text-classifier.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Tian</surname>
          </string-name>
          , Gptzero,
          <source>Retrieved Jan</source>
          <volume>25</volume>
          (
          <year>2023</year>
          )
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>A. M. Sarvazyan</surname>
            ,
            <given-names>J. Á.</given-names>
          </string-name>
          <string-name>
            <surname>González</surname>
            ,
            <given-names>M. Franco</given-names>
          </string-name>
          <string-name>
            <surname>Salvador</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chulvi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          , Overview of AuTexTification at IberLEF 2023:
          <article-title>Detection and Attribution of MachineGenerated Text in Multiple Domains</article-title>
          ,
          <source>in: Procesamiento del Lenguaje Natural</source>
          , Jaén, Spain,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes-y</surname>
          </string-name>
          <string-name>
            <surname>Gómez</surname>
          </string-name>
          ,
          <source>Overview of IberLEF 2023: Natural Language Processing Challenges for Spanish and other Iberian Languages, Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>M. M. Rahman</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Watanobe</surname>
          </string-name>
          ,
          <article-title>Chatgpt for education and research: Opportunities, threats</article-title>
          , and strategies,
          <source>Applied Sciences</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>5783</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>L. De Angelis</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Baglivo</surname>
            , G. Arzilli,
            <given-names>G. P.</given-names>
          </string-name>
          <string-name>
            <surname>Privitera</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>A. E.</given-names>
          </string-name>
          <string-name>
            <surname>Tozzi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <article-title>Chatgpt and the rise of large language models: the new ai-driven infodemic threat in public health</article-title>
          ,
          <source>Frontiers in Public Health</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <fpage>1567</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kashnitsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Herrmannova</surname>
          </string-name>
          , A. de Waard, G. Tsatsaronis,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fennell</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Labbé, Overview of the dagpap22 shared task on detecting automatically generated scientific papers</article-title>
          ,
          <source>in: Third Workshop on Scholarly Document Processing</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Shamardina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mikhailov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chernianskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fenogenova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saidov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Valeeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shavrina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Smurov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tutubalina</surname>
          </string-name>
          , E. Artemova,
          <article-title>Findings of the the ruatd shared task 2022 on artificial text detection in russian</article-title>
          ,
          <source>arXiv preprint arXiv:2206.01583</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Glazkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Glazkov</surname>
          </string-name>
          ,
          <article-title>Detecting generated scientific papers using an ensemble of transformer models</article-title>
          ,
          <source>arXiv preprint arXiv:2209.08283</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Maloyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nutfullin</surname>
          </string-name>
          , E. Ilyushin, Dialog-22
          <source>ruatd generated text detection</source>
          ,
          <source>arXiv preprint arXiv:2206.08029</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mitrović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Andreoletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Ayoub</surname>
          </string-name>
          ,
          <article-title>Chatgpt or human? detect and explain. explaining decisions of machine learning model for detecting short chatgpt-generated text</article-title>
          ,
          <source>arXiv preprint arXiv:2301.13852</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Flesch</surname>
          </string-name>
          ,
          <article-title>A new readability yardstick</article-title>
          .,
          <source>Journal of applied psychology</source>
          <volume>32</volume>
          (
          <year>1948</year>
          )
          <fpage>221</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Szigriszt</surname>
          </string-name>
          <string-name>
            <surname>Pazos</surname>
          </string-name>
          ,
          <article-title>Sistemas predictivos de legilibilidad del mensaje escrito: fórmula de perspicuidad (</article-title>
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fernández</surname>
          </string-name>
          <string-name>
            <surname>Huerta</surname>
          </string-name>
          , Medidas sencillas de lecturabilidad,
          <source>Consigna</source>
          <volume>214</volume>
          (
          <year>1959</year>
          )
          <fpage>29</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>R.</given-names>
            <surname>McAlpine</surname>
          </string-name>
          , From plain english to global english,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          , E. Loper,
          <article-title>Natural language processing with Python: analyzing text with the natural language toolkit, "</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Xi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey on transfer learning</article-title>
          ,
          <source>Proceedings of the IEEE</source>
          <volume>109</volume>
          (
          <year>2020</year>
          )
          <fpage>43</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>I.</given-names>
            <surname>Solaiman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brundage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Herbert-Voss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          , G. Krueger,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kreps</surname>
          </string-name>
          , et al.,
          <article-title>Release strategies and the social impacts of language models</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>09203</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized BERT pretraining approach</article-title>
          , CoRR abs/
          <year>1907</year>
          .11692 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1907</year>
          .11692. arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Palao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Oller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Penagos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Maria: Spanish language models</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          ). URL: https://upcommons.upc.edu/handle/2117/367156# .YyMTB4X9A-0.mendeley. doi:
          <volume>10</volume>
          .26342/2022-68-3.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          ,
          <source>in: PML4DC at ICLR</source>
          <year>2020</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          , arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>02116</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fevry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tsai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , S. Ruder,
          <article-title>Rethinking embedding coupling in pre-trained language models</article-title>
          , arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>12821</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Wu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Semantics-aware BERT for language understanding</article-title>
          ,
          <source>in: The Thirty-Fourth AAAI Conference on Artificial Intelligence</source>
          ,
          <source>AAAI</source>
          <year>2020</year>
          , The Thirty-Second
          <source>Innovative Applications of Artificial Intelligence Conference</source>
          ,
          <source>IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI</source>
          <year>2020</year>
          , New York, NY, USA, February 7-
          <issue>12</issue>
          ,
          <year>2020</year>
          , AAAI Press,
          <year>2020</year>
          , pp.
          <fpage>9628</fpage>
          -
          <lpage>9635</lpage>
          . URL: https://ojs.aaai.org/index.php/AAAI/article/view/6510.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>W. M.</given-names>
            <surname>Lim</surname>
          </string-name>
          , H. T. Madabushi, UoB at SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <article-title>Boosting BERT with Corpus Level Information (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sepúlveda-Torres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Saquete</surname>
          </string-name>
          , E. Lloret,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palomar</surname>
          </string-name>
          ,
          <article-title>Leveraging relevant summarized information and multi-layer classification to generalize the detection of misleading headlines</article-title>
          ,
          <source>Data &amp; Knowledge Engineering</source>
          <volume>145</volume>
          (
          <year>2023</year>
          )
          <article-title>102176</article-title>
          . URL: https:// www.sciencedirect.com/science/article/pii/S0169023X23000368. doi:https://doi.org/ 10.1016/j.datak.
          <year>2023</year>
          .
          <volume>102176</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <article-title>Ensemble stacking for machine learning</article-title>
          and
          <source>deep learning</source>
          ,
          <year>2021</year>
          . URL: https://www.analyticsvidhya.com/blog/2021/08/ ensemble-stacking
          <article-title>-for-machine-learning-and-deep-learning/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence : a modern approach, fourth edition</article-title>
          ,
          <source>Pearson series in artificial intelligence</source>
          , global ed. ed.,
          <string-name>
            <surname>Pearson</surname>
            <given-names>Education</given-names>
          </string-name>
          , Harlow,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L.</given-names>
            <surname>Biewald</surname>
          </string-name>
          ,
          <article-title>Experiment tracking with weights and biases, 2020</article-title>
          . URL: https://www.wandb. com/, software available from wandb.
          <source>com.</source>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Rajapakse</surname>
          </string-name>
          , Simple transformers, https://github.com/ThilinaRajapakse/ simpletransformers,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paszke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Massa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , G. Chanan,
          <string-name>
            <given-names>T.</given-names>
            <surname>Killeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gimelshein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antiga</surname>
          </string-name>
          , et al.,
          <article-title>Pytorch: An imperative style, high-performance deep learning library</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>