<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLPalma @ CLEF 2023 JOKER: A BLOOMZ and BERT Approach for Wordplay Detection and Translation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victor Manuel Palma Preciado</string-name>
          <email>victorpapre@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carolina Palma Preciado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <email>sidorov@cic.ipn.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto Politécnico Nacional de México</institution>
          ,
          <addr-line>Gustavo A. Madero, Ciudad de México</addr-line>
          ,
          <country country="MX">México</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Université de Bretagne Occidentale</institution>
          ,
          <addr-line>HCTI</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>The following work has the purpose of describing the participation in the JOKER 2023 track in the classification Task 1 in which it is asked to classify sentences that contain wordplay and the translation Task 3 in which this same type of sentence has to be translated into Spanish trying to maintain the wordplay and the sense in a certain way tropicalizing the sentence to Spanish language from English. Given the constant use of models such as GPT and T5, it was decided to take the direction of language models, making a slight finetuning of BLOOMZ &amp; mT5 models with which to address the problem of translation with a simple prompt and on the other hand the use of BERT as a model proposed for the classification task. The results were mixed since in the manual review of the BLOOMZ &amp; mT5 translated sentences not many were found to contain the Spanish pun, in the case of classification somewhat similar results were obtained given the structure of the dataset. It is believed that given their performance the models can still be optimized and improve the accuracy with which they classify, in the case of the translations perhaps another methodology could be chosen to improve the results obtained in this task, in general, the results are satisfactory for a first approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>since in other instances of humor, this model has demonstrated a certain level of effectiveness.
Therefore, this time we aim to test if it can really be used for different cases and if its use can be
generalized.</p>
      <p>
        An important aspect of language-based models, especially BLOOMZ &amp; mT5 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is they are
optimized to follow the next token predicted, in this case, following a prompt with instructions oriented
to the translation task. It is believed that such models could be useful to maintain the wordplay in the
translation. Although this type of tasks can become a little complicated since certain terms used in
different languages can make it difficult how to perceive the wordplay of the original sentence to the
translated one. This issue is later encountered, as these challenges can normally arise within a model.
      </p>
      <p>The use of homophones can complicate the task even more, as there may not be a similar word in
the translation that matches or has the same intention, altering in a certain way the interpretation of the
joke with the pun, thus changing the original meaning. It could be said that the pun is maintained, but
not the meaning. With the understanding that most puns become partial translations of what the original
sentence proposes, we can say that there are different phenomena present in the translation of puns
whether they are homophones, homographs, or this comes from a regional homophony of the sentence.</p>
      <p>The task of classification presents a different challenge because puns can be confused by those
sentences that unintentionally exhibit similar characteristics to a pun. Understanding the type of
sentences that our model needs to screen can give us a broader idea of what needs to be done. In this
case, the classification of puns appears to include certain sentences that emulate puns, but in reality,
they are not. This can hinder and confuse the model, as it introduces a certain amount of noise with
which we must deal.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Approach to the task</title>
      <p>
        The tasks proposed by JOKER in CLEF 2023 primarily focus on the automatic analysis of puns. In
this edition, a corpus consisting not only of English and French text but also of Spanish was available.
Based on previous successful works [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], it was decided to use the transformers approach [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Due to
time constraints, of the three available tasks, only Task 1, involving wordplay detection, and Task 3,
focusing on translation, were executed. It should be noted that tasks were performed solely on Spanish
and English texts from the training corpus provided by JOKER.
      </p>
      <p>
        The first task focused on pun detection, which can be solved as a binary classification task since it
has yes and no labels that indicate whether a text is a pun or not. Initially, it was considered to train a
multilingual model such as multilingual BERT [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which is pre-trained in 104 languages. However,
due to the low performance obtained during evaluation, it was decided to keep the model but apply a
separate approach.
      </p>
      <p>
        As a result, two models were trained: one specifically for English and the other for a mixture of
English-Spanish, both utilizing the same BERT architecture [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In the end, the models were trained
with BERT and multilingual BERT. For this task, the model was loaded from the Hugging Face
transformers module with the help of the Ktrain [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] wrapper, which facilitated the process of loading
and fine-tuning the model in a fast and simple manner.
      </p>
      <p>Due to the fact that transformers do not require extensive preprocessing, the training process was
relatively straightforward. However, during the fine-tuning phase, it was necessary to experiment with
different parameters to achieve the best results in validation accuracy and loss. In the end, the best
models were saved in a Keras H5 format for future reference.</p>
      <p>For the training of the BERT-like models, the following parameters were utilized: a batch size of 8,
3 training epochs, and a learning rate of 7e-6 for the BERT model trained on English data. On the other
hand, the parameters for BERT-multilingual trained on English-Spanish data were a batch size of 32, 3
training epochs, and a learning rate of 5e-5. Considering the text length from the JOKER corpus, a
maximum length of 70 tokens was declared.</p>
      <p>To develop Task 3, consisting of the translation of word plays from English to Spanish, BLOOMZ
&amp; mT5 was implemented, which is a more robust version of BLOOM, offering an improvement in task
performance. This model can follow human instructions through the use of a prompt such as the
examples used in the experimentation: “Translate into Spanish:”, “Translate the text into Spanish:”,
“Translate into Spanish the next text:”, and “Translate into Spanish the next sentence:”. It is worth
mentioning that the result obtained depends on the prompt used. In this case, four passes or runs were
made on the text to achieve a better result, each pass corresponding to one of the provided prompts.</p>
      <p>These iterations were necessary because in some passes the results were not favorable, either because
the model fail to produce a translation (resulting in blank output), generating incomplete translation, or
because the model produced more text than necessary (generating new text beyond the necessary
translation).</p>
      <p>
        The queries were conducted using the Hugging Face API, specifically the Inference API [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which
facilitates the execution of queries using machine learning models. As mentioned, multiple iterations
were performed due to the results obtained in each run, some sets of wordplays were either not translated
or did not yield a proper result. It was required to try different prompts to get the best possible result,
yet some puns failed to generate a translation, resulting in a blank query even after multiple iterations.
      </p>
      <p>Although the generation of the translation was fully automated, a sample of the results were
manually reviewed to identify potential errors. These errors were taken into consideration to determine
which texts needed a re-runs. For instance, in the case of the pun “What do you call a doctor who treats
retired soldiers? A veteran-arian.” the first two attempts did not produce any results. It was in the third
run that the text was successfully translated to “¿Qué es un doctor que trata a soldados retirados? Un
veterano-ario.”. Another example is the sentence “Recent survey revealed 6 out of 7 dwarf's aren't
happy.” which was generated on the second pass: “Una reciente encuesta reveló que 6 de cada 7 enanos
no son felices.”</p>
    </sec>
    <sec id="sec-3">
      <title>3. Resources employed</title>
      <p>Among the resources used to train and evaluate the models, Google's Colab environment was
employed. This platform enables Python programming and execution, while also providing easier use
of GPU. The use of GPU resources allowed for faster execution in the performance of the described
tasks, by enabling multiple simultaneous computations. The server used has the following
specifications: GPU NVIDIA-SMI 525.85.12, CUDA v12.0, and 25 of RAM.</p>
      <p>
        As part of the resources, it was decided to use a combination of Google Colaboratory and Petals [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
for Task 3 translation, since we did not have the computational capacity to handle models with
parameter size 10B, which requires a significant amount of computation beyond our capabilities. The
use of Petals becomes indispensable for those jobs that do not have enough computational capacity, as
it leverages the benefits of crowd computing to make inferences with large models in a relatively
easy/semi-efficient way. The platform offers the flexibility of an API and the power of PyTorch, for
seamless integration in the required tasks.
      </p>
      <p>
        On the other hand, for other BERT-like models that are less demanding in the use of resources, it is
necessary to consider the data volume and the desired width of tokenization. In the experimentation
process BETO [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], BERT, and BERT-multilingual models were tested for the classification task,
likewise, the training set was taken under three different approaches:
1. Spanish-only dataset.
2. English-only dataset.
      </p>
      <p>3. And a mixture of English-Spanish data.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>As can be seen in the tables below, the probabilities of each of the tasks and a brief explanation of
the cases are provided. Some of the sentences demonstrate a clear example of a pun, while others are
not so straightforward in terms of their intended meaning. This variability in pun structures contributes
to the challenge of accurately identifying puns. It is expected that the model may have some confusion
in finding the characteristics between sentences that lack a pun’s connotation but exhibit high similarity
to those that do. This type of concern seems to be overlooked by the model to some extent. It is logical
to evaluate the data set to get some insights into how the model works, which is exactly what is
presented in the following tables.</p>
    </sec>
    <sec id="sec-5">
      <title>Task 1: Detection of puns</title>
    </sec>
    <sec id="sec-6">
      <title>English corpus</title>
      <p>In this section, the results obtained for the binary classification of English and Spanish puns are
presented. As the test set labels were not provided, some inferences about the wordplays are made to
explain the model performance when evaluating this type of sentence.</p>
      <p>The BERT model trained on the English corpus consisting of 5,292 wordplays, was divided into a
training and validation set in a 80:20 distribution. Subsequently, the model was used to evaluate the test
set provided by JOKER, which comprised 3,183 sentences. The probability of a text containing a
wordplay was obtained, and the top 5 wordplays and non-wordplay were identified.</p>
      <p>In Table 1, it can be seen that a positive result indicating a wordplay involves a specific structure,
where a sentence contains references in pairs (the words indicating this event are marked in bold letters).
The highest results obtained a probability of around 0.99, where three of the puns with the highest
probability of being wordplay refer to a pun referring to lawyers, as an example:</p>
      <p>The lawyer asked a loaded question about guns.</p>
      <p>The pun talks about a loaded question, and it is easily understood that the play on the words lies in
the relationship between “loaded” and “guns”. This example provides a small idea of how the model
classifies such positive instances.</p>
      <p>Pun</p>
      <sec id="sec-6-1">
        <title>She was only a lawyer's daughter, but what a will to break.</title>
      </sec>
      <sec id="sec-6-2">
        <title>She was only an Attorney's daughter, but what a will to break.</title>
      </sec>
      <sec id="sec-6-3">
        <title>The lawyer asked a loaded question about guns.</title>
      </sec>
      <sec id="sec-6-4">
        <title>Once ice cream was invented the problem was licked.</title>
      </sec>
      <sec id="sec-6-5">
        <title>Our Boy Scouts' knot - tying class went off without a hitch.</title>
        <p>Regarding the results with a low probability of being puns, that indicate a non-wordplay sentence.
It can be observed that, unlike the previous table, these examples are not as obvious in terms of finding
the wordplay. Therefore, if even for humans it is challenging to identify this word matching, it can be
even more difficult for the model to understand the nuances that these sentences have with the use of
puns. The results in Table 2 show how these sentences do not contain wordplay, indicating that the
model correctly classified them as non-wordplay.</p>
        <p>Pun</p>
      </sec>
      <sec id="sec-6-6">
        <title>What's the name of that street in Paris? asked Tom quietly.</title>
      </sec>
      <sec id="sec-6-7">
        <title>I wouldn't marry you if you were the only woman on earth, said Tom quietly.</title>
      </sec>
      <sec id="sec-6-8">
        <title>Give me some pre - packed cheese slices, said Tom professionally.</title>
        <p>''Let's take a vacation in the south of France,'' said Tom loudly.</p>
      </sec>
      <sec id="sec-6-9">
        <title>Can you read music? the bandleader asked calmly.</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Spanish corpus</title>
      <p>The multilingual BERT model used to classify the Spanish test set (2,241 sentences) presents distinct
curiosities compared to the problems encounter in the English dataset. This may be attributed to the fact
that incorporating a mixture of datasets strengthened the model's ability to infer the type of wordplays,
but more so the type of redundant wordplays that was able to correctly classify. The example below
shows a sentence that was accurately classified as a wordplay:</p>
      <p>¿Sabes cuánto pago de alquiler por la frutería? - No, ¿cuánto? - Pimientos euros.</p>
      <p>In this instance, it is observed how the model was able to infer alquiler (rent), and connected in to
the word “euros” which corresponds to the word Pimientos (bell peppers). The intended meaning is to
convey Pimientos to quinientos (five hundred) in relation to a monetary amount, a very similar word to
the word pimientos. This wordplay is easy to understand, unlike other types of puns that are presented
in Table 3, like the sentence:</p>
      <p>Hey Jesús. - ¿Salió cara la cena? No, salió mala.</p>
      <p>To understand this pun, relate to religion, humans typically require prior knowledge about the
subject. It can be challenging to identify, as it often involves cultural references. However, it is among
the results with the highest probability of being a wordplay, which is correctly classified. As seen in the
table, the confidence in the probability is lower than that of the English model, ranging between 0.80
percent in the probability that it will be a wordplay.</p>
      <p>Pun</p>
      <sec id="sec-7-1">
        <title>Los jugadores de baloncesto, famosos por su mal carácter, pillan muchos rebotes</title>
        <p>¿Sabes cuánto pago de alquiler por la frutería? - No, ¿cuánto? - Pimientos euros.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Ejecutivo agresivo busca monedas antiguas para partirles la cara.</title>
      </sec>
      <sec id="sec-7-3">
        <title>Ejecutivo agresivo busca monedas antiguas para partirles la marca.</title>
      </sec>
      <sec id="sec-7-4">
        <title>Hey Jesús. - ¿Salió cara la cena? No, salió mala.</title>
        <p>In the case of the examples with lower probability, it can be noted that the presence of sentence
structures where the word order is altered can make it more difficult to find the wordplay. This is evident
in the example provided in Table 4:</p>
        <p>Un Señor Ruiz... que un ruiseñor.</p>
        <p>The use of Señor (Mr.) and then the surname Ruiz may create a pun that refers to a type of bird, the
ruiseñor (nightingale bird), but the model was unable to recognize the relationship between these two
statements, where only the order of the pun was altered. The other four sentences do not contain a pun,
this shows that most of the non-wordplay text where correctly identify as such.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Task 3: Translation of puns</title>
      <p>Translations posed an extremely interesting case due to the fact that translating a term that makes a
play of words in English may be difficult to interpret in Spanish. For a translation to fulfill the same
objective as its counterpart in Spanish, it is extremely important to pay attention to certain structures.
For example, the use of proper names that refer to a particular characteristic can be difficult to
understand when they are translated into Spanish.</p>
      <p>Understanding these nuances often becomes the pivotal part that modifies the behavior of what one
wants to translate but that on other occasions can play in favor of the translation by making sense in the
other language. Without forgetting that normally these words are quite related to the internal structure
of the word in both languages.</p>
      <p>These observations can be seen in Table 5, where different results are presented. The first sentence
“If there's one person you don't want to interrupt in the middle of a sentence it's a judge.” although
correctly translated to “Si hay una persona que no quieres interrumpir en medio de una oración, es un
juez.” the pun is not maintained, resulting in a literal translation.</p>
      <p>On the contrary, in the case of the sentence “I tried to learn how to drive a stick shift but couldn't
locate the manual.” that was correctly translated to “Intenté aprender a conducir con una caja de
cambios manual pero no pude encontrar el manual.”, maintained not only the sense, but also the
funniness of the wordplay (as indicated in bold letters).</p>
      <sec id="sec-8-1">
        <title>The bride's best friend is so proud, she's practically made of honor.</title>
      </sec>
      <sec id="sec-8-2">
        <title>La mejor amiga de la novia está tan orgullosa que parece que fuera la madrina.</title>
      </sec>
      <sec id="sec-8-3">
        <title>Some rappers are good but others are</title>
      </sec>
      <sec id="sec-8-4">
        <title>Ludacris.</title>
      </sec>
      <sec id="sec-8-5">
        <title>When the fog burns off it won't be mist.</title>
      </sec>
      <sec id="sec-8-6">
        <title>Algunos raperos son buenos, pero otros son</title>
      </sec>
      <sec id="sec-8-7">
        <title>Ludacris.</title>
      </sec>
      <sec id="sec-8-8">
        <title>Cuando se disipe la niebla, no será vapor.</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>5. Conclusions</title>
      <p>In conclusion, we can say that the approaches implemented to cope with the classification and
translation of text with puns is not misguided, perhaps more work could be done to strengthen and
further explore and improve the results in these tasks. Especially in the case of translation, as the aspect
of translating while maintaining the original context of the language to another makes this type of task
a very interesting one to delve into, and it calls for continued research to find better methods and
techniques.</p>
      <p>Moreover, the integration of more advanced language models, such as BLOOMZ &amp; mT5,
showcased the potential of the use of prompts to enhance the translation of puns, which are instructions
carried out by humans but this itself influences the result obtained. Regarding the classification of
wordplays, the utilization of wordplays with BERT-like showed good results but they could be
improved with better fine-tuning and exploring different variations of BERT-like architectures. Overall,
this study obtained good results that demonstrate the potential of transformer-based models and
language models in the classification and translation of puns.</p>
    </sec>
    <sec id="sec-10">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Liana</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Tristan</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <surname>Anne-Gwenn</surname>
            <given-names>Bosser</given-names>
          </string-name>
          , Victor Manuel Palma Preciado, Grigori Sidorov, and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Jatowt</surname>
          </string-name>
          .
          <year>2023</year>
          .
          <article-title>Overview of JOKER - CLEF-2023 track on Automatic Wordplay Analysis</article-title>
          . In Avi Arampatzis, Evangelos Kanoulas, Theodora Tsikrika, Stefanos Vrochidis, Anastasia Giachanou,
          <string-name>
            <given-names>Dan</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Aliannejadi</surname>
          </string-name>
          , Michalis Vlachos, Guglielmo Faggioli, Nicola Ferro (Eds.)
          <string-name>
            <surname>Experimental IR Meets Multilinguality</surname>
          </string-name>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF</source>
          <year>2023</year>
          )
          <article-title>J</article-title>
          . Cohen (Ed.), Special issue: Digital Libraries, volume
          <volume>39</volume>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Muennighoff</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutawika</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biderman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scao</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bari</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yong</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schoelkopf</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aji</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almubarak</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albanie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alyafeai</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Webson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raff</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Raffel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Crosslingual Generalization through Multitask Finetuning</article-title>
          . arXiv (Cornell University). https://doi.org/10.48550/arxiv.2211.01786
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Attention is All you Need</article-title>
          . En arXiv (Cornell University) (Vol.
          <volume>30</volume>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          ). Cornell University. https://arxiv.org/pdf/1706.03762v5
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Mahurkar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Patil</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>LRG at SemEval-2020 Task 7: Assessing the Ability of BERT and Derivative Models to Perform Short-Edits Based Humor Grading</article-title>
          . https://doi.org/10.18653/v1/
          <year>2020</year>
          .semeval-
          <volume>1</volume>
          .
          <fpage>108</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Victor</given-names>
            <surname>Manuel Palma Preciado</surname>
          </string-name>
          , Grigori Sidorov,
          <article-title>Carolina Palma Preciado Assessing WordplayPun classification from JOKER dataset with pretrained BERT humorous models</article-title>
          ,
          <source>JokeR: Automatic Wordplay and Humour Translation</source>
          , pages (
          <issue>1828-1833</issue>
          ), CLEF (
          <year>2022</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , M.-W.,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . CoRR, abs/
          <year>1810</year>
          .04805. Retrieved from http://arxiv.org/abs/
          <year>1810</year>
          .04805
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Pires</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlinger</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Garrette</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2019</year>
          ). How Multilingual is Multilingual BERT? https://doi.org/10.18653/v1/p19-
          <fpage>1493</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Arun</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Maiya</surname>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>ktrain: A Low-Code Library for Augmented Machine Learning</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .10703. BigScience Workshop. (
          <year>2022</year>
          ).
          <source>BLOOM (Revision 4ab0472)</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Inference</surname>
            <given-names>API - Hugging</given-names>
          </string-name>
          <string-name>
            <surname>Face</surname>
            . (s. f.). https://huggingface.co/inference-apiCañete,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Pérez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Spanish PreTrained BERT Model and Evaluation Data</article-title>
          .
          <source>In PML4DC at ICLR</source>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Borzunov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baranchuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dettmers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ryabinin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belkada</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chumachenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samygin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Raffel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Petals: Collaborative Inference and Fine-tuning of Large Models</article-title>
          . arXiv (Cornell University). https://doi.org/10.48550/arxiv.2209.01188
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Cañete</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>J.-H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Pérez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>SPANISH PRETRAINED BERT MODEL AND EVALUATION DATA</article-title>
          . workshop paper at PML4DC,
          <string-name>
            <surname>ICLR</surname>
          </string-name>
          <year>2020</year>
          . https://users.dcc.uchile.cl/~jperez/papers/pml4dc2020.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>