<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Sixth Workshop on Natural Language for Artificial Intelligence, November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Evaluating Text-To-Text Framework for Topic and Style Classification of Italian texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Papucci</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara De Nigris</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Miaschi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Università di Pisa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>TALIA S.r.l.</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Istituto di Linguistica Computazionale "A. Zampolli" (ILC-CNR), ItaliaNLP Lab</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>30</volume>
      <issue>2022</issue>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>In this paper, we propose an extensive evaluation of the first text-to-text Italian Neural Language Model (NLM), IT5 [1], on a classification scenario. In particular, we test the performance of IT5 on several tasks involving both the classification of the topic and the style of a set of Italian posts. We assess the model in two diferent configurations, single- and multi-task classification, and we compare it with a more traditional NLM based on the Transformer architecture (i.e. BERT). Moreover, we test its performance in a few-shot learning scenario. We also perform a qualitative investigation on the impact of label representations in modeling the classification of the IT5 model. Results show that IT5 could achieve good results, although generally lower than the BERT model. Nevertheless, we observe a significant performance improvement of the Text-to-text model in a multi-task classification scenario. Finally, we found that altering the representation of the labels mainly impacts the classification of the topic.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;transformers</kwd>
        <kwd>text-to-text</kwd>
        <kwd>t5</kwd>
        <kwd>bert</kwd>
        <kwd>topic classification</kwd>
        <kwd>style classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and Motivation</title>
      <p>
        Over the past few years, the text-to-text paradigm has become one of the most widely adopted
approach in the development of state-of-the-art Neural Language Models (NLMs) [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ]. The
basic idea of this paradigm, inspired by previous unifying frameworks for NLP tasks [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ], is
to consider each task as a text-to-text task, i.e. getting text as input data and producing new
text as output.
      </p>
      <p>
        This unifying framework has proven to be a particularly efective transfer learning method,
often outperforming previous models, e.g. BERT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], in data-poor settings. Nevertheless, few
works proposed systematic evaluations of such models in diferent classification scenarios
and in comparison with more traditional NLMs. Among these, [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] showed that T5 achieves
comparable, if not better performance, with previous state-of-the-art models on the most
popular NLP benchmarks, e.g. GLUE [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and SQuAD [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], instead, demonstrated that T5
      </p>
      <p>Values
Five ranges: 0-19, 20-29, 30-39, 40-49 and 50-100
M, F
Eleven possible categories: ANIME, AUTO-MOTO,
BIKES, CELEBRITIES, ENTERTAINMENT, NATURE,
MEDICINE-AESTHETIC, METAL-DETECTING,</p>
      <p>
        SMOKE, SPORTS, TECHNOLOGY
outperforms BERT in a document ranking task, especially in a data-poor setting with limited
training data. Inspecting the performance of 6 diferent NLMs on a sentiment analysis task, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
found that T5 is the second best performing model, next only to XLNet [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Whereas, focusing on languages other than English, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] compared the performance of their
IT5 with other multilingual and Italian models, showing e.g. that IT5 base outperforms BERT on
SQuAD-IT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], the extractive question answering task for the Italian language. Similar results
have been obtained by [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] while measuring the performance of their Brazilian Portuguese T5
model (PTT5) against the ones obtained with BERTimbau, a BERT model pre-trained on the
brWaC corpus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Comparing the models on three diferent evaluation tasks for the Portuguese
language (i.e. semantic similarity and entailment prediction [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and NER [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]), they showed
that PTT5 achieves competitive performance with BERTimbau, although the latter obtained
slightly better results.
      </p>
      <p>
        Building on these previous studies, in this work we propose an evaluation of the first
text-totext Transformer model developed for the Italian language, IT5 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], on several classification
tasks. More specifically, we performed our experiments on two diferent classification scenarios,
single-task and multi-task, and we compared the performance of IT5 against those obtained
with an Italian version of BERT. Furthermore, in order to verify the ability of the model in
a data-poor setting, we also tested its performance in a few-shot learning scenario. Finally,
following the approach devised by [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], we performed a more in-depth analysis to test the
impact of label representations in modeling the classification of the IT5 model.
      </p>
      <p>The remainder of the paper is organized as follows: in Sec. 2 and 3 we introduce the dataset
and the models used in our experiments, in Sec. 4 we describe the experimental setting, in Sec.
5 and 6 we discuss the obtained results and in Sec. 7 we conclude the paper.</p>
      <p>Contributions: In this paper we: i) proposed an extensive evaluation of IT5 performance on
three diferent classification tasks based on Italian sentences; ii) we tested the performance of
the model in diferent scenarios (single- and multi-task classification) and we compared them
with those obtained with another Transformer especially suited for classification tasks; iii)
we studied the behavior of the model in a data-poor setting by measuring its performance in
few-shot learning scenario; iv) we verified the impact of label modification on IT5’ performance.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data</title>
      <p>
        In order to perform our experiments, we relied on posts extracted from TAG-IT [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], the profiling
shared task presented at EVALITA 2020 [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The dataset, based on the corpus defined in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
consists of more than 10,000 posts written in Italian and collected from diferent blogs. Each
post is labeled with three diferent labels: age and gender of the writer and topic. The details
and the statistics about the dataset are reported in Table 1 and Figures 1, 2 and 3.
      </p>
      <p>As it can be noticed from the Figures, the Age variable presents a quite balanced distribution
among the five classes, especially for the three intervals between 30 and 100. For what concerns
the Gender attribute, we can observe that the majority of posts were written by male users,
thus determining a strongly unbalanced distribution of the two classes. The last variable, Topic,
presents 11 labels, with 3 of them (ANIME, SPORTS and AUTO-MOTO) having more than 2,500
posts each.</p>
      <p>
        In order to have enough data to fine-tune our pre-trained models, we decided to modify the
original task as defined in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Instead of predicting the three labels of a given collection of
texts (multiple posts), we fine-tuned our models to predict age, gender and topic from each
single post. Moreover, since a fair amount of sentences were quite short, we decided to remove
those shorter than 10 tokens. At the end of this process, we obtained a dataset consisting of
13553 posts as training set and 5055 posts as test set.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Models</title>
      <p>In what follows, we discuss more in detail the characteristics of the models used in our
experiments.</p>
      <p>
        IT5 We used the T5 base version pre-trained on the Italian language [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]1. In particular, the
model was trained on the Italian sentences extracted from a cleaned version of the mC4 corpus
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], a multilingual version of the C4 corpus including 107 languages. As discussed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
in order to compare diferent architectures (e.g. T5 and BERT), it would be ideal to analyze
models with meaningful similarities, e.g. having a similar number of parameters or amount of
computation to process an input-output sequence. Since T5 with n layers has approximately the
same number of parameters as a BERT with 2 layers but also the same amount of computational
cost of an -layers BERT, in order to achieve the fairest comparison of the two Transformers,
we decided to use the base version of IT5 (220M parameters).
      </p>
      <p>
        BERT In order to compare the performance of IT5 with that of another Transformer model
generically used in classification scenarios, we relied on a pre-trained Italian BERT. Specifically,
we used the base cased BERT (12 layers, 110M parameters) developed by the MDZ Digital
Library Team, available trough the Huggingface’s Transformers library [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]2. The model was
trained using Wikipedia and the OPUS corpus [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setting</title>
      <p>As we already introduced in Sec. 1, we performed our experiments on two diferent classification
scenarios: i) single-task and ii) multi-task classification. For what concerns the single-task
scenario, we both fine-tuned BERT and IT5 three times in order to create three diferent
singletask sequence classification models, one for each variable. To perform fine-tuning with the
BERT model, we converted the three target variables into numeric labels. On the other hand,
the target variables were verbalized empirically as follows for the IT5 model:
• Gender: values have been transformed in uomo and donna;
• Topic: values have been translated in Italian, written in lowercase and truncated into a
single word (e.g. MEDICINE-AESTHETIC into medicina), thus resulting in the following list:
anime, automobilismo, bici, sport, natura, metalli, medicina, celebrità, fumo, intrattenimento,
tecnologia;
• Age: values have been left unchanged.</p>
      <p>
        Moreover, following the Fixed-prompt LM tuning approach (see [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] for an overview), we added a
prefix to each input when fine-tuning the IT5 model. This approach implies providing a textual
template that is then applied to every training and test example. Fixed-prompt LM tuning
has been already successfully explored for text classification, allowing more eficient learning
[
        <xref ref-type="bibr" rid="ref28 ref29 ref30">28, 29, 30</xref>
        ]. In our experiments, we tested three diferent prefixes, one for each classification
task: "Classifica argomento" , "Classifica età" and "Classifica genere" .
      </p>
      <p>Concerning instead the multi-task classification, each sentence has been presented three
times during the training phase of the two models, each one with the appropriate label and, in
the case of IT5, with the appropriate prefix.
1https://huggingface.co/gsarti/it5-base
2https://huggingface.co/dbmdz/bert-base-italian-xxl-cased
Dummy (S)
Dummy (MF)
BERT Random
IT5 Random
BERT
IT5
MT BERT
MT IT5</p>
      <p>Macro
0.50
0.44
0.56
0.36
0.76
0.31
0.33
0.23
0.75
0.33</p>
      <p>Few-Shot Learning In order to evaluate the performance of IT5 also in a context with little
data available, we decided to carry out our classification experiments in a few-shot learning
scenario. Specifically, we divided the original dataset into 5 equal subsets (1/5 = 2,710, 2/5 =
5,420, 3/5 = 8,130, 4/5 = 10,840, 5/5 = 13,554) and then we monitored the performance trend of
both IT5 and BERT at increasing intervals of data samples: 0/5, 1/5, 2/5, 3/5, 4/5 and 5/5 of the
TAG-IT dataset.
4.1. Baseline and Evaluation
We relied on two diferent typology of models as baseline. The first one is based on two dummy
classifiers: i) most frequent classifier (Dummy (MF)), which always predict the most frequent
label for each input sequence and ii) stratified dummy classifier (Dummy(S)), that generates
predictions by respecting the class distribution of the training data. Moreover, in order to assess
the impact of the pre-training phase of the two Transformer models, we also used a BERT Italian
(BERT Random) and an IT5 model (IT5 Random) with randomly initialized weights.</p>
      <p>We used F-Score (macro and weighted) as evaluation metric for all the experiments.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>
        Single-task Classification results are reported in Table 2. As we can observe, Transformer
models outperformed the dummy baselines in almost all the classification tasks. The only
exception concerns the performance of IT5 on the Age prediction task, for which the stratified
dummy classifier obtained the same scores. It should be considered that the Age classification
task appears to be the most complex task, regardless of the model taken into account. In
fact, the best performing model (BERT) obtained only .11 points more than the baseline. The
complexity in predicting the age ranges could be due to the fact that the task requires more
sophisticated information rather than the simple identification of textual clues. On the other
hand, the classifiers that achieved best results are those trained to predict the gender and the
topic of each post. This result is in line with [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], where the authors suggested that textual
clues seem to be more indicative of these dimensions than age. Moreover, the higher scores
obtained for the gender classification task could also be indicative of the fact that, diferently
from the other two, gender prediction was cast as a binary task.
      </p>
      <p>When we look at the performance obtained by the randomly initialized BERT and IT5, we
note that the latter achieved results close to those of the pre-trained models. Indeed in some
cases, e.g. IT5 on the Age and Gender prediction tasks, the Random model gets better results.
This seems to suggest that the pre-training phase of IT5 did not allow the model to encode
enough useful information in order to improve its performance on the selected tasks. On
the other hand, the pre-training phase had a strong impact on BERT performance, since the
pre-trained model outperformed the Random one in all classification tasks.</p>
      <p>
        If we focus instead on the diferences between the two models, we can clearly notice that
BERT performed best in all three configurations. In particular, IT5 achieved fairly reasonable
results in comparison with BERT for simpler tasks, such as Gender and Topic classification.
For what concerns the Age prediction task instead, we observed a performance drop, with a
diference in terms of weighted F-Score of .17 points. A possible explanation for this behavior
could be due to the fact that, diferently from BERT, T5 has to produce the label by generating
open text, thus making the prediction more complex from a computational point of view. In
this regard, it is important to notice that for our experiments we relied on the base version of
IT5, which, despite being bigger in terms of parameters than BERT base, is still quite smaller
than the best-performing model (T5-11B) presented in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Moreover, it should be pointed out
that in some cases IT5 generated labels that did not belong to those defined in Sec. 4, but which
actually turned out to be more accurate than the original ones. This is the case, for instance, of
a few posts labelled with fumo (en. smoke) that were predicted instead by IT5 with the label
tabacco (en. tobacco). We will inspect more in detail this behavior in Sec. 6. We also found that
sometimes IT5 was not able to generate meaningful labels, but rather produced only punctuation
marks or single letters. Nevertheless, we only identified a few isolated cases of them (less than
5 for what concerns Topic classification), which had no real impact on the overall performance
of the model. We would like to also point out that the IT5 Random model does not generate
unexpected labels like the pre-trained one does. This could be another motivation for its better
performance in the two cases of Age and Gender classification.
      </p>
      <p>Multi-task Observing the results obtained in the multi-task setting, we notice a significant
increase in the performance of IT5. In fact, while BERT achieved a consistent boost only in
the Topic prediction scenario, T5 performances improve significantly in all classification tasks,
with an average improvement of around .06 points more (in terms of weighted F-Score) than
during single-task classification. This is particularly evident with regard to Topic and Age
classification, while the scores obtained for the Gender prediction task remain roughly the same.
This result could suggest that, besides having more data for the fine-tuning phase, the IT5 model
particularly benefits from learning multiple tasks at a time, thus improving its generalization
abilities.</p>
      <p>Few-Shot Learning Figures 4, 5 and 6 report the results obtained with the few-shot learning
classification scenario. As we can see, the trend is quite diferent between the two models. In
fact, while BERT performance shows a fairly regular increase across the 5 fractions of the dataset,
in the IT5 model we observe a quite constant improvement only for the Age prediction task.
Interestingly, for what concerns Topic and Gender classification, IT5 makes correct predictions
only after being exposed to 4/5 of the entire dataset. This behavior appears to be in line with
what we already noticed during multi-task classification, namely that having more data available
for the fine-tuning phase allows the model to perform better, and consequently, to obtain results
closer to those of BERT. This seems to be further suggested by the fact that, unlike IT5, BERT
obtains strong performance already from the early portions of the datasets but then it tends to
remain quite stable, showing an improvement of only a few points in the remaining portions.
This is especially the case of the Gender prediction task, where the accuracy of the BERT model
in predicting the correct labels is roughly the same (.84 in terms of weighted F-Score) even after
seeing 2/5 of the original dataset. Nevertheless, in the case of zero-shot learning, both models
are unable to correctly classify the posts occurring in the test set of the three datasets.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Label Analysis</title>
      <p>
        As described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], one of the issues of the Text-to-text framework applied in a classification
scenario is that the model could outputs text on a task that does not correspond to any of the
possible labels. However, as we already observed in the previous section, in some cases it seems
that IT5 was able to generate more appropriate labels that those originally defined for the task,
thus suggesting generalization abilities. For instance, as we can observe from the examples
in Table 3, the labels predicted for the three input posts are not among those expected for the
Topic prediction task. Nevertheless, by looking at the posts, the labels predicted by T5 might be
Model
IT5
IT5 shufled
considered more appropriate predictions.
      </p>
      <p>Inspired by such behavior, we decided to further investigate the generalization abilities of the
IT5 model by measuring the impact of diferent labels on model performance. More specifically,
we decided to produce a shufled version of each dataset by randomly replacing the labels with
each other. Results are reported in Table 4. As we can see, the most significant variations in
model performance concern the Topic and Age classification tasks. In particular, we can observe
a drastic performance drop for what concerns Topic, with a diference between the predictions
on correct and shufled labels of more than .24 points in terms of Weighted F-Score. Moreover,
it is interesting to note that the scores obtained with the shufled labels are also lower than
those obtained by the randomly initialized IT5 (0.17 vs. 0.34). This result seems to suggest that
the IT5 model is indeed able to learn some specific lexical correlations between the encoding
of the input tokens and of the labels during the fine-tuning phase and that these correlations
are no longer observable after the shufling process. This is also corroborated by the fact that,
when presented with shufled data, the model stopped generating new and more specific labels
for the input sequences.</p>
      <p>If we look instead at the results obtained with the Gender dataset, we can notice that shufling
the labels does not have a significant efect on the performance of the model. This is a clear
evidence that, unlike Topic, the Gender prediction task does not present a direct lexical connection
between the input sequence and the label. As a result, the model tends to memorize the
information available in the fine-tuning data rather than derive generalities exploiting the
knowledge learned during the pre-training phase.</p>
      <p>
        Finally, inspired by the work of [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], we conducted further analysis on the efect of the strings
used to represent labels on model performance. In particular, we decided to replace the labels
used for the Gender prediction task (i.e. uomo and donna) with the original tags defined in the
TAG-IT dataset, i.e. m and f. As shown in Table 5, modifying the label representation did not
afect the performance of IT5, which obtained basically the same results in both configurations.
This seems to confirm once again that for tasks that do not show an explicit relationship between
input samples and labels, the choice of the label largely does not afect model performance.
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>In this paper, we proposed an extensive evaluation of the first Italian text-to-text model, IT5,
on diferent classification tasks based on Italian sentences. Specifically, we chose to exploit
the TAG-it dataset in order to measure the performance of the model in diferent classification
scenarios.</p>
      <p>First, we evaluated IT5 in a high data setting, assessing its performance during single- and
multi-task classification and comparing them with the ones obtained by fine-tuning an Italian
version of BERT. Results showed that IT5 is able to achieve quite good results, especially in Topic
and Gender classification, and that its performance increases significantly when fine-tuned in a
multi-task manner. Nevertheless, we found that BERT outperformed IT5 in all classification
tasks.</p>
      <p>Next, we tested the model in a poor data setting by measuring its performance in a few-shot
learning scenario. Once again, IT5 achieved lower scores with respect to BERT, which obtained
satisfactory results even in a context with very few data available (e.g. 1/5 of the entire dataset).
A possible explanation of these results could be that given the high complexity of predicting the
correct label by generating open text, it may be necessary to employ bigger text-to-text models
to outperform models that are explicitly designed for solving classification tasks. Regardless
of the classification scenario, we noticed that, especially for the Topic prediction task, IT5
occasionally generated labels that were not among those defined in the TAG-it dataset and that
such labels often proved to be more indicative of the topic than the original ones. This result
suggested that the model is indeed able to identify lexical clues indicative of the topic although
in some cases it does not associate them with the labels that were originally defined for the task.</p>
      <p>Finally, we investigated the impact of modifying the classification labels on IT5 performance.
In particular, by shufling at random the values of the original labels, we found that the model
achieved generally lower scores and this is especially true for the classification of the topic.
Nevertheless, experimenting with the Gender prediction task, we found that the choice of label
representation does not afect significantly the model performance.
and the 11th International Joint Conference on Natural Language Processing (Volume 1:
Long Papers), Association for Computational Linguistics, Online, 2021, pp. 3816–3830. URL:
https://aclanthology.org/2021.acl-long.295. doi:10.18653/v1/2021.acl-long.295.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sarti</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Nissim, It5: Large-scale text-to-text pretraining for italian language understanding and generation</article-title>
          ,
          <source>ArXiv preprint 2203.03759</source>
          (
          <year>2022</year>
          ). URL: https://arxiv.org/abs/ 2203.03759.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Passaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Preface to the
          <source>Sixth Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          , in: D.
          <string-name>
            <surname>Nozza</surname>
            ,
            <given-names>L. C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          , M. Polignano (Eds.),
          <source>Proceedings of the Sixth Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2022</year>
          )
          <article-title>co-located with 21th International Conference of the Italian Association for Artificial Intelligence (AI*IA</article-title>
          <year>2022</year>
          ), November 30,
          <year>2022</year>
          , CEUR-WS.org,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          .,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Webson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sutawika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Alyafeai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chafin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stiegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Le</given-names>
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raja</surname>
          </string-name>
          , et al.,
          <article-title>Multitask prompted training enables zero-shot task generalization</article-title>
          ,
          <source>in: The Tenth International Conference on Learning Representations</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Aribandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schuster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. V.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. Q.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ni</surname>
          </string-name>
          , et al.,
          <article-title>Ext5: Towards extreme multi-task scaling for transfer learning</article-title>
          ,
          <source>in: International Conference on Learning Representations</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>McCann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Keskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <article-title>The natural language decathlon: Multitask learning as question answering</article-title>
          , arXiv preprint arXiv:
          <year>1806</year>
          .
          <volume>08730</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Keskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McCann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <article-title>Unifying question answering, text classification, and regression via span extraction</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>09286</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Language models are unsupervised multitask learners (????).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          ,
          <article-title>Glue: A multi-task benchmark and analysis platform for natural language understanding</article-title>
          ,
          <source>in: 7th International Conference on Learning Representations, ICLR</source>
          <year>2019</year>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rajpurkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lopyrev</surname>
          </string-name>
          , P. Liang, SQuAD:
          <volume>100</volume>
          ,000+
          <article-title>questions for machine comprehension of text</article-title>
          ,
          <source>in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Austin, Texas,
          <year>2016</year>
          , pp.
          <fpage>2383</fpage>
          -
          <lpage>2392</lpage>
          . URL: https://aclanthology.org/D16-1264. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D16</fpage>
          -1264.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Nogueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pradeep</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Document ranking with a pretrained sequence-tosequence model, in: Findings of the Association for Computational Linguistics: EMNLP 2020, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>708</fpage>
          -
          <lpage>718</lpage>
          . URL: https: //aclanthology.org/
          <year>2020</year>
          .findings-emnlp.
          <volume>63</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .findings-emnlp.
          <volume>63</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Pipalia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bhadja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shukla</surname>
          </string-name>
          ,
          <article-title>Comparative analysis of diferent transformer based architectures used in sentiment analysis</article-title>
          ,
          <source>in: 2020 9th International Conference System Modeling and Advancement in Research Trends (SMART)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>411</fpage>
          -
          <lpage>415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , J. Carbonell,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Xlnet: Generalized autoregressive pretraining for language understanding</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Croce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zelenanska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Basili</surname>
          </string-name>
          ,
          <article-title>Neural learning for question answering in italian</article-title>
          ,
          <source>in: International Conference of the Italian Association for Artificial Intelligence</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>389</fpage>
          -
          <lpage>402</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Carmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Piau</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Campiotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nogueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lotufo</surname>
          </string-name>
          ,
          <article-title>Ptt5: Pretraining and validating the t5 model on brazilian portuguese data</article-title>
          , arXiv preprint arXiv:
          <year>2008</year>
          .
          <volume>09144</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Wagner</surname>
          </string-name>
          <string-name>
            <surname>Filho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wilkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Idiart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Villavicencio</surname>
          </string-name>
          ,
          <article-title>The brWaC corpus: A new open resource for Brazilian Portuguese</article-title>
          ,
          <source>in: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ),
          <article-title>European Language Resources Association (ELRA), Miyazaki</article-title>
          , Japan,
          <year>2018</year>
          . URL: https://aclanthology.org/L18-1686.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.</given-names>
            <surname>Real</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fonseca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Gonçalo</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <article-title>The assin 2 shared task: a quick overview</article-title>
          ,
          <source>in: International Conference on Computational Processing of the Portuguese Language</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>406</fpage>
          -
          <lpage>412</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Seco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Vilela</surname>
          </string-name>
          ,
          <string-name>
            <surname>Harem:</surname>
          </string-name>
          <article-title>An advanced ner evaluation contest for portuguese</article-title>
          , in: quot; In Nicoletta Calzolari; Khalid Choukri; Aldo Gangemi; Bente Maegaard; Joseph Mariani; Jan Odjik; Daniel Tapias (ed)
          <source>Proceedings of the 5 th International Conference on Language Resources and Evaluation</source>
          (LREC'
          <year>2006</year>
          )
          <article-title>(Genoa Italy 22-</article-title>
          28 May
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Label representations in modeling classification as text generation, in: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th</article-title>
          <source>International Joint Conference on Natural Language Processing: Student Research Workshop</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Cimino</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Orletta</surname>
          </string-name>
          , Nissim, Tag-it - topic, age and gender prediction,
          <source>EVALITA</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Di</given-names>
            <surname>Maro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Croce</surname>
          </string-name>
          , L. Passaro,
          <year>Evalita 2020</year>
          :
          <article-title>Overview of the 7th evaluation campaign of natural language processing and speech tools for italian, in: 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . Final Workshop,
          <string-name>
            <surname>EVALITA</surname>
          </string-name>
          <year>2020</year>
          , volume
          <volume>2765</volume>
          , CEUR-ws,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maslennikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Labruna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cimino</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Dell'Orletta, Quanti anni hai? age identification for italian</article-title>
          ., in: CLiC-it,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Xue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Constant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siddhant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barua</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Rafel, mT5: A massively multilingual pre-trained text-to-text transformer, in: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>483</fpage>
          -
          <lpage>498</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .naacl-main.
          <volume>41</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .naacl-main.
          <volume>41</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delangue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cistac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Louf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Funtowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Davison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shleifer</surname>
          </string-name>
          , P. von Platen, C. Ma,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jernite</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Plu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Le</given-names>
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gugger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Drame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Lhoest</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rush</surname>
          </string-name>
          , Transformers:
          <article-title>State-of-the-art natural language processing</article-title>
          ,
          <source>in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          . URL: https://www.aclweb.org/anthology/2020.emnlp-demos.6. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .emnlp-demos.
          <volume>6</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nygaard</surname>
          </string-name>
          ,
          <article-title>The OPUS corpus - parallel</article-title>
          and free: http://logos.uio.no/opus, in
          <source>: Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC'04)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Lisbon, Portugal,
          <year>2004</year>
          . URL: http://www.lrec-conf.org/proceedings/lrec2004/pdf/320.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hayashi</surname>
          </string-name>
          , G. Neubig,
          <article-title>Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing</article-title>
          ,
          <source>arXiv preprint arXiv:2107.13586</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>T.</given-names>
            <surname>Schick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schütze</surname>
          </string-name>
          ,
          <article-title>Few-shot text generation with pattern-exploiting training</article-title>
          , arXiv preprint arXiv:
          <year>2012</year>
          .
          <volume>11926</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Schick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schütze</surname>
          </string-name>
          ,
          <article-title>Exploiting cloze-questions for few-shot text classification and natural language inference</article-title>
          ,
          <source>in: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics:</source>
          Main Volume,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>255</fpage>
          -
          <lpage>269</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .eacl-main.
          <volume>20</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .eacl-main.
          <volume>20</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Making pre-trained language models better few-shot learners</article-title>
          ,
          <source>in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>