<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CLiC-it</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Challenging specialized transformers on zero-shot classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>SerenaAuriemma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MauroMadeddu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MartinaMilian</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AlessandoBondiell</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AlessandroLenci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>LuciaPassar o</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Commons License Attribution 4.0 International</institution>
          ,
          <addr-line>CC BY 4.0</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dipartimento di Filologia, Letteratura e Linguistica, Università di Pisa</institution>
          ,
          <addr-line>Via Santa Maria, Pisa, 56126</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dipartimento di Informatica, Università di Pisa</institution>
          ,
          <addr-line>Largo B. Pontecorvo, 3 Pisa, 56127</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Domain Adaptation</institution>
          ,
          <addr-line>Transformers, Prompting, Zero-shot, Italian Bureaucratic Language, Public Administration</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>9</volume>
      <abstract>
        <p>industry. This paper investigates the feasibility of employing basic prompting systems for domain-specific language models. The study focuses on bureaucratic language and uses the recently introduced BureauBERTo model for experimentation. The experiments reveal that while further pre-trained models exhibit reduced robustness concerning general knowledge, they display greater adaptability in modeling domain-specific tasks, even under a zero-shot paradigm. This demonstrates the potential of leveraging simple prompting systems in specialized contexts, providing valuable insights both for research and curate in the fill mask task5][,1 where the model had to predict both random and in-domain masked words, we</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>wide variety of downstream tas1k, s2,[3, inter alia].</p>
      <p>Pre-trained Language Models (PLMs) have had a signwifi-anted to further inspect the domain lexical knowledge
cant impact on Natural Language Processing (NLP), aancdquired by this model during the domain adaptation.
the pre-train and fine-tune paradigm has become the prWe-e aimed at leveraging this knowledge to implement
dominant approach for applying efective models ontawo classification tasks in the PA domain, modeled as
prompt-based classification. Thus, we challenged the
efective with Large Language Models
(LLMs4)].[How</p>
      <p>However, one of the main concerns when workinmgodel to predict both the topics of PA texts, and the type
with PLMs is the paucity of annotated data, especialloyffgoerneric and PA-related named entities occurring in
specific domains, required to fine-tune the additional classe-ntences extracted from administrative documents.
sification layer on top of these models for downstreamWe conducted two prompting experiments for each
tasks, such as classification. Recently, prompt-based tutna-sk. We first adopted the Italian name of the
classificaing has started to afirm as a promising way to perfortmion classes as label words, then we associated in-domain
similar tasks, significantly reducing the need for antneor-ms to each class. We also compared BureauBERTo
tated data. This approach has been proven to be vwerityh an Italian generic PLM, UmBERTo (Sect3i)o. n</p>
      <p>Our findings show that in a zero-shot classification
exploiting the prompt-based tuning technique.
ever, it is often the case that LLMs are not availablesfcoenrario when the label words of each class are shallowly
low-resource languages, and that their performancerderlaast-ed to the content of the text or to the entity type fed
tically decreases when they are challenged on spetcoificthe model in the prompt template, both the generic
domains. Hence, we decided to test a domain-specificand the domain-specialized models perform poorly in the
model, BureauBERTo5[], a LM further pre-trained ocnlassification task. However, when the classes are
repreItalian bureaucratic texts (e.g., administrative acts,sbeanntke-d by multiple word labels semantically related to the
ing and insurance documents), in a zero-shot scenatreixot/entity to be classified, the PLMs improve their
performance by a wide margin. This gaining is particularly
evi</p>
      <p>Since BureauBERTo has shown to be particularlydaecn-t in the domain-adapted model BureauBERTo, which
0009-0006-6846-5826 (S. Auriemma); 0009-0002-7844-3963
CEUR
htp:/ceur-ws.org
ISN1613-073
© 2023 Copyright for this paper by its authors. Use permitted under Creative</p>
      <p>CEUR</p>
      <p>Workshop ProceedingsC(EUR-WS.org)
0000-0003-3426-6643 (A. Bondielli)0;000-0001-5790-4308 (A. Lenci); compared to when the same task is accomplished via
outperformed UmBERTo in both prompt-document
classification and prompt-entity typing tasks, suggesting that
the domain linguistic knowledge acquired by this model
during the additional pre-training phase could be
particularly useful in a prompt-based tuning scenario where
the model is much more reliant on its word knowledge,</p>
      <p>1See AppendixB for the plot of the model results in the fill-mask
task</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related work</title>
      <p>specific terms, to investigate how domain-related word
labels afect the model’s performance in diferent
classification tasks.</p>
      <p>PLMs have proven to be efective in NLP tasks re3-. Models
lated to specific domains, whether they were trained
from scratch6,[ 7], or further pre-trained on domaFinor our experiments, we decided to compare the
perdata 8[, 9, 10] with a Masked Language Modeling (MLMf).ormance of two PLMs, namely UmBERTo and
BuMore recently, the MLM training objective has been lerveearu-BERTo. UmBERTo2 is a RoBERTa-based language
aged to solve various NLP tasks reformulated as amsordtel trained on the Italian section of the OSCAR
corof cloze task, allowing the PLM to directly solve itpwuist,3ht-hat has been shown to perform well on
adminout any or with very few labelled examples. Oneisotfrative dat2a3][ compared to other generic PLMs of
the first works in this direction was proposed1b1y],[ the same size (110M parameters). BureauBER4T[o5] is a
who performed zero-shot learning using pre-trained LdMomsain-adapted model obtained by further pre-training
without fine-tuning on a dataset of training exampUlemsB.ERTo on Italian PA, banking, and insurance
docuWithin similar conditions, but using the larger GPmTe-n3,ts.
[4] achieved near state-of-the-art results for some
SuperGLUE [12] tasks. [13] showed that competitive
perfor</p>
      <p>4. Experimental settings
mance with those of GPT-3 can be achieved with much
smaller models like the 220M parameters ALBERT, by
performing some gradient-based fine-tuning of the mod4el.1. Prompting
using the labeled examples on a cloze task. Since thPerno,mpt-based classification requires a specific
temprompt-based learning has gained attention as a simppllaete to reformulate the original classification task as a
way to perform, among other tasks, zero-shot classicficlao- ze-task, where the text to be classified is fed to the
tion 1[4]. However, it’s essential to note that the permfoodre-l followed by a prompt sentence, such as “This
mance of prompt-based learning techniques scales w&lt;ittehxt&gt; is about [MASK]”. In this way, the model has
model size [15]. Consequently, general purpose Largteo predict the probability that a certain word is filled
Language Models (LLMs) with billions of parameteinrsthe “[MASK]” token. The mapping from the label
are typically used in prompt-learning experiments, evceanndidate word to a specific class is gained through the
for specialized domains such as the legal 1o6n]e. I[n verbalizer [13], which represents the original class names
contrast, for the biomedical and clinical dom1a7]ins,a[s a set of label words, greatly influencing the model
pershowed that smaller specialized models like BioB8E]RT f[ormance in the tas2k4[]. Hence, we decided to conduct
and Clinical BERT18[] outperform GPT-2 and T5 in aour prompt-based classification experiments in two
setfew-shot prompt based classification of medical texttisn.gs, using a standard and a cusvteormbalizer to better
The authors hypothesize that the advantage of the BuEnRdTe-rstand the correlation between the lexical
knowlbased models is possibly due both to their domain adeadpg-e of PLMs and the use of domain-related terms as the
tation and to their bidirectional MLM training objescettivoef,word labels in the promveprtbalizer.
which is more similar to the prompt template formaTthe first verbalizer, i.e., thbease-verbalizer simply uses
than those of auto-regressive and sequence-to-sequetnhceeItalian name of the classification classes as label words
models like GPT-2 and T5.1[9] reported a similar finding (e.g., Ambiente - “Environment” is the label word for the
even for the much larger GPT-3 over BioBERT. Nevertchlea-ssAmbiente - “Environment”), while the second
less, these approaches are constrained by the model invpeurtbalizer is maanual verbalizer that we constructed by
size, which limits the length of the conditioning inadpudting some synonyms of the class name and some
recontext and can significantly afect performan1c9e].[ lated PA terms as label words for each class, to better</p>
      <p>Although prompt-based classification with specializdeedpict the document classes and the entity types (in this
models has been explored for the medical and clincicaasle the label words for the cAlamsbsiente -
“Environdomains, to the best of our knowledge, this is the firmsetnt” are: ambiente - “environment”, natura - “nature”,
work that focuses on applying prompts to the Italianterardit-orio - “territory”, flora - “flora”,etc. ).
ministrative language and in a zero-shot classification
scenario. Additionally, a notable challenge in
promptbased approaches lies in their sensitivity to variations
in prompt templates and verbalize2r0s, 2[1, 22]. We
conducted experiments using diferent verbalizers, i.e., a 2https://github.com/musixmatchresearch/umberto
generic verbalizer and a custom verbalizer using domain3-https://oscar-corpus.com
4https://huggingface.co/colinglab/BureauBERTo
4.2. Datesets classification task as a masked language modeling
problem: &lt;text&gt;.Questo documento parla di &lt;mask&gt;.7
We evaluate the models in two tasks on two diferentThus, PLMs are challenged to infer the topic of the
datasets. For thperompt document classification , document by predicting the most appropriate label word
we used a subset of the ATTO corpu2s3][, which is a to represent the masked token in the prompt, following
collection of administrative documents annotatedtwheitdhocument text. Since the ATTO corpus contains
labels denoting topics. We filtered this dataset keeping</p>
      <p>only short documents of a maximum of 600 tokens, by
only those instances (2,811) that were annotated wsitethtaing the tokenizer’s truncation at 5128t,owkeenwsere
single topic label. able to feed the models the entire document in almost all</p>
      <p>For theprompt entity typing task, we used the PA-cases. Like with the prompt entity typing, we perform the
corpus of 2[5], a collection of 460 PA-documents with</p>
      <p>prompt-based classification twice. In the first experiment,
token-level annotations of Named Entities denoting botuhsed thebase verbalizer, where each class is linked to
we
general entities, such as persons, locations, organizatoinoenso,r few label words that correspond to the names of
and domain-specific entities, like legislative norms, acts,</p>
      <p>the classes in the original annotation of the ATTO corpus.
and PA-related organizations. For the second experiment, we use tmhaenual verbalizer,
which contains, in addition to the label wordsboafsethe
4.3. Evaluation metrics verbalizer, a collection of domain terms manually selected
We evaluated the performance of the models with caosmP-A representative topic labels for each class. The
complete list of the label words used in both verbalizers
mon classification metrics. is shown in Table1.9
4.4. Prompt entity typing</p>
    </sec>
    <sec id="sec-4">
      <title>5. Results and discussion</title>
      <p>We modeled the NER task introduced2b5y][as an entity
typing task. Entity typing can be considered a subtTaasbkle2 shows the results of prompting applied to the
of NER and focuses on entity classification. In otheenrtity typing task.
words, systems assign a label to an already extractIendthe first experiment, where a single class label is
entity. This task is often formulated to challenge systeumsesd (see Sec.4.4), UmBERTo almost doubled the results
at retrieving sub-categories organized in a hierarcohbitcaailned by BureauBERTo for F1 Micro (0.404 vs. 0.263)
structure (e.g., an entity corresponding to a personamnady Macro Average (0.335 vs. 0.201). Surprisingly, for
be specified as director, major, lawyer, etc.) As 2in5][, a domain entity likAeCT, BureauBERTo missed all the
we asked models to identify only coarse-grained entietnietsi:ties, whereas UmBERTo obtained a low but higher
generic ones, such as personPsE(R), locationsL(OC), and score (0.140). For thLeAW entity, UmBERTo overpasses
organizationsO(RG); and related to the administratBivuereauBERTo, as well. We may suppose that this is due to
domain: law referenceLsA(W), administrative actAsC(T), the fact that UmBERTo was trained on Common-Crawl,
and PA organizationOsP(A). which also contains legal and administrative texts in its</p>
      <p>We prompted the models by giving as input a sentenIctealian section. Very high results are obtained by
Umand an entity occurring in it, asking to predict the eBnEtRitTyo forPER entities, reaching 0.827 in our zero-shot
type in place of a masking token. The resulting tescme-nario. On the contrary, both models obtain very low
plate is:&lt;text&gt;. In questa frase, &lt;entity&gt; è un results foLrOC, OPA, andORG. These two latter classes
esempio di &lt;mask&gt;.5 are very similar to each other: ORG refers to
organi</p>
      <p>As anticipated, we verbalized the entities in two wzaaytsi.ons in general, comprising firms and associations,
In the first experiment, we provided an Italian translatwiohnereasOPA can be considered as a subclassOoRfG,
of the entity or a single word representing the entityacnldarsse.fers to organizations within the Public
AdministraIn the second experiment, we expanded most of the labteioln, such as municipal departments. Such overlapping
words by including synonyms and other terms relatedmatyo impact on classification.
the various class6es. For what concerns the second experiment, we added
to the prompt also highly distinctive words for each class.
4.5. Prompt document classification In this case, we notice a better ability of BureauBERTo
to recognize domain-specific entities suchAaCsT, LAW,
For the recognition of the topics in PA documents, we
designed the following template to model the document
7In English:&lt;text&gt;. This document is about &lt;mask&gt;.</p>
      <p>8512 is the maximum number of tokens that these Transformers
5In English:&lt;text&gt;. In this sentence, &lt;entity&gt; is an models can receive as input.
example of &lt;mask&gt;. 9See AppendixA for the English translation of the label words
6Both verbalizers for entity typing are in AppAen.dix for document classification.</p>
      <p>Basic Labels +In-domain Lexicon
ambiente, natura, territorio, flora, fauna, animali, clima, inquinamento, rifiuti, igiene,
caccia, pesca, verde, ecologia, agricoltura, acque
avvocatura, avvocati, giustizia, legale, ricorso, giudici, Tribunale, Corte, Appello,</p>
      <p>Assise, notifica, atti, Albo, Pretorio, protocollo
Bandi-Contratti</p>
      <p>bandi, contratti, bando, contratto, gara, appalto, assunzione, liquidazione
commercio, economia, attività, economica, beni, commerciare, vendite, acquisti,
commercianti, confesercenti
cultura, turismo, sport, culturale, turisti, musei, arte, cinema, vacanze, spettacolo,
scuola, manifestazioni
demografico, popolazione, abitanti, residenti, censimento, anagrafe, residenza,
domicilio, cittadinanza, leva
edilizia, costruzioni, cantiere, ristrutturazione, planimetrie, residenziale
personale, risorse, umane, assunzioni, lavoro, part-time
servizi, informazioni, informativi
finanza, euro, finanziario, contabilità, contabile, copertura, rimborsi, pagamenti,
versamenti, bilancio, spese, sanzioni, multe, tributi, retribuzioni, emolumenti
sociale, leva, militare, disabili, protezione, civile, invalidi
Pubblica-Istruzione</p>
      <p>istruzione, istituto, scolatisco, scuola, insegnante, formazione, educazione
urbanistica</p>
      <p>urbanistica, trasporti, trasporto, trafico, circolazione, veicoli, viabilità, viaria
and OPA. However, despite the general improvemenatdded for thPeER entity class, i.eg.eneralità -
“particuin recognizing such classes, we notice that it perfolarrms”s andnominativo - “name”.
worse than UmBERTo for traditional entities. This eTxh- e results in Tabl3eshow that the performance of
periment based on the comparison of general-purpoUsmeBERTo increases not only for tPhEeR entities but
language models and domain-adapted ones has yieldtehdat the ablation improves the F1-score oOfRtGhcelass
compelling insights. Generally, both types of modaeslswell. Whereas UmBERTo reaches the highest
perdemonstrate enhanced performance when enriched wfiotrhmances for overall F1 Micro Avg, the deletion of
indomain-specific terms within their prompts. However, idtomain lexicon from the verbalizer seems to penalize
is evident that the domain-adapted model outperfoBrumrseauBERTo in the recognitionPoEfR entities.
Folthe general-purpose model, exhibiting an improvemelnotwing the trend observed in UmBERTo, the ablation
of more than twofold (0.516 vs 0.368 for Macro Averaimgpeacts the model’s ability to properly recognize the
F1 score). This significant boost in performance sugo-ther classes. Despite this, the adapted model still
obgests that the domain-adapted model is likely to be mtoariened higher results on the in-domain entity classes:
attuned and proficient in leveraging domain-specific teAr-CT, LAW, andOPA further solidifying the advantages of
minology. domain-adapted models in specialized contexts. Finally,</p>
      <p>Nevertheless, it is important to acknowledge tithaistworth noting that we observed a high variability
domain-specific terms may wield less influence over of results according to diferent prompts and verbalizer
generic entities such PaEsR. With the in-domain lex-configurations, as shown in the ablation study. In fact,
icon added to the verbalizer, UmBERTo fails to recogdneizleting the in-domain lexicon related to one of the entity
anyPER entity. By looking at the confusion matrixcfloarsses afected the performance achieved by the models
UmBERTo, we observed that the model identifies almosotn all the others, due to wrong classifications (e.g.,
peoall the people’s names aOsRG entities. Thus, we carriedple names confused with location addresses or company
out an ablation study by deleting the in-domain tenrammses). Therefore, future investigations into prompt
tuning are necessary and can lead to further interetsytpiensgoccurring in administrative sentences.
insights. We compared the results obtained in these two tasks</p>
      <p>Regarding the prompt document classification experbiy- the PA-specialized model BureauBERTo with those
ments, whose results are summarized in ta4,bwlee ob- of the domain-agnostic model UmBERTo. Our findings
served a similar trend. When only one or few wosrhdow that by enriching with domain terms the set of
labels are used to represent a topic class, both the gewnoeridclabels encoded in the promveprtbalizer both
modand the domain-specialized models obtained a ratherelloswdemonstrated enhanced performances. Moreover,
accuracy (0.22 vs. 0.09) and Macro Average F1 scoreBsureauBERTo exhibited an improvement over UmBERTo
(0.16 vs. 0.06). In this case, UmBERTo outperformed Buo-f +0.06 Weighted Average F1 score in the document
clasreauBERTo in almost all classes, with the exceptiosnificoaftion (0.51 vs. 0.57) and of more than twofold in the
Cultura, Turismo e Sport - ‘Culture, tourism, and entity typing task (0.516 vs. 0.368 for Macro Average F1
sports’, Demografico - ‘Demographics’, andPerson- score), meaning that the domain adapted model is more
ale - ‘Personnel’. Looking into the details of the scorpersoficient in leveraging domain-specific terminology.
obtained by UmBERTo in its most recognizable classesThese results underscore the importance of tailoring
( Pubblica istruzione - ‘Public Education’, Edilizia language models to specific domains to unlock their full
- ‘Constructions’ andUrbanistica - ‘Urban plan- potential and address the nuanced challenges posed by
ning’), we speculate that the single-word labels usedivtoerse subject matters. However, it is also worth
mendefine these classes provided a suficient cue to enable thteioning that we noticed a high variability in the task
model to appropriately recognize these topics. This irseisnults according to diferent prompting and diferent
line with the fact that the UmBERTo pre-training colrapbuesl words. In particular, when the label words adopted
included texts extracted from Italian municipalitiest’owdebepict a certain topic class are, within the domain
conpages, which often refer to such topics. text, semantically related to the label words of another</p>
      <p>On the other hand, in the second experiment, whecrlaess, the models’ classification output seems to be biased
we manually added to the promvpetrbalizer a set of in favor of one of the two classes.
salient PA-related terms to depict the document topIincsconclusion, our study underscores the critical need
at a finer-grained level, we observed a significant imf-or a thorough exploration of prompt engineering,
parprovement in the overall performance of both modteiclsu.larly in the context of the entity typing task. This
The benefits of a custom-made set of domain-relateimdperative arises not only from the potential to augment
terms are particularly evident for the specialized mtohdeeplredictive capabilities of models, but also from the
BureuBERTo, which reached a better accuracy (0.60nvese.d to consolidate the knowledge related to general
en0.54) and Weighted Average F1-score (0.57 vs. 0.51) thatnity classes. Notably, the Public Administration (PA)
doUmBERTo. It appears that the model adapted to themdaoi-n exhibits distinctive characteristics, both in terms of
main may possess heightened sensitivity, enabling itrteoferencing entity names within documents and
employefectively capitalize on the contextual cues ofered binyg domain-specific terminology. Notably, the identified
domain-specific terms. However, by performing a classp-atterns within the PA domain deviate from the broader,
wise comparison between the two experimental settinggens,eral-purpose Italian style, indicating the necessity for
we observed that for some classes that shared a commtaoinlored, domain-specific prompt experimentation.
domain lexicon, such aPsubblica Istruzione - ‘Public This investigative efort shed the linguistic intricacies
Education’ andCultura, Turismo e Sport - ‘Culture, that exert an impact on Transformer model performance.
Tourism, and Sports’, orServizi finanziari - ‘Finan- Our findings, as revealed in the ablation study on entity
cial services’ andBandi e Contratti - ‘Tenders and linking, emphasize the pivotal importance of delving into
Contracts’ the models’ classification could have beenthe interplay among diferent entity classes present in
influenced in favor of one of the two classes, due to thdeaitrasets. A nuanced analysis of how these classes interact
topic descriptor lexical overlap. These findings confirmand potentially overlap is indispensable for honing the
the necessity of further inquiry into the efect of lexmiocdaell’s ability to distinguish between them in a
domainspecificity on prompt-based classifications, especially fosrpecific context.
domain-adapted models. To conclude, this leads us to surmise as a future
direction for our work a further inspection of how
domainadapted PLMs encode in their embedding the semantics
6. Conclusion and future work of domain-related terms and how this information relates
to their performance in prompt-based tasks.</p>
      <p>In this paper, we propose a zero-shot prompt tuning
classification approach for solving two tasks related to
the Italian PA domain: the classification of documents
according to their topic and the recognition of the entity</p>
      <p>P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>
        0.455
0.618
0.524
Associates, Inc., 2019. URLh: ttps://proceedings.ne[24] T. Gao, A. Fisch, D. Chen, Making pre-trained
lanu
        <xref ref-type="bibr" rid="ref6">rips.cc/paper_files/paper/2019</xref>
        /file/4496bf24afe7f guage models better few-shot learners, in:
Proceedab6f046bf4923da8de6-Paper.p.df ings of the 59th Annual Meeting of the
Associa[13] T. Schick, H. Schütze, It’s not just size that mat- tion for Computational Linguistics and the 11th
Inters: Small language models are also few-shot ternational Joint Conference on Natural Language
learners, in: Proceedin
        <xref ref-type="bibr" rid="ref9">gs of the 2021</xref>
        Confer- Processing (Volume 1: Long Papers), Association
ence of the North American Chapter of the As- for Computational Lin
        <xref ref-type="bibr" rid="ref9">guistics, Online, 2021</xref>
        , pp.
sociation for Computational Linguistics: Human 3816–3830. URL: https://aclantholo
        <xref ref-type="bibr" rid="ref9">gy.org/2021</xref>
        .ac
Language Technologies, Association for Compu- l-long.29.5doi:10.18653/v1/2021.acl- long.29
tational Lin
        <xref ref-type="bibr" rid="ref9">guistics, Online, 2021</xref>
        , pp. 2339–2352. 5.
      </p>
      <p>
        URL: https://aclantholo
        <xref ref-type="bibr" rid="ref9">gy.org/2021</xref>
        .naacl-mai.n.1[8255] L. C. Passaro, A. Lenci, A. Gabbolini, Informed
doi:10.18653/v1/2021.naacl-main.185. PA: A NER for the italian public administration
[14] R. Puri, B. Catanzaro, Zero-shot text classifica- domain, in: R. B. andMalvina Nissim, G. Satta (Eds.),
tion with generative language models, Comput- Proceedings of the Fourth Italian Conference on
ing Research Repository (CoRR) abs/1912.10165 Computational Linguistics (CLiC-it 2017),
        <xref ref-type="bibr" rid="ref6">Rome,
(2019</xref>
        ). URL: http://arxiv.org/abs/1912.101.65 Italy, December 11-13, 2017, volume 2006 oCfEUR
arXiv:1912.10165. Workshop Proceedings, CEUR-WS.org, 2017. URL:
[15] B. Lester, R. Al-Rfou, N. Constant, The power of https://ceur-ws.org/Vol-2006/paper048..pdf
scale for parameter-eficient prompt tunin
        <xref ref-type="bibr" rid="ref9">g, in:
Proceedings of the 2021</xref>
        Conference on Empirical
Methods in Natural Lan
        <xref ref-type="bibr" rid="ref9">guage Processing, 2021</xref>
        , App.. Label Words
3045–3059.
[16] F. Yu, L. Quartey, F. Schilder, Legal promptinTga:ble5 shows the verbalizer for entity typing. T6able
Teaching a language model to think like a lawcyoern,tains the English version of the verbalizer adopted
arXiv preprint arXiv:2212.01326 (2022). for the document classification (see Ta1bfloer the Italian
[17] S. Sivarajkumar, Y. Wang, Healthprompt: A zervoe-rsion).
      </p>
      <p>shot learning paradigm for clinical natural language
processing., in: AMIA... Annual Symposium
proceedings. AMIA Symposium, volume 2022, 2022, pp.</p>
      <p>972–981.
[18] E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng,</p>
      <p>D. Jin, T. Naumann, M. McDermott, Publicly
available clinical bert embeddings, arXiv preprint
arXiv:1904.03323 (2019).
[19] M. Moradi, K. Blagec, F. Haberl, M. Samwald, Gpt-3
models are poor few-shot learners in the biomedical
domain, arXiv preprint arXiv:2109.02555 (2021).
[20] E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang,</p>
      <p>L. Wang, W. Chen, et al., Lora: Low-rank
adaptation of large language models, in: International</p>
      <p>
        Conference on Learnin
        <xref ref-type="bibr" rid="ref9">g Representations, 2021</xref>
        .
[21] Y. Lu, M. Bartolo, A. Moore, S. Riedel, P. Stenetorp,
      </p>
      <p>Fantastically ordered prompts and where to find
them: Overcoming few-shot prompt order
sensitivity, in: Proceedings of the 60th Annual Meeting
of the Association for Computational Linguistics
(Volume 1: Long Papers), 2022, pp. 8086–8098.
[22] D. Trautmann, A. Petrova, F. Schilder, Legal prompt
engineering for multilingual legal judgement
prediction, arXiv preprint arXiv:2212.02199 (2022).
[23] S. Auriemma, M. Miliani, A. Bondielli, L. C.
Passaro, A. Lenci, Evaluating pre-trained transformers
on italian administrative texts, in: Proceedings of
1st Workshop AIxPA (co-located with AIxIA 2022),
2022.
persona (person), generalità (particulars), nominativo (name)
luogo (place), località (locality)
organizzazione (organization), azienda (firm), società (corporation),
associazione (association), compagnia (company)
legge (law), norma (rule), decreto (decree), legislativo (legislative)
atto (act), delibera (resolution), determina (decision), deliberazione
(deliberation), regolamento (regulation)
uficio (ofice)
environment, nature, land, flora, fauna, animals, climate,
pollution, waste, hygiene, hunting, fishing, green, ecology,
agriculture, water
advocacy, attorneys, justice, legal, appeal, judges,
courthouse, court, appello, assise, notification, acts, albo,
pretorio, protocol
tenders, contracts, notice, contract, tender, hiring,
liquidation
trade, economy, business, economic, goods, trade, sales,
purchases, merchants, confesercenti
culture, tourism, sports, cultural, tourists, museums, art,
cinema, vacations, entertainment, school, events
demographics, population, inhabitants, residents, census,
registry, residence, domicile, citizenship, conscription
building, construction, yard, renovation, planimetry,
residential
personnel, resources, human, hiring, work, part-time
education, institute, school, teacher, training, education
services, information, informative
finance, euro, financial, accounting, accountant, coverage,
refunds, payments, disbursements, budget, expenses,
penalties, fines, taxes, wages, emoluments
welfare, conscription, military, disabled, protection,
civilian, disability
urban planning, transportation, transports, trafic,
circulation, vehicles, roadway</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Processing Systems</source>
          , volume
          <volume>30</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Inc.</surname>
          </string-name>
          ,
          <year>2017</year>
          . URL: https://proceedings.neurips.
          <article-title>cc/pap This research has been funded by the Project “ABI2LE er_files/paper/2017/file/3f5ee243547dee91fbd053c (Ability to Learning)”, Regione Toscana (POR Fesr 2014- 1c4a845aa-Paper.pd</article-title>
          .f
          <year>2020</year>
          )
          <article-title>; by PNRR - M4C2 - Investimento 1.3</article-title>
          ,
          <issue>Partenariato</issue>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Howard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          ,
          <article-title>Universal language model fineEsteso PE00000013 -</article-title>
          “
          <string-name>
            <surname>FAIR - Future Artificial</surname>
          </string-name>
          <article-title>Intelligence tuning for text classification</article-title>
          ,
          <source>in: Proceedings of the Research” - Spoke</source>
          <volume>1</volume>
          “
          <string-name>
            <surname>Human-centered</surname>
            <given-names>AI</given-names>
          </string-name>
          ”,
          <article-title>funded by 56th Annual Meeting of the Association for Computhe European Commission under the NextGeneration tational Linguistics (Volume 1: Long Papers), AssoEU programme; and partially supported by: TAILOR, a ciation for Computational Linguistics, Melbourne, project funded by EU Horizon 2020 research and innova-</article-title>
          <source>Australia</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>328</fpage>
          -
          <lpage>339</lpage>
          . URLh: ttps://aclantho tion programme
          <source>under GA No 952215. logy.org/P18-103.1doi:10</source>
          .18653/v1/
          <fpage>P18</fpage>
          -1031.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>References Pre-training of deep bidirectional transformers for</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>language understanding</article-title>
          ,
          <source>in: Proceedings of the [1]</source>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          , 2019 Conference of the North American Chap-
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>tention is all you need</article-title>
          , in: I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          , tics:
          <source>Human Language Technologies</source>
          , Volume 1
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Garnett</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information tational Linguistics</source>
          , Minneapolis, Minnesota,
          <year>2019</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . URL: https://aclanthology.org/N19
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          -
          <fpage>1423</fpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1423. [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tinn</surname>
          </string-name>
          , H. Cheng, M. Lucas,
          <string-name>
            <given-names>N.</given-names>
            <surname>Usuyama</surname>
          </string-name>
          , [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D. X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Poon</surname>
          </string-name>
          , Domain-
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Henighan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Ramesh,
          <article-title>on Computing for Healthcare (HEALTH) 3 (</article-title>
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hesse</surname>
          </string-name>
          , M. Chen,
          <volume>1</volume>
          -
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Sigler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Litwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chess</surname>
          </string-name>
          , J. Clark,[8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>So</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          ers, in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hadsell</surname>
          </string-name>
          , Bioinformatics
          <volume>36</volume>
          (
          <year>2020</year>
          )
          <fpage>1234</fpage>
          -
          <lpage>1240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin (Eds.), Advances in Neural In-[9]
          <string-name>
            <given-names>I.</given-names>
            <surname>Chalkidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fergadiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Malakasiotis</surname>
          </string-name>
          , N. Ale-
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>formation Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>tras</given-names>
          </string-name>
          , I. Androutsopoulos,
          <article-title>Legal-bert: The mup-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Associates</surname>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
          <article-title>URLh:ttps: pets straight out of law school</article-title>
          , arXiv preprint
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          //proceedings.neurips.cc/paper_files/paper/2020/fi arXiv:
          <year>2010</year>
          .
          <volume>02559</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>le/</surname>
            1457c0d6bfcb4967418bfb8ac142f64a-Paper..pdf [10]
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Licari</surname>
          </string-name>
          , G. Comandè,
          <article-title>Italian-legal-bert: A pre[5</article-title>
          ]
          <string-name>
            <given-names>S.</given-names>
            <surname>Auriemma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Madeddu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Miliani</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Bondielli, trained transformer language model for italian law,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>UmBERTo to the Italian bureaucratic language</article-title>
          , in: edge Management for Law Workshop (KM4LAW),
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Falchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giannotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Boldrini</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Rinzivillo</surname>
          </string-name>
          , S. Colantonio (Eds.),
          <source>Proceedings[o1f1]</source>
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>shops co-located with the 3rd CINI National Lab multitask learners (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>AIIS Conference on Artificial Intelligence (Ital</source>
          [I1A2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Pruksachatkun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nangia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <year>2023</year>
          ), volume
          <volume>3486</volume>
          ofCEUR Workshop Proceedings,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bowman</surname>
          </string-name>
          , Super-
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          .org, Pisa, Italy,
          <year>2023</year>
          , pp.
          <fpage>240</fpage>
          -
          <lpage>248</lpage>
          . URL:
          <article-title>glue: A stickier benchmark for general-purpose</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3486</volume>
          /42. p.
          <article-title>df language understanding systems</article-title>
          , in: H. Wal[6]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <article-title>Scibert: A pretrained lach</article-title>
          , H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
          </string-name>
          , F.
          <string-name>
            <surname>d'</surname>
          </string-name>
          Alché-Buc,
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          arXiv:
          <year>1903</year>
          .
          <volume>10676</volume>
          (
          <year>2019</year>
          ).
          <source>formation Processing Systems</source>
          , volume
          <volume>32</volume>
          , Curran
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>