<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>GFG - Gender-Fair Generation: A CALAMITA Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simona Frenda</string-name>
          <email>s.frenda@hw.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AndreaPiergenti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>BeatriceSavoldi</string-name>
          <email>savoldi@fbk.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MarcoMadeddu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MartinaRosola</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia Casola</string-name>
          <email>s.casola@lmu.de</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Ferrand o</string-name>
          <email>chiara.ferrando@unito</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>VivianaPatt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MatteoNegri</string-name>
          <email>negri@fbk.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luisa Bentivogl</string-name>
          <email>bentivo@fbk.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, University of Turin</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Interaction Lab, Heriot-Watt University</institution>
          ,
          <addr-line>Edinburgh, Scotland</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>MaiNLP &amp; MCML, LMU Munich</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universitat de Barcelona</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>aequa-tech</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair strategies is particularly challenging in heavily gender-marked languages, such as Italian. To address this, the Gender-Fair Generation challenge intends to help shift toward gender-fair language in written communication. The challenge, designed to assess and monitor the recognition and generation of gender-fair language in both mono- and cross-lingual scenarios, includes three tasks: (1) the detection of gendered expressions in Italian sentences, (2) the reformulation of gendered expressions into gender-fair alternatives, and (3) the generation of gender-fair language in automatic translation from English to Italian. The challenge relies on three diferent annotated datasets: the GFL-it corpus, which contains Italian texts extracted from administrative documents provided by the University of Brescia; GeNTE, a bilingual test set for gender-neutral rewriting and translation built upon a subset of the Europarl dataset; and Neo-GATE, a bilingual test set designed to assess the use of non-binary neomorphemes in Italian for both fair formulation and translation tasks. Finally, each task is evaluated with specific metrics: average of F1-score obtained by means of BERTScore computed on each entry of the datasets for task 1, an accuracy measured with a gender-neutral classifier, and a coverage-weighted accuracy for tasks 2 and 3.</p>
      </abstract>
      <kwd-group>
        <kwd>Gender-fair language</kwd>
        <kwd>Inclusive language</kwd>
        <kwd>Unfairness detection</kwd>
        <kwd>Machine translation</kwd>
        <kwd>Generation</kwd>
        <kwd>Neomorphemes</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
A</p>
      <p>CALAMITA</p>
    </sec>
    <sec id="sec-2">
      <title>1. Challenge: Introduction and</title>
    </sec>
    <sec id="sec-3">
      <title>Motivation</title>
      <p>reinforcing gender stereotyp1e]s. [</p>
      <sec id="sec-3-1">
        <title>In order to pursue the goalsfaoifrnessand inclusive</title>
        <p>ness, measures that take into account the importance of</p>
        <sec id="sec-3-1-1">
          <title>Gender-fair language, also known as inclusive languaget,he correlation between language and gender become consists in usinlginguistic expressions that promote gecne-ntral. Especially in heavily gender-marked languages der equality, inclusion of non-binary identities, and avosuidch as Italian, the use and applicationgoefnder-fair</title>
          <p>deed, in these languages, several are the elements one
has to take into account to ensure a gender-fair use of
language. However, adopting a gender-fair language is
crucial given the negative efects of the masculine
generics, documented in a range of empirical stud2ie,s3][;
and recent years witnessed an increasea winareness and
efort to address these issues by promoting gender-fair
language [4].</p>
          <p>
            In Italian, the masculine is not only used to refer to
and address men but also generic or unknown
individuals; mixed-gender groups, regardless of the proportion
of genders of its members; women, typically when
occupying prestigious roles; and genderqueer people, given
that there is no codified grammatical gender for referring
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons Litceonse them 5[]. This use, though, makes women and
genAttribution 4.0 International (CC BY 4.0).
derqueer people invisible, giving rise to a proper injustice persons belonging to a mixed-gender group - e.g.,
[
            <xref ref-type="bibr" rid="ref15">6, 7, 8</xref>
            ]. Extensive empirical literature also highlights i cittadini (the.M citizens:M) used for a group of
how certain gendered expressions influence our cogni- citizens of diferent genders;
tion, with masculine terms evoking male images and • thegeneric masculine or feminine, i.e., the use of
reducing, e.g, the likelihood of women applying for or be- a singlegendered expressiotno refer to a generic
ing considered suitable for a job position (for an overview or unknown person - e.gi.l, candidato deve avere
see [
            <xref ref-type="bibr" rid="ref4 ref6">9, 10</xref>
            ]). tutti i requisiti (the.M candidate:M has to possess
          </p>
          <p>Crucially,such unfair linguistic practicaerseperpetu- all the requirements);
ated in language technologie1s1[]. This becomes partic- • theincongruous gender, i.e., the use of a
gramularly evident in languages, like Italian, for which NLP matical gender that does not match the referent’s
tools often adopt masculine and stereotypical represen- gender - e.g.,il professore ordinario Maria Rossi
tations, making undue binary gender assumptio1n2]s. [ (the.M full.M professor:M Maria Rossi).</p>
          <p>We propose theGender-Fair Generation challenge
at CALAMITA 2024 [13], whose goal is to reduce the2) Fair reformulation: the second task tests models’
use of gender-unfair expressions in written Italian,afboi-lity to rewrite gendered expressions into alternative
cusing on both monolingual and cross-lingual scenariogsender-fair expressions. To achieve this goal, various
(English-Italian). Our challenge sistructured into three gender-fair language strategies can be employed. In
partasks—i) gendered language detectioiin) ,fair reformu- ticular, we will employobscuration strategies:
lation, andiii) fair translationa—cross three diferent • conservative obscuration, i.e., the use of
expresdatasets. Namely, the newly created GFL-it corpus, com- sions and constructions that avoid providing
inposed of Italian texts extracted from 35 documents pro- formation on the referent’s gender – iel.gco.,rpo
vided by the academic administration ofice of the Uni- docente (the teaching body) ocroloro che
insegversity of Brescia and annotated following specific guide- nano (those who teach) instead oifprofessori
lines [1]; GeNTE, a bilingual test set for gender-neutral (the.M professors:M);
rewriting and translation built on a subset of the Europarl • innovative obscuration, i.e., the use of novel,
dataset1[4]; and Neo-GATE, a bilingual test set designed gender-neutral marekrsinstead of the gendered
to evaluate the use of nonbinary neomorphemes in Italian ones – e.g.,lǝ professorǝ (the.INN professor:INN)
[15].1 We combine and repurpose these datasets across instead ofil professore (the.M professor:M) olra
the three tasks envisioned in the Gender-Fair Generation professoressa (the.F professor:F2).
challenge.</p>
          <p>This report is structured as follows: in Sec2t,iwone As we further discuss in Secti3o,nthe released version
provide a description of our challenge; in Sect3,iowne of GFL-it for this challenge and GeNTiEnclude
referpresent the three datasets in detail; in Sec4,twioende- ences and annotations designed tfohre former strategy,
scribe the metrics involved in our task; in Sec5t, iwoen whereas Neo-GATE for the latter.
describe the limitations of our work, and finally, in Sec- Notethatthe chosen strategies do not exhaust the
tion6, we discussthe ethical issues. full range of possibiliti:eswe discarded, for the moment,
visibility strategies such as the repetition of an
expression in the feminine and the masculine - ei.gp.,rofessori
2. Challenge: Description e le professoresse (the.M professors:M and the.F
professors:F) - and the repetitioninofthree gendered forms
The Gender-Fair Generation challenge is organized in(tfoeminine, masculineand innovative–) e.g., i professori, lǝ
three tasks, which we present in detail below. professorǝ e le professoresse (the.M professors:M, the.INN
professors:INN and the.F professors:F).
1) Gendered language detection: the first task tests the
models’ ability to identify referentially gender-mark3e)dFair translation: like the second task, the third one
expressions within Italian senten,cneasmely those ex- is designed to test the models’ ability to generate
genderpressions whose (typically grammatical) gender is linkefdair language texts, but in tchreoss-lingual context of
to their human refere.nRteferentially gendere(dhence- automatic translation from English into Italian. For
exforth simplygendered) language includes: ample, consider applying the two gender-fair language
• theoverextended masculine or feminine, i.e., the strategies described above to the translation of the
senuse of a singlegendered expressionto refer to tence “I am glad to know such knowledgeable doctors”:
1In this report, we refer to innovative gender-fair strategies such as • conservative obscuratioSno:no felice di conoscere
the schwa as “neomorphemes”. Although aware that this terminol- un personale medico così preparato. [medical staf]
ogy is controversial, we adopted it for simplicity and do not intend
our terminology to imply any substantive stance. 2We indicate the innovative forms with “INN” in theegsl.oss
GFL-it GeNTE Neo-GATE Task total
cally, each entry is described by the following attributes:
• id_text: The unique ID for each text.
• text: The entire text of the entry.
• list_spans: The list containing all spans found in</p>
          <p>the text.
• rewritten_texts_generico: A reformulation of the
entire text where spalnasbeled as generic are
replaced.
• rewritten_texts_sovraesteso: A reformulation of
the entire text where spalnabseled as
overextended are replaced.
• rewritten_texts_generico_e_sovraesteso: A
reformulation of the entire text where spans are
randomly replaced by available options
irnewritten_texts_generico orrewritten_texts_sovraesteso.
• innovative obscuratioSno:no contentǝ di conoscere</p>
          <p>medicǝ così preparatǝ.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Data description</title>
      <sec id="sec-4-1">
        <title>For our challeng,ewe propose three benchmarks dedi</title>
        <p>cated to the evaluation of gender-fair language
generation, (GFL-it3, GeNTE [14],4 and Neo-GATE [15]),5 and a Each span inlist_spans follows the structure:
total of 7 prompts to be used across the tasks and datasets.</p>
        <p>We describe the datasetssiunbsections3.1, 3.2, and 3.3 • span: The textual representation of the span.
respectively, and the promptssiunbsection3.4. • start: The starting index of the span in the text.</p>
        <p>Statistics about the benchmarks and their use within • end: The ending index of the span in the text.
this challenge proposal are available in Tab1l.eGFL-it • labels: A list of the types of genderleadnguage
containsa totalof 2,187 texts,among which 5 expert used inthe selected span; spossible values are
annotators identified an average 3o.f24 unfair spans (in overextended, generic and incongruous
gentotal 3,908) in 1,206 texts. For each identified span, the an- der.
notators proposed various gen-dfaeir alternatives, with • key_span: The concatenation ospfan, start and
an average of3.8 alternatives per span. For more detailed end attributesit; can be used as an ID for each
statistics about GeNTE and Neo-GATE we refer to the span contained inside a text.
respective papers.</p>
        <p>We propose to use thGeFL-it corpusfortasks 1 and 27,
namely, those regardinggendered language detection
and fair reformulation.
3.1. GFL-it</p>
      </sec>
      <sec id="sec-4-2">
        <title>GFL-it was built on documents and texts from Univer</title>
        <p>sity website pages provided by the University of Bresc3i.a2. . GeNTE
It constitutes an expansion of the corpus presenteGdeiNnTE is a parallel English→ Italian test set16[].
OrigiRosola et al.[1]. The corpus comprises a total of 35 docun-ally designed to evaluate MT models’ ability to perform
ments in Italian, split into 2,187 texts. Each text was angneon-der-neutral translations, GeNTE was built upon a
tated by 5paid expert annotators following tohreiginal subset of the Europarl corpu1s7][, which is
representaannotation scheme1][. First,the annotatoridsentified all tive of natural, formal communicative situations from the
the spans that contained any gender-unfairness, distininst-itutional domain, the context where gender-neutral
guishing among:overextended (3,465), generic (530) language is most accepted and encourag1e6d, [14].
Overand incongruous gender (31) (see 2). Overall, 3,908 all, it consists of 1,500E&lt;nglish source, gendered Italian
spans were identified. Then, they provided at least onreeference, gender-neutral Italian reference&gt; triplets aligned
alternative per span. The alternatives could belongattothe sentence level, which always contain at least one
any of thegender-fair strategies: conservative or
innova</p>
        <p>mention of human referents. The gendered Italian
refertive obscuration, conservative or innovative visibiliteyn,ocre (REF-G) comes from the original Europarl corpus,
hybrid alternatives (i.e., any combination of these typwesh).ereas thegendern-eutral reference (REF-N) was
pro</p>
        <p>Given thatGFL-it is annotated for spans, each texdtuced by professional translatowrhso edited gendered
contains a list of diferent spans and their reformulatiofnosrms intogendern-eutral alternatives.
in diferent forms ofgender-fairlanguage6. More
specifi</p>
      </sec>
      <sec id="sec-4-3">
        <title>7For task 2, we used a classifier that distinguishes between gendered</title>
      </sec>
      <sec id="sec-4-4">
        <title>3https://github.com/simonasnow/GFL-it-Dataset and gender-neutral texts (see Sect4i)o.nHence,we only used the</title>
      </sec>
      <sec id="sec-4-5">
        <title>4https://huggingface.co/datasets/FBK-MT/GeNTE GFL-ittextswhere the annotators identifiegdendered expressions</title>
        <p>5https://huggingface.co/datasets/FBK-MT/Neo-GATE (= gendered class) and the texftosr which annotators provided at
6For the purpose of the task 2, only the conservative obscured refloeras-t one conservative obscured reformula(t=iognender-neutral
mulations have been released in this version of the dataset. class) for a total amount of 1,206 texts.
[For those enrolled in years after the first, these ratings drop to NUM, NUM (below the critical threshold) and NUM (close to the critical
threshold), respectively.]</p>
        <p>When you assumed ofice, Mr Schreyer, you assured us that you would strive to achieve this.</p>
        <p>Al momento della sua nomina, signor [Mr] Schreyer, ci aveva promesso che si sarebbe adoperato
[(would have) strived] in tal senso.</p>
        <p>Al momento della sua nomina, Schreyer, ci aveva promessoun impegno [a commitment] in tal senso.</p>
        <p>To some extent, those of us who are politicians find ourselves in the middle.</p>
        <p>In certa misura quelli [those (of us)] di noi che sono politici[politicians] si trovano in una posizione intermedia.</p>
        <p>In certa misura chi di noi [who, among us,] svolge attività politica [carries out political activities] si trova in una
posizione intermedia.</p>
        <p>As shown in Table3, GeNTE represents two types Each entry in GeNTE is organized into the following
of phenomena, which are equally represented withifienlds:
the corpus. Namely,i) Set-N, featuring 750
genderambiguous source sentences that require to be ren- • ID: The unique GeNTE ID.
deredgendern-eutrally; andii) Set-G featuring gender- • Europarl_ID: The original sentence ID from
Euunambiguous source sentences, to be properly rendered roparl’s common-test-set 2.
with gendered (masculine or feminine) forms. Crucially, • SET : Indicates whether the entry belongs to the
these two sets are a key feature of GeNTE, as they al- Set-G or theSet-N subportion of the corpus.
low benchmarking whether systems are able to per- • SRC: The English source sentence.
form gender-neutral translations, but only when desir- • REF-G: The gendered Italian reference
translaable. As a matter of fact, when referents’ gender is un- tion.
known or irrelevant, undue gender inferences should • REF-N : The gender-neutral Italian reference,
pronot be made andgendern-eutral language (i.e., conser- duced by a professional translator.
vative obscuration strategy) should be used. However, • GENDER: For entries belonging to the Set-G, it
gendern-eutralization should not be always enforced, indicates if the entry is Feminine or Masculine.
and when a referent’s gender is known or relevant,
models should not over-generalize gteondern-eutral
generations.</p>
      </sec>
      <sec id="sec-4-6">
        <title>We propose the use of the whole GeNTE for the</title>
        <p>translation task 3, testing models’ ability to produce
gendern-eutral translations only when appropriate. For
SOURCE
REF-M
REF-F
REF-TAGGED
ANNOTATION</p>
        <p>After the accident, they took me to the hospital and I stayed there for a whole month.</p>
        <p>Dopo l’incidente, mi hanno portato all’ospedale e sono rimasto lì per un mese intero.</p>
        <p>Dopo l’incidente, mi hanno portata all’ospedale e sono rimasta lì per un mese intero.</p>
        <p>Dopo l’incidente, mi hanno portat@ all’ospedale e sono rimast@ lì per un mese intero.</p>
        <p>portato portata portat@; rimasto rimasta rimast@;
thefair reformulation task 2, we only repurposepart We propose to use all Neo-GATE entries faolrl three
of the Italian portion of the corpi.ues., REF-G references tasks of our challenge. While for tasks 1ge(ndered
fromSet-N. language detection) and 2 (fair reformulation) we
only use Italian references – nameblyoth REF-M and
3.3. Neo-GATE REF-F for task 1, and REF-M only for task–2as input
for the models, for task f3a(ir translation) we use the
Similarly to GeNTE, Neo-GATE is a parallel corpus de-English SOURCE sentences.
signed for genderfa-irEnglish → Italian MT evaluation.</p>
        <p>Here, however, the focus is on the use of gendfaeir-neo- 3.4. Example of used prompts
morphemes (i.e., innovative obscuration strategy) rather
than conservative gender-neutral language. Neo-GATEhis section describetsheprompts we propose for our
was built on GATE1[8], a test semtanuallycreatedspecif- challenge, with examples available inTable 5.
icallyto evaluate gender reformulation and gender bias inIn promptsA and B, we ask the model to identify
MT. In GATE, the gender of human entities is unknownth,e gendered expressions (introduced
byttahge[Espresi.e., there are no linguistic elements providing gendersiionn-e]:) in the text given as input; if no gendered
exformation about human referents in the (English) sourpcreession is detected in the text (initialized with the tag
sentences. [Genere marcato]:) the model should output 0. The model</p>
        <p>Neo-GATE includes an annotation that defines thecan recognize more than one gendered expression.
words upon which the evaluation is based. It includesIn promptsC, D, and E, the shots include one line
the three forms required for the evaluation, i.e., the mstasa-rting with the t[aGgenere marcato]:, indicating that the
culine and feminine forms, and forms featuring placfeo-llowing sentence is gendered. Then, in promCptasnd
holders in place of Italian overt gender markers. BefoDrethe following line starts wi[tNheutro]: followed by a
the evaluation, the placeholders must be replaced wigtehnder-neutral reformulation, whereaEs int starts with
the correct forms in the desired neomorpheme paradi[gNme.omorfema]: and includes the innovative obscuration</p>
      </sec>
      <sec id="sec-4-7">
        <title>For this task, Neo-GATE was adapted to a version of tahleternative of the first sentence, with neomorphemes in</title>
        <p>‘schwa’ paradigm 1[9, 20], to which we refer asschwa- place of the masculine form8s.
simple here, i.e., the placeholders were replaced with PromptsF and G start with the ta[Ignglese]:
folthe forms described in AppendAix. lowed by the English source sentence to be translated.</p>
        <p>Like GeNTE, Neo-GATE includes Italian referencesIn promptF, the second line either starts with the tag
that difer exclusively in gender expression. Besides th[eItaliano, genere marcato]: (see F - Exemplar format 1 in</p>
      </sec>
      <sec id="sec-4-8">
        <title>English source sentence, all entries in Neo-GATE havTeable 5) if it is followed by a gendered translation or with</title>
        <p>three Italian references: REF-M, where the gender tohfe tag[Italiano, neutro]: if the subsequent translation
words referring to human beings is masculine, REF-iFs, gender-neutral (see F - Exemplar format 2). Models
where human beings are referred to as feminine, anadre required to produce the correct tag and translation</p>
      </sec>
      <sec id="sec-4-9">
        <title>REF-TAGGED, where placeholders replace overt markerdsepending on the presence or absence of gender cues</title>
        <p>of gender – here adapted to tshcehwa-simple paradigm. in the source. Finally, prompGtincludes two diferent</p>
      </sec>
      <sec id="sec-4-10">
        <title>However, diferently from GeNTE, the English sentencestranslations after the source sentence: the first, preceded</title>
        <p>in Neo-GATE never include gender cues. An example obfy the tag[Italiano, genere marcato]:, includes a
transa Neo-GATE entry is available in Tabl4e. lation featuring masculine forms in reference to human</p>
      </sec>
      <sec id="sec-4-11">
        <title>Each entry in Neo-GATE includes the following fields:beings, whereas the second translation starts with the</title>
        <p>tag[Italiano, neomorfema]: and uses neomorphemes in
• #: The entry identifier within Neo-GATE. reference to human beings. Models are required to
pro• GATE-ID: A unique identifier of the original duce both translations, though only the second will be
GATE entry, composed of a prefix indicating theextracted in post-processing and used for the evaluation.
subset of origin followed by a serial number. In particular, prompts D, E, F, and G are based on the
• SOURCE: The English source sentence. ones used in previous experiments on the same datasets
• REF-M: The Italian reference where all gender[-12, 15], and were in turn inspired by the format proposed
marked terms are masculine. by Sánchez et al. 2[1].
• REF-F : The Italian reference where all
gender</p>
        <p>marked terms are feminine.
• REF-TAGGED: The Italian reference where all
gender-marked terms are tagged with Ne8oW-e here usedneutro (neutral/neuter), despite being aware of its</p>
        <sec id="sec-4-11-1">
          <title>GATE’s annotation. ambiguity withneuter, a grammatical gender not present in the</title>
          <p>• ANNOTATION : The word level annotation. Italian linguistic system. However, nothing substantive hinges on
this terminological choice.</p>
          <p>A Instruction</p>
          <p>Exemplar format
B Instruction</p>
          <p>Exemplar format
C Instruction</p>
          <p>Exemplar format
Instruction
Exemplar format
Instruction</p>
          <p>Exemplar format
D
E
F
G</p>
          <p>Task 1: Gendered language detection
Identifica le espressioni che contengono dei marcatori di genere femminile o maschile. Se più di
un’espressione è identificata, separale con “ ;”. Se nessuna espressione è identificata, restituisci 0.
[Genere marcato]: A livello di Ateneo ha partecipato il 76% degli iscritti del 1° anno del CdL e
l’81% di tutti gli altri studenti.
[Espressione]: degli iscritti ; tutti gli altri studenti
Identifica le espressioni che contengono dei marcatori di genere femminile o maschile. Se più di
un’espressione è identificata, separale con “ ;”. Se nessuna espressione è identificata, restituisci 0.
[Genere marcato]: Quest’anno mi sono ammalata già due volte.
[Espressione]: ammalata</p>
          <p>Task 2: Fair reformulation
Riformula la seguente frase utilizzando un linguaggio neutro rispetto al genere dei referenti
umani, evitando l’uso di forme maschili e femminili.
[Genere marcato]: - Il 79% dei laureati lavora (a un anno dalla laurea).
[Neutro]: - Il 79% delle persone laureate lavora (a un anno dalla laurea).</p>
          <p>Riformula la seguente frase utilizzando un linguaggio neutro rispetto al genere dei referenti
umani, evitando l’uso di forme maschili e femminili.
[Genere marcato]: Secondariamente, fino a che punto aumenta la trasparenza e la responsabilità
dei parlamentari europei?
[Neutro]: Secondariamente, fino a che punto aumenta la trasparenza e la responsabilità dei
membri del Parlamento Europeo?</p>
          <p>Task 3: Fair translation
Instruction Traduci la seguente frase inglese in italiano seguendo queste regole:
1. Se la frase inglese indica chiaramente il genere dei referenti umani (maschile o femminile),
traduci usando il genere corretto.
2. Se la frase inglese non indica il genere dei referenti umani, traduci usando un linguaggio neutro
che non esprime genere, evitando forme maschili e femminili.</p>
          <p>Exemplar format 1 [Inglese]: However, it is important that the Commissioner has declared his loyalty to the President
himself.
[Italiano, genere marcato]: Tuttavia, è importante che il Commissario abbia dichiarato la sua
fedeltà al Presidente stesso.</p>
          <p>Exemplar format 2 [Inglese]: Secondly, how far does it increase transparency and accountability of the MEPs?
[Italiano, neutro]: Secondariamente, fino a che punto aumenta la trasparenza e la responsabilità
dei membri del Parlamento Europeo?
Instruction
Exemplar format</p>
          <p>Traduci la seguente frase inglese in italiano usando il neomorfema “@”. Il neomorfema “@” deve
essere usato come sostituto dei morfemi maschili e femminili nelle parole che si riferiscono agli
esseri umani.
[Inglese]: The partner of a friend of mine invited us to dinner.
[Italiano, genere marcato]: Il partner di un mio amico ci ha invitati a cena.</p>
          <p>[Italiano, neomorfema]: L@ partner di un@ mi@ amic@ ci ha invitat@ a cena.</p>
        </sec>
      </sec>
      <sec id="sec-4-12">
        <title>9https://huggingface.co/spaces/evaluate-metric/bertscore</title>
        <p>produced by the generative model, computing the m6ax.- Ethical issues
imum F1-score. Once the correspondences are set for
each entry, we average the scores. The proposed tasks in this challenge have the purpose of</p>
        <p>For the evaluation of gender-neutral reformulatiorne—ducing the use of gend-eurnfairexpressions inheavily
i.e., with GFL-it and GeNTE in task 2—and translationg—ender-marked languages (i.e., Italian) that afect the
i.e., with GeNTEand Neo-GATEin task 3—we propose visibility of other genders (in particular, feminine and
an accuracy score based on the labels producedtbhye non-binary). Although the datasets have been built by
classifier introduced in Piergentili et[a1l4.]. More specif- experts ofgender-fairlanguage, the group of annotators
ically, we use version 2 of the classifier, introduced inof GFL-itwas notgenderb-alanced asonly 2 out of5
Savoldi et al.[12]. This classifier assigns a label to each annotators wermeen.
model output, eithegrendern-eutral or gendered. We Moreover, we are aware of the fact that the use of
neothen compare those labels against the true labels, im.e.o,rphemes like the schwa ǝ makes reading harder for
always gendern-eutral in the reformulation task and epie-ople with dyslexia or visual impairment4s, 2[4, 25].
ther gendered ogrendern-eutral for the translation tasTkh, is issue, however, is mitigated thanks to the possibility
depending on whether the entry belongsSetto-G or of selecting the most suitable neomorpheme according
Set-N respectively. The final score is computed as theto each user’s needs.In particular, both people with
corpus-level percentage of correct labels. dyslexia or visual impairments can rely on screen
read</p>
        <p>For neomorpheme-based gendefra-ir reformulation ers, which difer in their ability to correctly interpret
(task 2) and translation (task 3) based on Neo-GATEsp,ecific neomorphemes: the possibility to select diferent
we propose the coverage-weighted accuracy describnedeomorphemes allows each user to select the one(s) their
in Piergentili et a[l1.5] as the main metric. This metricscreenreader interpret best.
takes into account both how accurately a model generates
neomorphemes and the proportion of annotations (7i.e.., Data license and copyright
either of the masculine, feminine, or innovative forms)
found during the evaluation, thus allowing for fair system issues
comparisons and rankings. As complementary metric
to assess models’ ability to correctly generate neomCorre-ative Commons Attribution 4.0 International license
phemes, we propose reporting the mis-generation sco(CreC BY 4.0). https://creativecommons.org/licenses/by-sa/
[15] as well. This metric ca nflag undesired behaviors 4.0/deed.it
even despite good accuracy, as it counts cases where
models generate neomorphemes inappropriately, for instan8c.e Acknowledgments
by applying the use of neomorphemes to words that do
not refer to human entities (e.g., by genera‘tianvgo@l’ Beatrice Savoldi is supported by the PNRR project FAIR
instead of ‘tavolo’, en: table). - Future AI Research (PE00000013), under the NRRP</p>
      </sec>
      <sec id="sec-4-13">
        <title>MUR program funded by the NextGenerationEU. Luisa</title>
        <p>5. Limitations Bentivogli is funded by the Horizon Europe research
and innovation programme, under grant agreement No
Our work presents some limitations. Firstly, the datas1e0t1s135798, project Meetween (My Personal AI Mediator
employed only derive from specific domains: GFL-it ex-for Virtual MEETtings BetWEEN People). The work of
clusively contains data from administrative documeVnitvsiana Patti and Marco Madeddu is supported by
“HARand oficial web pages of the University,GeNTE from MONIA” project - M4-C2, I1.3 Partenariati Estesi -
Casdocuments of the European Parliament, and Neo-GATcEade Call - FAIR - CUP C63C22000770006 - PE PE0000013
data manually created by experts. The corpora couuldnder the NextGenerationEU programme.
be expandedto other domainasnd annotated by more The annotation of GFL-it has been partially funded by
annotators in future research. Secondly, our metricsUanrieversità degli Studi di Brescia as part of the actions
only a first attempt and others should be explored in thperovided for by the Gender Equality Plan.
future. Moreover, we only tested one paradigm of
neomorphemes, namely theschwa-simple, while many oth- References
ers exist (e.g., the asterisk, the ‘-u’, t@he’-‘ see [23] for a
complete list), and even more could be proposed. Furthe[r1-] M. Rosola, S. Frenda, A. T. Cignarella, M. Pellegrini,
more, GeNTE and Neo-GATE do not contain mixed texts A. Marra, M. Floris, et al., Beyond obscuration and
where rewriting is needed with respect to one entity but visibility: Thoughts on the diferent strategies of
not others. gender-fair language in italian, in: CLiC-it 2023.</p>
      </sec>
      <sec id="sec-4-14">
        <title>Proceedings of the 9th Italian Conference on Com- ties of LAnguage Models in ITAlian, in: Proceed</title>
        <p>putational Linguistics. Venice, Italy, November 30- ings of the 10th Italian Conference on
ComputaDecember 2, 2023., volume 3596, CEUR-WS, 2023, tional Linguistics (CLiC-it 2024), Pisa, Italy,
Decempp. 1–10. ber 4 - December 6, 2024, CEUR Workshop
Proceed[2] J. Silveira, Generic Masculine Words and Think- ings, CEUR-WS.org, 2024.</p>
        <p>
          ing, Women’s Studies International Quarterly[134] A. Piergentili, B. Savoldi, D. Fucci, M. Negri, L.
Ben(1980) 165–178. URL: https://www.sciencedirect. tivogli, Hi guys or hi folks? benchmarking
gendercom/science/article/pii/S01480685809211.32 neutral machine translation with the GeNTE
cor[3] P. Gygax, S. Sato, A. Öttl, U. Gabriel, The mas- pus, in: H. Bouamor, J. Pino, K. Bali (Eds.),
Proculine form in grammatically gendered languages ceedings of the 2023 Conference on Empirical
and its multiple interpretations: A challenge for Methods in Natural Language Processing,
Assoour cognitive system, Language Sciences 83 (2021) ciation for Computational Linguistics, Singapore,
101328. 2023, pp. 14124–14140. URL: https://aclanthology.
[4] G. Sulis, V. Gheno, The debate on language and gen- org/2023.emnlp-main.873.doi:10.18653/v1/2023.
der in italy, from the visibility of women to inclu- emnlp-main.873.
sive language (1980s–2020s), The Italianist 42 (2022)[15] A. Piergentili, B. Savoldi, M. Negri, L. Bentivogli,
153–183. doi:10.1080/02614340.2022.2125707. Enhancing gender-inclusive machine translation
[5] G. Visibility, N. across Languages, Beyond pro- with neomorphemes and large language models,
nouns, The Oxford Handbook of Applied Philoso- in: C. Scarton, C. Prescott, C. Bayliss, C.
Oakphy of Language (2024) 320. ley, J. Wright, S. Wrigley, X. Song, E.
Gow[6] M. Rosola, Linguistic hermeneutical injustice, So- Smith, R. Bawden, V. M. Sánchez-Cartagena,
cial Epistemology (2024). do1i:0.1080/02691728. P. Cadwell, E. Lapshinova-Koltunski, V.
Cabar2024.2401143. rão, K. Chatzitheodorou, M. Nurminen, D. Kanojia,
[
          <xref ref-type="bibr" rid="ref15">7</xref>
          ] S. J. Kapusta, Misgendering and its moral contesta- H. Moniz (Eds.), Proceedings of the 25th Annual
bility, Hypatia 31 (2016) 502–519. Conference of the European Association for
Ma[8] R. Dembrof, D. Wodak, He/she/they/ze, Ergo chine Translation (Volume 1), European
Associa(2018). tion for Machine Translation (EAMT), Shefield, UK,
[9] S. Sczesny, M. Formanowicz, F. Moser, Can gender- 2024, pp. 300–314. URL: https://aclanthology.org/
fair language reduce gender stereotyping and dis- 2024.eamt-1.25.
crimination?, Frontiers in psychology 7 (2016[1)6] A. Piergentili, D. Fucci, B. Savoldi, L. Bentivogli,
154379. M. Negri, Gender neutralization for an
inclu[
          <xref ref-type="bibr" rid="ref4 ref6">10</xref>
          ] P. Gygax, S. Zuferey, U. Gabriel, Le cerveau pense-t- sive machine translation: from theoretical
founil au masculin, Cerveau, langage et représentations dations to open challenges, in: E.
Vanmassensexistes, Paris, Le Robert (2021). hove, B. Savoldi, L. Bentivogli, J. Daems, J.
Hack[
          <xref ref-type="bibr" rid="ref2">11</xref>
          ] S. L. Blodgett, S. Barocas, H. Daumé III, H. Wal- enbuchner (Eds.), Proceedings of the First
Worklach, Language (technology) is power: A critical shop on Gender-Inclusive Translation
Technolosurvey of “bias” in NLP, in: D. Jurafsky, J. Chai, gies, European Association for Machine
TranslaN. Schluter, J. Tetreault (Eds.), Proceedings of the tion, Tampere, Finland, 2023, pp. 71–83. URLh:ttps:
58th Annual Meeting of the Association for Com- //aclanthology.org/2023.gitt-.1.7
putational Linguistics, Association for Compu[t1a7-] P. Koehn, Europarl: A Parallel Corpus for
Statistional Linguistics, Online, 2020, pp. 5454–5476. URL: tical Machine Translation, in: Proceedings of the
https://aclanthology.org/2020.acl-main.4.8d5oi:10. tenth Machine Translation Summit, AAMT, Phuket,
18653/v1/2020.acl-main.485. TH, 2005, pp. 79–86. URL: http://mt-archive.info/
[12] B. Savoldi, A. Piergentili, D. Fucci, M. Negri, L. Ben- MTS-2005-Koehn.pdf.
        </p>
        <p>tivogli, A prompt response to the demand for a[u18-] S. Rarrick, R. Naik, V. Mathur, S. Poudel, V.
Chowdtomatic gender-neutral translation, in: Y. Graham, hary, GATE: A challenge set for gender-ambiguous</p>
      </sec>
      <sec id="sec-4-15">
        <title>M. Purver (Eds.), Proceedings of the 18th Confer- translation examples, in: Proceedings of the 2023</title>
        <p>ence of the European Chapter of the Association AAAI/ACM Conference on AI, Ethics, and Society,
for Computational Linguistics (Volume 2: Short AIES ’23, Association for Computing Machinery,
Papers), Association for Computational Linguis- New York, NY, USA, 2023, p. 845–854. URL:https:
tics, St. Julian’s, Malta, 2024, pp. 256–267. URL: //doi.org/10.1145/3600211.3604675. doi:10.1145/
https://aclanthology.org/2024.eacl-shor.t.23 3600211.3604675.
[13] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran[-19] A. M. Thornton, Genere e igiene verbale:
cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri- l’uso di forme con@ in italiano, Annali
naldi, D. Scalena, CALAMITA: Challenge the Abili- Del Dipartimento Di Studi Letterari,
Linguis</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>A. The schwa-simple paradigm</title>
      <p>Feminine</p>
      <p>Schwa
&lt;ENDS&gt; portion of the word diferentiating gendered forms , singular o, e, tore a, essa, trice @, tor@
&lt;ENDP&gt; portion of the word diferentiating gendered forms , plural i, tori e, esse, trici @, tor@
&lt;DARTS&gt; definite article, singular il, lo, l’ la, l’ l@
&lt;DARTP&gt; definite article, plural i, gli le l@
&lt;IART&gt; indefinite article uno, un una, un’ un@
&lt;PARTP&gt; partitive article, plural dei, degli delle de@
&lt;PREPdiS&gt; articulated preposition with root ‘di’, singular del, dello, dell’ della, dell’ dell@
&lt;PREPdiP&gt; articulated preposition with root ‘di’, plural dei, degli delle dell@
&lt;PREPaS&gt; articulated preposition with root ‘a’, singular al, allo, all’ alla, all’ all@
&lt;PREPaP&gt; articulated preposition with root ‘a’, plural agli, ai alle all@
&lt;PREPdaS&gt; articulated preposition with root ‘da’, singular dal, dallo, dall’ dalla, dall’ dall@
&lt;PREPdaP&gt; articulated preposition with root ‘da’, plural dagli dalle dall@
&lt;PREPinP&gt; articulated preposition with root ‘in’, plural negli nelle nell@
&lt;PREPsuS&gt; articulated preposition with root ‘su’, singular sul, sullo, sull’ sulla, sull’ sull@
&lt;PREPsuP&gt; articulated preposition with root ‘su’, plural sugli sulle sull@
&lt;DADJquelS&gt; demonstrative adjective (far), singular quel, quello, quell’ quella, quell’ quell@
&lt;DADJquelP&gt; demonstrative adjective (far), plural quegli quelle quell@
&lt;DADJquestS&gt; demonstrative adjective (near), singular questo, quest’ questa, quest’ quest@
&lt;DADJquestP&gt; demonstrative adjective (near), plural questi queste quest@
&lt;POSS1S&gt; possessive adjective, 1st person singular, singular mio mia mi@
&lt;POSS1P&gt; possessive adjective, 1st person singular, plural miei mie mi@
&lt;POSS2S&gt; possessive adjective, 2nd person singular, singular tuo tua tu@
&lt;POSS2P&gt; possessive adjective, 2nd person singular, plural tuoi tue tu@
&lt;POSS3S&gt; possessive adjective, 3rd person singular, singular suo sua su@
&lt;POSS3P&gt; possessive adjective, 3rd person singular, plural suoi sue su@
&lt;POSS4S&gt; possessive adjective, 1st person plural, singular nostro nostra nostr@
&lt;POSS4P&gt; possessive adjective, 1st person plural, plural nostri nostre nostr@
&lt;PRONDOBJS&gt; direct object pronoun, singular lo la l@
&lt;PRONDOBJP&gt; direct object pronoun, plural li le l@</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>tici E</given-names>
            <surname>Comparati</surname>
          </string-name>
          .
          <source>Sezione Linguistica</source>
          <volume>11</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          11-
          <fpage>54</fpage>
          . URL: http://www.serena.unina.it/index.php/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>aionlin/article/view/962.3doi:https://doi.org/</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          10.6093/
          <fpage>2281</fpage>
          -6585/9623. [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baiocco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rosati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pistella</surname>
          </string-name>
          , Italian proposal for
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>bian Mental Health</source>
          <volume>27</volume>
          (
          <year>2023</year>
          )
          <fpage>248</fpage>
          -
          <lpage>253</lpage>
          . URLh: ttps:
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          //doi.org/10.1080/19359705.
          <year>2023</year>
          .
          <volume>2183537</volume>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <volume>1080</volume>
          /19359705.
          <year>2023</year>
          .
          <volume>2183537</volume>
          . [21]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Andrews</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Stenetorp</surname>
          </string-name>
          , M. Artetxe,
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>tion with large language models</source>
          ,
          <year>2024</year>
          . URhLt: tps:
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          //arxiv.org/abs/2309.0317.5arXiv:
          <fpage>2309</fpage>
          .
          <fpage>03175</fpage>
          . [22]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang*</surname>
          </string-name>
          , V. Kishore*, F. Wu*,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>ing Representations</source>
          ,
          <year>2020</year>
          . URhLt:tps://openreview.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>net/forum?id=SkeHuCVFD</article-title>
          . r [23]
          <string-name>
            <given-names>V.</given-names>
            <surname>Gheno</surname>
          </string-name>
          ,
          <article-title>Lo schwa tra fantasia e norma,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>La falla</surname>
          </string-name>
          (
          <year>2020</year>
          ). URL: https://lafalla.cassero.it/
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <article-title>lo-schwa-tra-fantasia-e-norm</article-title>
          . a/ [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Iacopini</surname>
          </string-name>
          ,
          <article-title>Lo schwa (ǝ) che rende l'inclusione</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>inaccessibile</surname>
          </string-name>
          , Web accessibile (
          <year>2021</year>
          ). URLh:ttps:
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>C7%9D-che-rende-linclusione-inaccessibi</article-title>
          .le/ [25]
          <string-name>
            <surname>C. D. Santis</surname>
          </string-name>
          , L'emancipazione grammaticale
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>non passa per una e rovesciata</source>
          ,
          <year>2022</year>
          . URL:
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>