<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How To Build Competitive Multi-gender Speech Translation Models For Controlling Speaker Gender Translation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Gaido</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dennis Fucci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Negri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luisa Bentivogli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Trento</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>When translating from notional gender languages (e.g., English) into grammatical gender languages (e.g., Italian), the generated translation requires explicit gender assignments for various words, including those referring to the speaker. When the source sentence does not convey the speaker's gender, speech translation (ST) models either rely on the possibly-misleading vocal traits of the speaker or default to the masculine gender, the most frequent in existing training corpora. To avoid such biased and not inclusive behaviors, the gender assignment of speaker-related expressions should be guided by externally-provided metadata about the speaker's gender.1 While previous work has shown that the most efective solution is represented by separate, dedicated gender-specific models, the goal of this paper is to achieve the same results by integrating the speaker's gender metadata into a single “multi-gender” neural ST model, easier to maintain. Our experiments demonstrate that a single multi-gender model outperforms gender-specialized ones when trained from scratch (with gender accuracy gains up to 12.9 for feminine forms), while fine-tuning from existing ST models does not lead to competitive results.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;gender bias</kwd>
        <kwd>gradient reversal</kwd>
        <kwd>speech translation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>researcher – it: Sono una/un giovane ricercatrice/tore). In
this case, text-to-text machine translation (MT) models
Spurred by growing concerns about fairness in language mostly output masculine forms, while direct (or
end-totechnologies, research on understanding and mitigating end) speech-to-text translation (ST) systems partly rely
gender bias in automatic translation is gaining traction on the biological cue of the speaker’s vocal traits to
as[1]. The bias of automatic systems is extremely evident sign gender [4, 5]. However, direct ST models are still
when it comes to ambiguous sentences or expressions, largely biased toward producing masculine forms, and,
where there are no explicit cues in the source content most importantly, biological aspects are related to the
about the correct gender1 assignment of a referent (e.g., sex rather than to the gender of an individual. Hence,
en: The doctor arrived – it: Il/La dottore/essa è arrivato/a). their exploitation is not inclusive of all people, harming
In this setting, the state-of-the-art neural models often several groups such as transgenders [6].
choose the masculine forms or perpetuate stereotypical As a solution, [7] proposed to leverage external
metaassignments, as they reflect the condition statistically data about the speaker’s gender to control the gender
more likely based on their (biased) training data [2, 3]. assignment of words referred to the speaker. Specifically,</p>
      <p>This situation frequently occurs when the source lan- they investigated two approaches: i) the development
guage is genderless or employs notional gender, express- of two separate gender-specialized models, fine-tuned on
ing gender in a limited set of parts of speech, and the gender-specific data as also proposed later in MT [ 8],
target language follows a grammatical gender system, and ii) a single multi-gender model, where the speaker
embedding gender distinctions throughout a broad inven- gender is a tag fed to a single model as in multilingual
tory of parts of speech. Focusing on the case in which the systems [9]. While the second solution would be
prefersource language is English, a notional gender language, able (as the specialized solution involves the higher cost
and the target language is Italian, a grammatical gender of maintaining two separate models), the experiments
language, a frequent instance of this condition is repre- in [7] demonstrate that specialized models outperform
sented by first-person references, i.e. by the words and the multi-gender approach by a large margin in terms of
expressions referred to the speaker (e.g., en: I am a young gender accuracy.</p>
      <p>In light of the above, in this paper we address the
following research questions: i) why do specialized models
outperform multi-gender ones? ii) Can we build
competitive multi-gender systems? Through experiments on
English-Italian translation of TED talks, we show that
the low accuracy of multi-gender models comes from the
initialization with the weights of a gender-unaware ST
CLiC-it 2023: 9th Italian Conference on Computational Linguistics,
Nov 30 — Dec 02, 2023, Venice, Italy
$ mgaido@fbk.eu (M. Gaido); dfucci@fbk.eu (D. Fucci);
negri@fbk.eu (M. Negri); bentivo@fbk.eu (L. Bentivogli)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs 1IhStpTN:/c1e6u1r3-w-0sh.o7r3groCCuoEmgmUhooRnsuLWticetonhsreekApstthraibopupteiornP,4wr.0oeIncteuerensadetiiotnnhagles(CwC(CBoYEr4dU.0)R.ge-WndSer.otrog)indicate the
preferred linguistic expression of gender and not the gender identity.
system and the inability to override the behavior of the decoder where the language is represented as a tag
prebase ST model (i.e., the reliance on vocal cues) during the pended to the text [13, 14, 9]. In the case of one-to-many
ifne-tuning stage. We also try to address this problem multilingual models, this means that the &lt;bos&gt; token
with two solutions: i) a contrastive loss that penalizes the is replaced with a token that indicates the language, so
extraction of gender cues from speech input, and ii) alter- that Eq. 1 becomes:
ing vocal properties of training data to misalign gender
cues with gender tags and gender translations.2 Despite softmax(((); LID, 0, ..., − 1)) (2)
the slight improvements brought by these solutions in
gender accuracy and overall translation quality, none of
them efectively close the performance gap with the
specialized solution. However, training multi-gender models
from scratch yields competitive results, outperforming
the specialized approach with gender accuracy gains of
up to 12.9 points for feminine translations. Therefore, we
recommend building multi-gender models from scratch,
while building them on top of existing systems remains
an open research question.
where LID is the identifier of the desired target language.</p>
      <p>In direct ST, [15] demonstrated the efectiveness of
this solution, also known as “target forcing”, while [16]
proposed other methods to integrate the language
information into the architecture. Thanks to its simplicity
and efectiveness, target forcing is currently the most
widespread method to build multilingual ST systems
[17, 18], also when using large pre-trained textual models
such as mBART [19] to initialize the ST decoder [20, 21].</p>
      <p>In line with this trend, [7] obtained their best
multigender models with target forcing. As such, we build
multi-gender models using target forcing with the F and
M tags representing the two grammatical genders instead
of the language identifiers.</p>
      <sec id="sec-1-1">
        <title>3Although this paper does not aim at perpetuating a binary</title>
        <p>vision of gender, in this work we limit to the feminine and masculine
2Our code is released open source under Apache 2.0 Licence at: categories for the sake of simplicity, as the available benchmarks
https://github.com/hlt-mt/FBK-fairseq/ currently do not cover the non-binary case.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.2. Gender Information Removal</title>
        <p>In this section, we introduce the basic concepts useful
for understanding the rest of the paper. First, we provide
an overview of the methods proposed in the literature to
integrate language tags into neural multilingual
translation models (§2.1), from which multi-gender models draw
inspiration. Then, we present how gender information
has been removed from neural representations through
adversarial training in previous works (§2.2), from which
we derive our solution presented in §3.1.</p>
        <p>With the goal of fairer technology that does not rely
on spurious cues reflecting stereotypical biases in the
available data, researchers have tried to build systems
that achieve “equalized odds” among diferent
demographic groups [22]. Formally, this means that, given
an attribute  representing the belonging to one of the
 demographic groups, the predicted probability (ˆ )
2.1. Tags Integration in Multilingual of a fair system should be independent of the variable ,</p>
        <p>Models i.e. (ˆ ) = (ˆ |), ∀ ∈ . The variable  is named
State-of-the-art models in MT and ST are sequence-to- the protected attribute and in the context of gender bias
sseivqeueTnrcaensmfoordmelesrmdaedceodoefran[1e0n].coTdheer aauntdoarengaruestosirveegrdees-- lSioteirnattuhries wreoprrkeswenetcsotnhseidgeernde=ro{fth,ein}v.3olved person.
coder predicts the next-token probability over a prede- The first attempts to achieve equalized odds across
ifned vocabulary at every iteration by looking at the genders in neural systems have focused on deep neural
encoder output and at the previously generated tokens, network (DNN) classifiers [ 23, 24, 25, 26]. In this line
which are pre-pended a special token named beginning of of work, the last hidden representation of the DNN is
sentence (&lt;bos&gt;). Formally, the probability  () over passed both to a linear layer that predicts the
classificathe vocabulary  at time step  is: tion scores ˆ and to a linear layer (the discriminator)
devoted to predicting the protected attribute . The DNN
softmax(((); &lt;bos&gt;, 0, ..., − 1)) (1) is then trained in an adversarial manner [27], i.e. it is
alternatively trained to i) predict  (while keeping the
where E is the encoder, D is the decoder, X is the input shared DNN freezed) and ii) predict ˆ while minimizing
sequence, and  is the token generated at the -th time the ability to predict  (keeping the protected attribute
step. classification layer freezer). As this training procedure is</p>
        <p>While early attempts to build multilingual MT models often unstable, similar practices based on minmax
optiwere based on training dedicated encoders and decoders mizations have been proposed [28], even with
discriminafor each language [11, 12], nowadays the preferred solu- tors made of functions diferent from linear projections
tion is a model made of a single universal encoder and
[29], or using more than one discriminator [30]. [31] also function [34], whose output is averaged over the
tempoproposed methods to automatically extract the protected ral dimension to obtain a single vector representing the
attributes in case they have not been provided. logit4 of the discriminator.</p>
        <p>Such adversarial training procedures can be seen as an Furthermore, we experiment with assigning dedicated
extension of the gradient reversal [32], where the train- class weights to the loss of the discriminator, as a
countering alternatively freezes the base model (to refine the measure to the class imbalance between female and male
discriminator) and the discriminator (inverting its loss speakers in the training data. Specifically, we assigned
to train the base model to be unable to discriminate). In the weights ( , ) proportionally to the inverse of
fact, the gradient reversal layer, applied to the hidden the frequency of each class ( , ) in the training data:
representations before feeding them to the discriminator,
is an identity function in the forward pass, while inverts {︃ ∝ 1 ,  ∝ 1
ftahcetgorra d.ieBnyt innatmhienbgacktwhearhdidpdasesn, srceaplriensgenittbaytioanp,osittihvee  * +  +  * + = 1
discriminator, and  the gradient reversal layer, this resulting in  = 1.4,  = 0.8 in our case.
means that:</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. Audio Manipulation</title>
        <p>() = , ∇(()) = −  ∇()
(3)</p>
        <p>Our second approach aims to break the correlation
bewhere  is a hyperparameter that can either be fixed or tween the vocal characteristics of the speaker on one
updated according to the following equation: side and the gender tag and target translation on the
other. To this aim we manipulate part of the training
 = 1 + 2−  − 1 (4) data using the Opposite pitch manipulation strategy by
[35]. The amount of data that is manipulated at each
where  is the ratio between the number of parameter up- iteration (epoch) is controlled with a hyperparameter, ,
dates performed and the total number of updates needed which determines the probability of altering an utterance,
to complete the training. regardless of whether it is produced by a male or female
speaker. The manipulation is performed by altering two
crucial acoustic parameters distinguishing between male
3. Solutions for Multi-gender and female voices [36, 37]:  0 and formants. In
particuModels lar, we first estimate the ˜0 median of the  0 contour of
the considered speech segment. Then, we sample a new
To create multi-gender ST models that solely rely on the ˜0 ′ median of the desired output audio from a normal
gender tag, ignoring spurious cues related to speakers’ distribution whose mean and standard deviation depend
vocal traits, we test two approaches. First, we try to create on the target gender: for feminine voices, we use 250
gender-invariant encoder representations by adding a Hz as the mean and 17 as the standard deviation so that
gradient-reverted discriminator on the speaker’s gender the sampled value is between 199 Hz and 301 Hz with
(§3.1). Second, we manipulate the input audio by altering 99.7% probability; for masculine voices, the mean is 140
the speaker’s pitch, so that the correlation between the Hz and the standard deviation is 20 to obtain a 99.7%
gender tag (and output text) and the speaker’s vocal traits probability range within 80 Hz and 200 Hz. Once ˜0 and
is lost (§3.2). ˜0 ′ are defined, we compute a scaling factor  as the
ratio ˜0 ′/˜0 . Lastly, the original  0 contour is scaled by
3.1. Gradient Reversal the  factor, while the formants are scaled by 1.2 when
converting from male to female voices, or by 0.8
otherwise. This perturbation is applied independently to each
sample during each training epoch, so as to maximize
the variability of the training data.</p>
        <p>As seen in §2.1, the decoder of a multi-gender model
has three inputs: the encoder output, a tag
representing the speaker’s gender, and the previously generated
tokens. As we want the decoder to have the tag as the
only source of information about the speaker’s gender,
we propose to create encoder outputs that do not
convey any information regarding the speaker’s gender by
adding a gradient-reverted discriminator on top of the
encoder, motivated by the success of this approach in MT
with sentences where there is a single referent whose
gender has to be determined [33]. The discriminator is
made of two fully-connected layers with ReLU activation</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experimental Settings</title>
      <sec id="sec-3-1">
        <title>Our ST models are composed of a Conformer [38] en</title>
        <p>coder with 12 layers and a Transformer [10] decoder with
6 layers. We used the Conformer implementation by [39],</p>
      </sec>
      <sec id="sec-3-2">
        <title>4The logit is the vector of raw predictions before a function</title>
        <p>(commonly, the softmax) that maps it into probabilities.</p>
        <p>Gender Accuracy (↑)
1M 1F-Tag M
Specialized
Multi-gender
+ gradient reversal
+ gradient reversal weighted
+ audio manipulation (50%)
+ audio manipulation (80%)
Multi-gender
+ gradient reversal
+ gradient reversal weighted
+ audio manipulation (50%)
+ audio manipulation (80%)</p>
        <p>BLEU (↑)
which does not contain bugs related to the presence of value for which the loss on the validation set does not
padding. The embedding size was 512, and the dropout explode during training.
was set to 0.1. We optimized label-smoothed cross
entropy using Adam. The learning rate followed the Noam Audio Manipulation. In our experiments, we tested
scheduler with 25,000 warmup updates and a maximum two values (0.5 and 0.8) for the hyperparameter , which
value of 2− 3. We train for 50,000 updates and average controls the probability of manipulating a speech
segthe last 7 checkpoints. ment. In the first case, 50% of the data is manipulated,</p>
        <p>We train our models on MuST-C [40], an ST corpus leading to a complete loss of correlation between gender
built from TED data, for which is also available the an- tags and vocal traits (50% of the samples with the F tag
notation of the gender of the speaker [7]. We extract 80 would exhibit frequency characteristics typical of
mascufeatures with log mel-filterbank from the input audio and line voices, and 50% of the samples with the M tag would
normalize them with cepstral mean and variance [41]. have frequency characteristics typical of feminine voices).
The target text is encoded into subwords with 8,000 BPE In the second case, instead, the correlation between the
merge rules [42] learned on the training set. We evalu- gender tag and the vocal traits is negative, to counteract
ate on the MuST-SHE benchmark [4], which contains a the patterns learned by a gender-unaware ST model. In
section (“Category 1”) dedicated to assessing the gender any case, as the training data is imbalanced (70% of the
assignment of words referring to the speaker. We com- samples are uttered by male speakers, and 30% by female
pute SacreBLEU5 [43] on the whole MuST-SHE test set speakers), and the manipulation probability is the same
to evaluate the translation quality of our models and gen- for segments uttered by male and female speakers, the
der accuracy [7] on the feminine and masculine sections gender imbalance in the training data is not mitigated.
of “Category 1” to evaluate the ability of each model to
correctly assign gender to words referring to the speaker.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Results</title>
      <p>Gradient Reversal. The loss of the auxiliary
speakerclassification task is summed to the loss on the decoder
output scaling it by a 0.5 factor. For the gradient reversal
layer, we tested both fixed values of  and controlling
its value with  . For fine-tunings, we set  = 10, so as
to give similar weight to the gender classification loss
and the cross-entropy loss for the translation. When
training from scratch, instead, despite diferent attempts
the training is unstable and diverges unless lambda is set
to a fixed, small value, where its contribution is negligible.
We report results for  = 0.5, which is the highest 
5case:mixed|ef:no |tok:13a|smooth:exp|version:2.0.0</p>
      <sec id="sec-4-1">
        <title>We investigate the performance of multi-gender mod</title>
        <p>els trained in two diferent ways: i) fine-tuning a base,
gender-unaware ST model, and ii) training from scratch.</p>
        <p>In both cases, we study the efect of the introduction of
the discriminator with gradient reversal and of the audio
manipulation techniques. Table 1 presents BLEU and
gender accuracy scores (separately for segments spoken
by female (1F) and male (1M) speakers) for all the models,
comparing them with the specialized models. To assess
the inclusivity of our solution in cases where speakers
exhibit vocal traits that do not conform with traditional
gender perceptions, we also report gender accuracy for
tests where the gender tag is inverted compared to the weight of the F class is increased in the discriminator. In
original audio segment (1F-Tag M and 1M-Tag F). In these addition, in both cases the translation quality sufers a
instances, the gender translation is expected to align considerable drop. In the case of audio manipulation, the
with the gender tag, and we use the “wrong” reference translation quality drop is lower (although still present),
of MuST-SHE, which swaps the speaker’s references to as well as the diferences in terms of gender accuracy.
the opposite gender. They do not provide, though, any benefit compared to
the simple multi-gender training.</p>
        <p>Fine-tuning. The fine-tuned models from a base ST In summary, the results demonstrate that the
previsystem consistently yield lower scores compared to the ous finding about the low performance of multi-gender
specialized systems. The simple multi-gender model per- models is due to the adoption of a fine-tuning strategy.
forms 1.4 BLEU points worse than the specialized models In this setting, the model cannot efectively override the
in terms of overall translation quality. However, when reliance on speakers’ vocal traits of the gender-unaware
audio manipulation and especially gradient reversal tech- base ST model. In addition, techniques aimed at avoiding
niques are employed during fine-tuning, the performance the exploitation of speakers’ vocal traits seem inefective.
gap is reduced by up to half. Regarding gender accu- However, training the multi-gender model efectively
racy, the multi-gender model achieves considerably lower solves the problem and the model is capable of following
scores than the specialized models, confirming previous the indication given by the gender tag, outperforming
ifndings from [ 7]. This indicates that a fine-tuned multi- even the specialized strategy by up to 12.9 gender
accugender model struggles to accurately follow the gender racy (1M-Tag F).
tag for gender translation. The accuracy gap is
particularly high when the tag conflicts with vocal traits (-16.3 in 6. Conclusions
1F-Tag M, -9.0 in 1M-Tag F), where multi-gender models
show below-chance accuracy for feminine forms, being
below 50%. Both gradient reversal and audio
manipulation techniques seem to further bias the model towards
masculine forms. This likely indicates that the reduced
ability to rely on the speakers’ vocal traits is not
compensated by the model looking at the gender tag, rather it
strengthens its tendency to default to the most frequent
masculine forms. The only technique that consistently
improves both masculine and feminine translations
compared to the simple fine-tuned multi-gender model is the
introduction of audio manipulation with high probability
(80%). However, the gains (2.5 for 1F and 3.1 for 1M, and
4.4 for 1F-Tag M and 0.7 for 1M-Tag F) are limited and
the gap with specialized models remains large.</p>
        <p>In this paper, we studied the efect of diferent training
strategies to build multi-gender ST models, i.e. models
that are informed of the gender of the speaker by an
explicit gender tag. Focusing on English-Italian translations,
we demonstrated that the low accuracy of multi-gender
models shown by previous work stems from the their
initialization with gender-unaware ST system weights and
the inability of efectively overriding the reliance on
vocal cues during fine-tuning. On the other hand, training
multi-gender models from scratch proved to be an
efective solution, outperforming the approach based on the
creation of two gender-specialized models. As training
from scratch is not always feasible, we also experimented
with two methods to enhance the reliance on the gender
tag in fine-tuned multi-gender models: penalizing the
Training from scratch. Unlike the fine-tuned models, extraction of gender cues from speech input, and altering
the multi-gender models trained from scratch yield com- the vocal properties of the speakers in the training data
parable or even higher results than the specialized mod- to avoid the alignment between biological cues and
genels. This suggests that when an ST model is trained from der tags and translations. While these solutions partially
scratch with gender tags, it learns to efectively follow improved gender accuracy and overall translation quality
them. Specifically, the simple multi-gender model trained in fine-tuned multi-gender models, they did not close the
from scratch achieves comparable translation quality to gap with specialized models. Therefore, further research
the specialized system (-0.2 BLEU) and significantly out- is needed in this direction.
performs it in gender accuracy, with gains ranging from
0.2 (1M) to 12.5 (1F-Tag M). As in the fine-tuning case,
neither gradient reversal nor audio manipulations increase Acknowledgments
the reliance of the model on the tag and the resulting
models are more biased toward masculine forms. In fact, This work is part of the project “Bias Mitigation and
the multi-gender with gradient reversal reaches the high- Gender Neutralization Techniques for Automatic
Transest accuracies in producing masculine forms (93.2 in 1M lation”, which is financially supported by an Amazon
and 94.1 in 1F-Tag M), while sufering substantial drops Research Award AWS AI grant. We acknowledge the
in feminine accuracy. This efect is reduced when the support of the PNRR project FAIR - Future AI Research
ral Language Processing, Association for
Computational Linguistics, Online and Punta Cana,
Dominican Republic, 2021, pp. 1640–1654. URL: https:
//aclanthology.org/2021.emnlp-main.123. doi:10.</p>
        <p>18653/v1/2021.emnlp-main.123.
[9] M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu,
[1] B. Savoldi, M. Gaido, L. Bentivogli, M. Ne- Z. Chen, N. Thorat, F. Viégas, M. Wattenberg,
gri, M. Turchi, Gender bias in machine G. Corrado, M. Hughes, J. Dean, Google’s
multranslation, Transactions of the Association tilingual neural machine translation system:
Enfor Computational Linguistics 9 (2021) 845– abling zero-shot translation, Transactions of the
874. URL: https://aclanthology.org/2021.tacl-1.51. Association for Computational Linguistics 5 (2017)
doi:10.1162/tacl_a_00401. 339–351. URL: https://aclanthology.org/Q17-1024.
[2] M. O. R. Prates, P. H. C. Avelar, L. C. Lamb,
Assessing gender bias in machine translation: a case [10] dAo.i:V1a0s.w1a1n6i2,/Nt.aSchla_zae_er0,0N0.6P5a.rmar, J. Uszkoreit,
study with google translate, Neural Computing and L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin,
Applications 32 (2020) 6363–6381. doi:10.1007/ Attention is all you need, in: I. Guyon, U. V.
s00521-019-04144-6. Luxburg, S. Bengio, H. Wallach, R. Fergus,
[3] W. I. Cho, J. W. Kim, S. M. Kim, N. S. Kim, On S. Vishwanathan, R. Garnett (Eds.),
ProceedMeasuring Gender bias in Translation of Gender- ings of the 31st International Conference on
neutral Pronouns, in: Proceedings of the First Work- Neural Information Processing Systems, Curran
shop on Gender Bias in Natural Language Process- Associates Inc., Long Beach, USA, 2017. URL:
ing, Association for Computational Linguistics, Flo- https://proceedings.neurips.cc/paper/2017/file/
rence, Italy, 2019, pp. 173–181. URL: https://www. 3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
aclweb.org/anthology/W19-3824. doi:10.18653/ [11] O. Firat, K. Cho, Y. Bengio, Multi-way,
multilinv1/W19-3824. gual neural machine translation with a shared
at[4] L. Bentivogli, B. Savoldi, M. Negri, M. A. Di Gangi, tention mechanism, in: Proceedings of the 2016
R. Cattoni, M. Turchi, Gender in Danger? Evaluat- Conference of the North American Chapter of the
ing Speech Translation Technology on the MuST- Association for Computational Linguistics: Human
SHE Corpus, in: Proceedings of the 58th Annual Language Technologies, Association for
ComputaMeeting of the Association for Computational Lin- tional Linguistics, San Diego, California, 2016, pp.
guistics, Association for Computational Linguis- 866–875. URL: https://aclanthology.org/N16-1101.
tics, Online, 2020, pp. 6923–6933. URL: https://www.</p>
        <p>aclweb.org/anthology/2020.acl-main.619. [12] dOo. iF:1ir0at.,1B8. 6Sa5n3k/avr1an/,NY1.6A-l-1o1n0a1iz.an, F. T. Yarman
Vu[5] M. Gaido, M. A. Di Gangi, M. Negri, M. Turchi, On ral, K. Cho, Zero-resource translation with
multiKnowledge Distillation for Direct Speech Transla- lingual neural machine translation, in:
Proceedtion, in: Proceedings of CLiC-IT 2020, Online, 2021. ings of the 2016 Conference on Empirical Methods
URL: http://ceur-ws.org/Vol-2769/paper_28.pdf. in Natural Language Processing, Association for
[6] L. Zimman, Transgender language, transgen- Computational Linguistics, Austin, Texas, 2016, pp.
der moment: Toward a trans linguistics, in: 268–277. URL: https://aclanthology.org/D16-1026.
K. Hall, R. Barrett (Eds.), The Oxford Handbook
of Language and Sexuality, 2020. doi:10.1093/ [13] dRo.iS:1e0n.n1r8ic6h5, 3B/.vH1/aDd1d6o-w1,0A2.6.Birch, Improving
oxfordhb/9780190212926.013.45. Neural Machine Translation Models with
Mono[7] M. Gaido, B. Savoldi, L. Bentivogli, M. Negri, lingual Data, in: Proceedings of the 54th Annual
M. Turchi, Breeding Gender-aware Direct Speech Meeting of the Association for Computational
LinTranslation Systems, in: Proceedings of the guistics (Volume 1: Long Papers), Association for
28th International Conference on Computational Computational Linguistics, Berlin, Germany, 2016,
Linguistics, International Committee on Compu- pp. 86–96. URL: https://aclanthology.org/P16-1009.
tational Linguistics, Barcelona, Spain (Online),
2020, pp. 3951–3964. URL: https://www.aclweb.org/ [14] dTo.-iL:1.H0.a1,J8. 6N5ie3h/uve1s/,AP1.W6-a1ib0e0l,9T.oward multilingual
anthology/2020.coling-main.350. doi:10.18653/ neural machine translation with universal encoder
v1/2020.coling-main.350. and decoder, in: Proceedings of the 13th
Interna[8] P. K. Choubey, A. Currey, P. Mathur, G. Dinu, tional Conference on Spoken Language
TranslaGFST: Gender-filtered self-training for more accu- tion, International Workshop on Spoken Language
rate gender in translation, in: Proceedings of the Translation, Seattle, Washington D.C, 2016. URL:
2021 Conference on Empirical Methods in Natu- https://aclanthology.org/2016.iwslt-1.6.
[15] H. Inaguma, K. Duh, T. Kawahara, S. Watanabe, ciates Inc., Red Hook, NY, USA, 2016, p. 3323–3331.</p>
        <p>Multilingual end-to-end speech translation, in: [23] A. Beutel, E. H. Chi, J. Chen, Z. Zhao, Data
de2019 IEEE Automatic Speech Recognition and Un- cisions and theoretical implications when
adverderstanding Workshop (ASRU), 2019, pp. 570–577. sarially learning fair representations, 2017. URL:
doi:10.1109/ASRU46091.2019.9003832. https://arxiv.org/pdf/1707.00075.pdf.
[16] M. A. Di Gangi, M. Negri, M. Turchi, One-to-many [24] B. H. Zhang, B. Lemoine, M. Mitchell, Mitigating
multilingual end-to-end speech translation, in: unwanted biases with adversarial learning, in:
Pro2019 IEEE Automatic Speech Recognition and Un- ceedings of the 2018 AAAI/ACM Conference on
derstanding Workshop (ASRU), 2019, pp. 585–592. AI, Ethics, and Society, AIES ’18, Association for
doi:10.1109/ASRU46091.2019.9004003. Computing Machinery, New York, NY, USA, 2018,
[17] C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, p. 335–340. URL: https://doi.org/10.1145/3278721.</p>
        <p>J. Pino, Fairseq S2T: Fast speech-to-text model- 3278779. doi:10.1145/3278721.3278779.
ing with fairseq, in: Proceedings of the 1st Con- [25] Y. Elazar, Y. Goldberg, Adversarial removal of
deference of the Asia-Pacific Chapter of the Associa- mographic attributes from text data, in:
Proceedtion for Computational Linguistics and the 10th ings of the 2018 Conference on Empirical Methods
International Joint Conference on Natural Lan- in Natural Language Processing, Association for
guage Processing: System Demonstrations, Associ- Computational Linguistics, Brussels, Belgium, 2018,
ation for Computational Linguistics, Suzhou, China, pp. 11–21. URL: https://aclanthology.org/D18-1002.
2020, pp. 33–39. URL: https://aclanthology.org/2020. doi:10.18653/v1/D18-1002.</p>
        <p>aacl-demo.6. [26] L. Gao, H. Zhan, A. Chen, V. Sheng, Mitigate gender
[18] E. Salesky, M. Wiesner, J. Bremerman, R. Cat- bias using negative multi-task learning, 2022. URL:
toni, M. Negri, M. Turchi, D. W. Oard, M. Post, https://doi.org/10.21203/rs.3.rs-2024101/v1. doi:10.
The Multilingual TEDx Corpus for Speech Recog- 21203/rs.3.rs-2024101/v1.
nition and Translation, in: Proc. Inter- [27] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu,
speech 2021, 2021, pp. 3655–3659. doi:10.21437/ D. Warde-Farley, S. Ozair, A. Courville, Y.
BenInterspeech.2021-11. gio, Generative adversarial networks, Commun.
[19] Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, ACM 63 (2020) 139–144. URL: https://doi.org/10.</p>
        <p>M. Ghazvininejad, M. Lewis, L. Zettlemoyer, Mul- 1145/3422622. doi:10.1145/3422622.
tilingual denoising pre-training for neural ma- [28] S. Ravfogel, M. Twiton, Y. Goldberg, R. D.
Cotchine translation, Transactions of the Associa- terell, Linear adversarial concept erasure, in:
tion for Computational Linguistics 8 (2020) 726– K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari,
742. URL: https://aclanthology.org/2020.tacl-1.47. G. Niu, S. Sabato (Eds.), Proceedings of the 39th
Indoi:10.1162/tacl_a_00343. ternational Conference on Machine Learning,
vol[20] D. Liu, T. Binh Nguyen, S. Koneru, E. Yavuz Ugan, ume 162 of Proceedings of Machine Learning
ReN.-Q. Pham, T. Nam Nguyen, T. Anh Dinh, search, PMLR, 2022, pp. 18400–18421. URL: https:
C. Mullov, A. Waibel, J. Niehues, KIT’s multilin- //proceedings.mlr.press/v162/ravfogel22a.html.
gual speech translation system for IWSLT 2023, in: [29] S. Ravfogel, F. Vargas, Y. Goldberg, R. Cotterell,
Proceedings of the 20th International Conference Adversarial concept erasure in kernel space, in:
on Spoken Language Translation (IWSLT 2023), As- Proceedings of the 2022 Conference on Empirical
sociation for Computational Linguistics, Toronto, Methods in Natural Language Processing,
AssociCanada (in-person and online), 2023, pp. 113–122. ation for Computational Linguistics, Abu Dhabi,
URL: https://aclanthology.org/2023.iwslt-1.6. United Arab Emirates, 2022, pp. 6034–6055. URL:
[21] E. Gow-Smith, A. Berard, M. Zanon Boito, https://aclanthology.org/2022.emnlp-main.405.</p>
        <p>I. Calapodescu, NAVER LABS Europe’s multilingual [30] X. Han, T. Baldwin, T. Cohn, Diverse
adverspeech translation systems for the IWSLT 2023 low- saries for mitigating bias in training, in:
Proresource track, in: Proceedings of the 20th Interna- ceedings of the 16th Conference of the European
tional Conference on Spoken Language Translation Chapter of the Association for Computational
Lin(IWSLT 2023), Association for Computational Lin- guistics: Main Volume, Association for
Compuguistics, Toronto, Canada (in-person and online), tational Linguistics, Online, 2021, pp. 2760–2765.
2023, pp. 144–158. URL: https://aclanthology.org/ URL: https://aclanthology.org/2021.eacl-main.239.
2023.iwslt-1.10. doi:10.18653/v1/2021.eacl-main.239.
[22] M. Hardt, E. Price, N. Srebro, Equality of oppor- [31] S. Shao, Y. Ziser, S. B. Cohen, Erasure of unaligned
tunity in supervised learning, in: Proceedings of attributes from neural representations,
Transacthe 30th International Conference on Neural Infor- tions of the Association for Computational
Linguismation Processing Systems, NIPS’16, Curran Asso- tics 11 (2023) 488–510. URL: https://aclanthology.
org/2023.tacl-1.29. doi:10.1162/tacl_a_00558. recognition, Speech Communication 25 (1998) 133–
[32] Y. Ganin, V. Lempitsky, Unsupervised Domain 147. URL: https://www.sciencedirect.com/science/
Adaptation by Backpropagation, in: F. Bach, article/pii/S0167639398000338. doi:https://doi.
D. Blei (Eds.), Proceedings of the 32nd International org/10.1016/S0167-6393(98)00033-8.
Conference on Machine Learning, volume 37 of [42] M. A. Di Gangi, M. Gaido, M. Negri, M. Turchi, On
Proceedings of Machine Learning Research, PMLR, target segmentation for direct speech translation,
Lille, France, 2015, pp. 1180–1189. URL: https:// in: Proceedings of the 14th Conference of the
Asproceedings.mlr.press/v37/ganin15.html. sociation for Machine Translation in the Americas
[33] E. Fleisig, C. Fellbaum, Mitigating Gender Bias in (Volume 1: Research Track), Association for
MaMachine Translation through Adversarial Learning, chine Translation in the Americas, Virtual, 2020,
2022. arXiv:2203.10675. pp. 137–150. URL: https://aclanthology.org/2020.
[34] V. Nair, G. E. Hinton, Rectified linear units improve amta-research.13.</p>
        <p>restricted boltzmann machines, in: Proceedings [43] M. Post, A call for clarity in reporting BLEU scores,
of the 27th International Conference on Interna- in: Proceedings of the Third Conference on
Mational Conference on Machine Learning, ICML’10, chine Translation: Research Papers, Association
Omnipress, Madison, WI, USA, 2010, p. 807–814. for Computational Linguistics, Belgium, Brussels,
[35] D. Fucci, M. Gaido, M. Negri, M. Cettolo, L. Ben- 2018, pp. 186–191. URL: https://www.aclweb.org/
tivogli, No Pitch Left Behind: Addressing Gen- anthology/W18-6319.
der Unbalance in Automatic Speech Recognition
through Pitch Manipulation, in: IEEE Automatic
Speech Recognition and Understanding Workshop
(ASRU), Taipei, Taiwan, 2023.
[36] R. O. Coleman, A comparison of the
contributions of two voice quality characteristics to the
perception of maleness and femaleness in the
voice, Journal of Speech &amp; Hearing Research
19(1) (1976) 168–180. doi:https://doi.org/10.</p>
        <p>1044/jshr.1901.168.
[37] J. M. Hillenbrand, M. J. Clark, The role of f0 and
formant frequencies in distinguishing the voices
of men and women, Attention Perception &amp;
Psychophysics 71(5) (2009) 1150–1166. doi:10.3758/</p>
        <p>APP.71.5.1150.
[38] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang,</p>
        <p>J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu,
R. Pang, Conformer: Convolution-augmented
Transformer for Speech Recognition, in:
Proceedings of the 21st Annual Conference of the
International Speech Communication Association,
International Speech Communication Association,
Shanghai, China (Online), 2020, pp. 5036–5040.</p>
        <p>doi:10.21437/Interspeech.2020-3015.
[39] S. Papi, M. Gaido, A. Pilzer, M. Negri, When Good
and Reproducible Results are a Giant with Feet of
Clay: The Importance of Software Quality in NLP,
2023. arXiv:2303.16166.
[40] R. Cattoni, M. A. Di Gangi, L. Bentivogli,</p>
        <p>M. Negri, M. Turchi, Must-c: A
multilingual corpus for end-to-end speech translation,
Computer Speech &amp; Language 66 (2021) 101–
155. URL: https://www.sciencedirect.com/science/
article/pii/S0885230820300887. doi:https://doi.</p>
        <p>org/10.1016/j.csl.2020.101155.
[41] O. Viikki, K. Laurila, Cepstral domain segmental
feature vector normalization for noise robust speech</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>