<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>When You Doubt, Abstain: A Study of Automated Fact-checking in Italian Under Domain Shift</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Valer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan Ramponi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Tonelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler (FBK), Digital Humanities Unit - Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Trento, Department of Information Engineering and Computer Science - Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Data for building fact-checking models for Italian is scarce, often contains ambiguous claims, and lacks textual diversity. This makes it hard to reliably apply such tools in the real world to support fact-checkers' work. In this paper, we propose a categorization of claim ambiguity and label the largest Italian test set based on it. Moreover, we create challenge sets across two axes of variation: genres and fact-checking sources. Our experiments using transformer-based semantic search show a large drop in performance under domain shift, and indicate the benefit of models' abstention in case of lacking evidence.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Automated fact-checking</kwd>
        <kwd>claim ambiguity</kwd>
        <kwd>domain shift</kwd>
        <kwd>models' abstention</kwd>
        <kwd>semantic search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>when applied on texts reflecting diferent genres (e.g.,
from news headlines to posts on social media).</p>
      <p>
        Countering the spread of mis/disinformation is one of the In this paper, we aim to advance automated
factmajor challenges of our society, but human fact-checkers checking in Italian by examining claim ambiguity in the
struggle to cope with the increasing amount of content largest, publicly-available test set to date, and providing
being published. On these bases, in recent years auto- means to measure and mitigate the impact of domain
mated fact-checking has gained increasing attention in shift along genres and sources dimensions of variation.
the NLP community, resulting in a significant body of Our study shows that automated fact-checking in Italian
works and initiatives, e.g., the Fact Extraction and VERi- is still far from being reliably applied in the real-world,
ifcation Workshop (FEVER), at its 7th edition in 2023 [
        <xref ref-type="bibr" rid="ref15">1</xref>
        ]. and indicates the benefit of models’ abstention in case of
      </p>
      <p>Research eforts in NLP for automated fact-checking lacking evidence for verification.
span over a plurality of tasks, from claim detection to
verdict prediction and justification production [ 2]. Never- Contributions i) We propose a categorization of claim
theless, languages other than English, one of them being ambiguity, ii) annotate the Italian test portions of
XItalian, are mostly overlooked in current NLP research on Fact according to it, and iii) create challenge test sets
the topic. Specifically, little work has been done to build for studying automated fact-checking in Italian under
doannotated corpora for the Italian language, which is cur- main shift. We further iv) assess performance shift using
rently included in just a handful of multilingual datasets, transformer-based semantic search, v) highlighting the
i.e., X-Fact [3] and FakeCovid [4]. To exacerbate the benefit of abstention in the case of insuficient evidence.
problem, most datasets for automated fact-checking not
only include underspecified claims for which verdicts are
hard-to-impossible to be determined [5], but also typi- 2. Fact-checking Data
cally lack domain diversity, making it dificult to
ascertain the reliability of the resulting fact-checking systems</p>
      <sec id="sec-1-1">
        <title>Among the fact-checking datasets comprising Italian, we</title>
        <p>select X-Fact [3] for our study since it represents a more
diversified set of topics and comprises a larger amount
of claims in the Italian language than FakeCovid [4].</p>
        <p>X-Fact contains 31,189 non-English textual claims
from 25 languages, among which are 1,513 Italian claims
based on Pagella Politica (PP)1 and Agenzia Giornalistica</p>
      </sec>
      <sec id="sec-1-2">
        <title>1Pagella Politica website: https://pagellapolitica.it/</title>
        <p>Italia (AGI)2 fact-checks. The original veracity labels for the Italian in-domain and out-of-domain test sets of
Xthe claims derive from diferent sources in multiple lan- Fact accordingly, expanding the observed causes of
amguages, and therefore have been homogenized by Gupta biguity beyond e.g., underspecification due to ill-defined
and Srikumar (2021) [3] to a fixed label set, i.e., true, terms [6] and pronouns [7].
mostly-true, partly-true, mostly-false, false, as well as
complicated for cases whose original labels have been found Reasons for claim ambiguity The reasons why a
hard to be mapped to the proposed label set.3 claim may be ambiguous are identified based on a
pre</p>
        <p>The data is structured into training, development, and liminary assessment of the test portions of X-Fact and
test splits. Both the training and development sets have of past literature [6, 7]. In the following, we provide
been extracted from PP and include 943 and 125 claims, ambiguity classes ordered by decreasing severity and
respectively. The test set instead comprises an in-domain accompanied by definitions and examples:
portion (190 claims from PP) and an out-of-domain one
(255 claims from AGI). We remove instances marked as 1. Missing information: the claim does not
concomplicated by Gupta and Srikumar (2021) [3] from the tain information that calls for verification:
test set since they do not provide any information about “Di Battista e la guerra in Afghanistan.” [En:
claim veracity. As a result, while the in-domain test por- “Di Battista and the war in Afghanistan.” ]:
tion remains the same (i.e., 190 claims), the size of the mostly-true
out-of-domain test portion decreases to 160 claims due to
the filtering of 95 claims (i.e., 37.3%).</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Annotation and Challenge Sets</title>
      <sec id="sec-2-1">
        <title>In this section, we present our proposed categorization of</title>
        <p>claim ambiguity and describe the annotation process of
the Italian test sets of X-Fact according to it (Section 3.1).
Moreover, we detail the creation of challenge test sets
aimed at studying the performance of automated
factchecking in Italian under domain shift (Section 3.2).</p>
        <sec id="sec-2-1-1">
          <title>3.1. Claim Categorization</title>
          <p>Textual claims that undergo fact-checking may be hard
or even impossible to verify due to ambiguity and
underspecified language. When it comes to datasets for
automated fact-checking, the additional context that has
been used by fact-checkers is typically not included,
making labels for claims in such decontextualized conditions
to change from concrete verdicts (e.g., true, false) to being
unverifiable [ 5]. Moreover, claims and associated verdict
labels that are derived from fact-checking websites and
included in most datasets may cause further ambiguity.
Indeed, the claim often corresponds to the headline of
the article describing the statement that has been
verified, but the verdict label typically refers to the latter,
and thus the claim-label pair may not match the original
statement-label association (see the “Discordant label”
ambiguity class described further on).</p>
          <p>In this section, we provide a categorization of the
reasons why a claim may be ambiguous4 and annotate
2Agenzia Giornalistica Italia website: https://www.agi.it/
3We leave out the label other from our discussion since it is present
only in some non-Italian subsets which are not part of this study.
4In the remainder of this paper, we use “ambiguity” as a broad term
that also includes underspecified language.
2. Lack of context: the claim does not provide
enough context (e.g., who, when, and where) or
contains ill-defined terms and pronouns, and thus
can not be unambiguously verified:
“Siamo al nono mese consecutivo di riduzione
degli sbarchi.” [En: “We are in the ninth
consecutive month of reduced arrivals by sea.” ]: true
3. Discordant label: the fact-checked statement
has been rewritten in a negated form or as its
opposite, but the label reflects the veracity of the
original statement:
“No, la Banca d’Italia non è controllata dalle
banche private.” [En: “No, the Bank of Italy is
not controlled by private banks.” ]: partly-true
4. Claim as question: the fact-checked statement
has been rewritten as a question. Although it
may preserve the information necessary for
factchecking purposes, this alteration does not
represent an actual claim:
“Davvero la triplice sede del Parlamento
europeo costa oltre 200 milioni di euro l’anno?”
[En: “Does the triple seat of the European
Parliament really cost over 200 million euros per
year?” ]: partly-true
5. No ambiguity: the claim is unambiguous and
therefore presents suficient information for
automated fact-checking purposes:5
“In Italia ci sono 18 milioni di persone a rischio
povertà.” [En: “In Italy there are 18 million
people at risk of poverty.” ]: true</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>5Note that real-world facts and the subsequent claims are often time-,</title>
        <p>space-, and culture-dependent. We relax the “perfect unambiguity”
requirement in the context of this work.
Annotating claim ambiguity We manually annotate
claims for ambiguity on both Italian test sets of X-Fact
following the proposed categories. We focus our eforts
on test portions since these represent data that should be
used to reliably assess automated fact-checking systems.
This results in 350 annotated claims (i.e., 190 in-domain
and 160 out-of-domain, cf. Section 2). If more than one
ambiguity class is applicable for a given claim, the more
severe one is chosen. For instance, if a claim falls under
both “claim as question” and “lack of context” categories,
then the latter is applied. Annotation is carried out by
a native speaker of Italian. Since the ambiguity classes
are rather straightforward, no double annotation was
performed. The distribution of annotated claims among
classes and test set portions is presented in Table 1.</p>
        <sec id="sec-2-2-1">
          <title>3.2. Creation of Challenge Test Sets</title>
          <p>A typical assumption in most machine learning
algorithms is that training and test data follow the same
underlying distribution [8]. This is reflected by datasets in
which diversity in textual types is rather limited, which
makes it hard to assess the performance of automated
fact-checking into the wild, such as under genre shift (i.e.,
from article headlines to user-generated content on social
media). Although X-Fact includes in-domain and
out-ofdomain sets, these mainly reflect diferent fact-checking
sources rather than textual genres.</p>
          <p>To provide the research community with means to
investigate and mitigate the impact of genre shift on
automated fact-checking in Italian, we extend X-Fact with
new challenge test sets. We rewrite the subset of claims
from the in-domain and out-of-domain Italian test sets
which exhibit suficient information for fact-checking
purposes (i.e., those in bold in Table 1, namely “claim as
question” and “no ambiguity”, totalling 117 claims for the
in-domain test set and 137 claims for the out-of-domain
• News-like: “Il M5S si conferma una delle
principali forze politiche in Europa, secondo Di Maio.”
[En: “The M5S is confirmed as one of the main
political forces in Europe, according to Di Maio.” ]
• Social-like: “Il #M5S è traa i partiti maggiori
d’Europa!!!” [En: “The #M5S is amongg the major
parties in Europe!!!” ]</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>As a result, two in-domain (nl and sl, 117 claims each)</title>
        <p>and two out-of-domain (nl and sl, 137 claims each) test
sets are created as two variants of the original test sets.
Detailed statistics for all the subsets are in Table 2.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experiments</title>
      <sec id="sec-3-1">
        <title>In this section, we present the setup for our experiments (Section 4.1), details on model selection (Section 4.2), and further provide a discussion of the results (Section 4.3) and an error analysis (Section 4.4).</title>
        <sec id="sec-3-1-1">
          <title>4.1. Experimental Setup</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>We conduct experiments on automated fact-checking</title>
        <p>along two axes of variation: source (in-domain vs
outof-domain) and genre (news-like vs social-like), using the
data splits presented in Table 2.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Method All experiments employ a semantic search</title>
        <p>
          method for evidence retrieval based on
SentenceBERT [9], followed by majority-driven veracity
classification. Compared to standard sentence classification, e.g.,
using BERT [
          <xref ref-type="bibr" rid="ref20">10</xref>
          ], this makes automated fact-checking
more transparent, since instances used to determine the
veracity label of input claims can be further inspected
and ultimately shown to the end user.
        </p>
        <p>Formally, given an input claim  ∈  , where  is a test
set among test-(nl|sl)(|) (cf. Table 2, bottom), the
Id
train
dev
test
test
test-nl
test-nl
test-sl
test-sl</p>
        <p>Subset description</p>
        <p>Source
Training set
Development set
Test set (in-domain)
Test set (out-of-domain)
test as news-like
test as news-like
test as social-like
test as social-like</p>
        <p>PP
PP
PP
AGI
PP
AGI
PP
AGI</p>
        <p>Genre
nl
nl
nl
nl
nl
nl
sl
sl</p>
        <p>
          T
goal is to find the most relevant claim(s) {1, ..., } ∈ , for reference. Specifically, the latter does not include
 ≤ | |, where  is the evidence set (i.e., union of test(|) as part of the evidence set .
train, dev, and test(|);6 cf. Table 2, top), and 
represents the maximum number of most similar claims Label mapping Since test and its challenge test
to retrieve from it. In order to retrieve such evidence sets have only true and false classes, all dataset labels
claims, 1, ..., | | and 1, ..., || are all assigned an em- have been mapped as follows: {true, mostly-true} → true,
bedding  and  , respectively, using a pre-trained {partly-true,8 mostly-false, false} → false.
multilingual SentenceTransformers model7 with
default hyperparameters [11]. Then, the cosine similar- Metrics We use macro F1 score for evaluation to
acity between the input claim embedding  and those count for the unbalanced class distribution in the test
of all evidence claims 1 , ..., || is computed, i.e., sets (cf. ood ones, Table 2). When computing the F1 score,
( ,  ),  = 1, ..., ||. If ( ,  ) &gt;  , abstention is counted as a wrong prediction, because on
where  is a similarity threshold in [
          <xref ref-type="bibr" rid="ref15">0, 1</xref>
          ], the claim  is a controlled setup the model has access to relevant
eva candidate for determining the veracity label of . All idence and thus should not abstain. We also measure
candidate claims are sorted by similarity score and the the correct (cor), error (err), and abstention (abs) rates,
most recurring label among the top  evidence claims by respectively counting the cases in which the model
is finally assigned to the input claim . If no evidence correctly or wrongly predicts a veracity label, or abstains.
claim is found or there is a tie among label counts from
retrieved claims, then the model abstains. We believe that
the possibility to abstain, rather than forcibly assigning 4.2. Model Selection
a label, is highly desirable in real-world scenarios, since Our method depends on two hyperparameters: the
maxit is not always possible to assess the veracity of a claim. imum number  of evidence claims to be retrieved and
the threshold  . We tested values of  in the search space
Settings In order to isolate the impact on performance {1, 2, 3, 4, 5} across all axes of variation (i.e., data splits
of genres and sources from the actual availability of rele- in Table 2, bottom) and thresholds  ∈ [0.30, 0.85] (with
vant evidence (i.e., verified claims about the input claim’s step 0.05),9 finding that  = 1 gives on average the best
topic), we mainly focus on experiments in a controlled macro F1 across all configurations (cf. Figure 1). 10 As a
setting. This ensures that information for verification of result, we use  = 1 in the rest of this paper, and present
each input claim is available in the evidence set .
Nevertheless, we also present results in a non-controlled setting
6This makes it sure that veracity information for input claims is
actually available, thus allowing us to study the impact of sources
and genres in a controlled setting.
7We use the paraphrase-multilingual-MiniLM-L12-v2
multilingual model since preliminary experiments with Italian models
resulted in worse performance. We hypothesize this is due to the
pre-training data composition and size used by the latter.
8Indeed, the partly-true label is used in PP for claims that are wrong
but based on a grain of truth.
9The range is motivated by preliminary experiments: we found that
values  &lt; 0.30 and  &gt; 0.85 are not informative, since the
method retrieves almost all or no claims, respectively.
10Interestingly, we observe that  = 2 and  = 4 values result
to low F1 scores. This is because retrieving an even number of
evidence claims leads to a higher probability of abstention, as true
and false evidence claims may be in equal number, and abstention
is considered as an error for the F1 score in the controlled setting.
        </p>
        <p>Abstention helps in reducing errors When the
model abstains (considering  = 1), there are no
instances in  that are “similar enough” (i.e.,  ) to the
input claim. Intuitively, this reduces the impact of
erroall  values in the aforementioned range for discussing neous predictions “when in doubt”. Figure 2 (dashed lines)
the trade-of between errors and abstention. provides insights into the impact of abstention on
formerly cor and err percentages across all configurations,
4.3. Results and Discussion as  varies. We can see that up to  ∼= 0.6, abstention
has the great advantage of reducing err while negligibly
We present the results across sources, genres, and setups impacting cor. Even in the most challenging test set (i.e.,
in Figure 2, highlighting the trade-of between abstention test-sl, Figure 2d), cor predictions for  = 0.6 are
and correct/wrong veracity label predictions. more than double (i.e., 69%, Table 3) than the incorrect
ones (i.e., 31% err+abs, Table 3). The trade-of between
Genre shift has a large impact on performance By reducing err and preserving cor becomes evident with
comparing results on test-nl with those of test-sl  ≳ 0.7, for which abstention comes mainly at the
ex(Figure 2a and 2b) and results on test-nl with that of pense of cor. In the non-controlled setting (Figure 2,
test-sl (Figure 2c and 2d), we see a substantial drop bottom), on the other hand, this aspect is hard to assess
in cor and an increase in err on social-like test sets. We due to spurious factors. By looking at Table 3, we can
present selected results for  = 0.6 in Table 3, i.e., the also see that error rates on ood sets compared to id ones
threshold for which, on average, the abs ratio still has a do not increase, and actually moderately decrease on the
higher impact on err than cor before mainly impacting nl genre (i.e., from 0.10 to 0.08).
cor (cf. Figure 2). The F1 scores largely drop from 0.86
to 0.74 and from 0.82 to 0.67 when testing the model on
data derived from the same or a diferent fact-checking 4.4. Error Analysis
source, respectively. Such findings attest the impact of
genres on the performance in case of available evidence
for veracity prediction. This is confirmed in the
noncontrolled setup (results not shown for brevity), albeit
the F1 score exhibits a smaller drop due to confounding
reasons such as the lack of relevant evidence.</p>
        <p>We collect abs and err predictions across test sets in the
controlled setting (with  = 1,  = 0.6) and perform
a manual analysis. As shown in Table 4, 33.0% has the
correct evidence retrieved at first ( ( ,  ) &gt;  ),
but this is later discarded because wrong evidence has
higher similarity to the input. In the remaining cases,
the method fails to retrieve the correct evidence, and
Fact-checking sources do matter, too By looking at thus either wrongly predicts the label (22.9%) or abstains
the results on test-nl (Figure 2c) and test-sl (Fig- (44.0%). Among the 55.9% (61) err only, 13.1% (8) is
ure 2d), we can observe that not only the cor percentages actually based on a correctly-retrieved relevant claim, but
drop earlier compared to the in-domain source counter- because of claim ambiguity in train and dev sets (e.g.,
parts (i.e., Figure 2a and 2b), but also that errors accumu- discordant labels), the prediction is wrong. In particular,
late in the presence of multiple dimensions of variation, such case accounts for 22.2% (8 out of 36) of errors with
i.e., source and genre (cf. Figure 2d vs Figure 2b). The evidence. This gives a concrete measure of the impact of
F1 score drops from 0.86 to 0.82 and from 0.74 to 0.67 on ambiguity on the automated fact-checking task.
nl and sl genres, respectively (Table 3). This is again As regards the performance shift on social-like sets
confirmed in the non-controlled setting (Figure 2, bottom). compared to news-like ones, we observe that this is
(a) Results on test-nl.</p>
        <p>(b) Results on test-sl.</p>
        <p>(c) Results on test-nl.</p>
        <p>(d) Results on test-sl.
(e) Results on test-nl.</p>
        <p>(f) Results on test-sl.</p>
        <p>(g) Results on test-nl.</p>
        <p>(h) Results on test-sl.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Related Work</title>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <sec id="sec-5-1">
        <title>This work has received financial support from the European Union’s Horizon Europe research and innovation program under grant agreement No. 101070190 (AI4Trust).</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion References</title>
      <sec id="sec-6-1">
        <title>In this work, we show that domains do have a large</title>
        <p>impact on performance of automated fact-checking for
Italian, and the faculty of abstention may be considered
to cope with lack of suficient evidence. Moreover, we
contribute to the community by classifying claim
ambiguity in the largest Italian test set to date and distributing
Italian challenge test sets reflecting diversified domains.
Future work includes complementing challenge sets with
further versions by multiple annotators as well as
automating claim ambiguity assessment. Moreover, the
confidence level of the classifier could be investigated and
measured with tailored metrics to improve automated
fact-checking reliability in handling uncertain cases.</p>
        <p>In general, as suggested by Schlichtkrull et al. (2023)
[16], it would be important to assess the system
eficacy with its intended users, in order to evaluate any
unforeseen harm possibly caused by actual applications
of the technology. In the future, we therefore plan to
test our system by including it in the workflow adopted
by professional fact-checkers to verify possible cases of
mis/disinformation.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Linguistics</surname>
          </string-name>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>4512</fpage>
          -
          <lpage>4525</lpage>
          . URL: https:
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          //aclanthology.org/
          <year>2020</year>
          .emnlp-main.
          <volume>365</volume>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <volume>18653</volume>
          /v1/
          <year>2020</year>
          .emnlp-main.
          <volume>365</volume>
          . [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Carrella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miani</surname>
          </string-name>
          , S. Lewandowsky, IRMA: the
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>335-million-word Italian coRpus for studying Mis-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          informAtion,
          <source>in: Proceedings of the 17th Con-</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>tia</surname>
          </string-name>
          ,
          <year>2023</year>
          , pp.
          <fpage>2339</fpage>
          -
          <lpage>2349</lpage>
          . URL: https://aclanthology.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          org/
          <year>2023</year>
          .eacl-main.
          <volume>171</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          eacl-main.
          <volume>171</volume>
          . [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babulkov</surname>
          </string-name>
          , G. Da San Martino, P. Nakov,
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>checked claims</article-title>
          ,
          <source>in: Proceedings of the 58th An-</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Linguistics</surname>
          </string-name>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>3607</fpage>
          -
          <lpage>3618</lpage>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>332</volume>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <volume>18653</volume>
          /v1/
          <year>2020</year>
          .acl-main.
          <volume>332</volume>
          . [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hardalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chernyavskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Koychev</surname>
          </string-name>
          , D. Il-
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>Proceedings of the 2nd Conference of the Asia-</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>tional Linguistics and the 12th International Joint</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>ume 1: Long Papers)</article-title>
          , Association for Computa-
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>tional Linguistics</surname>
          </string-name>
          , Online only,
          <year>2022</year>
          , pp.
          <fpage>266</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          URL: https://aclanthology.org/
          <year>2022</year>
          .aacl-main.
          <volume>22</volume>
          . [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Simonsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lioma</surname>
          </string-name>
          , I. Au-
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>for Computational Linguistics</source>
          <volume>10</volume>
          (
          <year>2022</year>
          )
          <fpage>746</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          763. URL: https://aclanthology.org/
          <year>2022</year>
          .tacl-
          <volume>1</volume>
          .
          <fpage>43</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>doi:10</source>
          .1162/tacl_a_
          <fpage>00486</fpage>
          . [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schlichtkrull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ousidhoum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          , The
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>arXiv:2304.14238</source>
          (
          <year>2023</year>
          ). URL: https://arxiv.org/abs/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>