<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Informed Pre-Training for Critical Error Detection in English-German</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lisa Pucknat</string-name>
          <email>lisa.pucknat@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maren Pielka</string-name>
          <email>maren.pielka@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafet Sifa</string-name>
          <email>rafet.sifa@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LWDA'22: Lernen</institution>
          ,
          <addr-line>Wissen, Daten, Analysen</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Machine Translation, Quality Estimation, Critical Error Detection</institution>
          ,
          <addr-line>Informed Machine Learning</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents two data augmentation methods for pre-training, to find critical errors in machine translations. This includes an alignment approach used in traditional machine translation and an imitation method, mimicking the structure of the data. Both methods are adapted to a binary classification. Our approach achieves competitive results on the WMT'21 critical error detection (CED) dataset while only using 0.06% of datapoints in comparison to the first placement.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Finding critical errors for machine translations in the scope of quality estimation (QE) is a new
research field, introduced during the WMT’21 shared task: Quality Estimation 1. It aims to
create a supervisory system for critical translation errors independent of the translation model.
The organizers of the shared task define a critical error to fall into five categories, which are
deviation in toxicity, health- or safety risks, named entities, sentiment polarity or negation and
deviation in units/time/date/numbers. The necessity of the task is motivated by health, safety,
legal, reputation, religious or financial concerns. Generally, the task is a binary classification
on whether a sentence and its machine translation contain at least one critical error, without
a respective gold translation. In contrast to other QE tasks, translation errors are tolerated
if they are not critical. Due to the novelty of the task and the newly introduced dataset for
benchmarking, there is only limited research available [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, we present two new
approaches with data augmented pre-training. We utilize pre-training methods which were
successfully applied to machine translation and natural language inference and adapt them to
the CED task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This includes aligning the languages and mimicking the structure of the
CED dataset. For this, we rely only on a parallel corpus and lists of synonyms and antonyms,
eliminating the need for human annotators, which can be costly.
CEUR
Workshop
Proceedings
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Pre-training large language models, including multilingual settings, has led to many
state-ofthe-art results in natural language processing tasks [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ]. Diferent QE tasks also benefited
from further pre-training with artificial data that approximated the structure of the given
training dataset, starting with the common component of an initial parallel corpus [
        <xref ref-type="bibr" rid="ref10 ref11 ref8 ref9">8, 9, 10, 11</xref>
        ].
Explicitly for the CED task, the best performing approach on English-German sentences [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
is a multimetric-multilingual pre-training proposed by [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Here, large amounts of parallel
sentences in diferent languages were gathered, totaling to about 72.3 million examples. By
generating new machine translations for each sentence pair and automatically calculating common
QE metrics, a large dataset for pre-training was created. Other approaches included feature
extraction [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], incorporation of high quality machine translations of source sentences [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and
incorporation of uncertainty features into the fine-tuning process [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Regarding informed
machine learning, prior information assists in the learning process. So-called prototypes, which
are representative for the dataset, yield a beneficial efect [
        <xref ref-type="bibr" rid="ref16">16, 17</xref>
        ]. Recreating a dataset with
simpler structure, could be seen as a prototype approach.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>For the downstream task, we use the English-German split of the CED WMT’21 competition
dataset2, which is composed out of sentences from the Jigsaw Toxic Comment Classification
Challenge3 and Wikipedia comments. Translations are created by the ML50 multilingual
translation model by FAIR [18]. The dataset consists of a total of 12,000 sentences, splitting
into 10 000 from the train dataset and respectively 1000 for each valid and test set. Exemplary
sentences and translations of the dataset are displayed in table 1.</p>
      <p>For data-augmentation and pre-training, we decided to use a parallel corpus composed out of
subtitles from TED Talks, which are short presentations with innovative, semi-scientific topics.
They resemble the conversational, less complex nature of the sentences in the CED dataset. We
extract them from OPUS4, which is an open collection of parallel corpora. In its raw state, it
totals to approximately 189 000 sentence pairs. We want to incorporate as much information as
possible from the CED dataset into the creation of the new dataset, for which reason we only
select sentences with a similar sentence length. This corresponds to sentences with a maximum
length of 20 and 24 for English and German, respectively, determined by the 75% percentile of
sentence lengths from the CED dataset. We obtain structurally similar sentences by tolerating
a maximum deviation of 10% between the source and the translated sentence lengths, since
English and German sentence lengths are highly correlated [19].</p>
      <p>2https://github.com/sheffieldnlp/mlqe-pe/tree/master/data/catastrophic_errors/
3https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge
4https://opus.nlpl.eu/</p>
      <sec id="sec-3-1">
        <title>So what you’re saying is the internet police</title>
        <p>AREN’T coming to my door???
What the fuck? Ur a dick, urv never even
been 2 helensville ya mommas boi!!!
This is a fact for which reliable, published
sources do exist!</p>
      </sec>
      <sec id="sec-3-2">
        <title>Also was du sagst ist , dass die Internet- NOT polizei nicht zu meiner Tür kommen ? ? ? Was ist das ? Ur ein dick , urv noch nie 2 ERR helensville ya mommas boi ! ! !</title>
        <p>Das ist eine Tatsache , für die zuverlässige NOT
, veröfentlichte Quellen existieren !</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>
        For the alignment approach, an adaption of the pre-training method for machine translation is
utilized [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Aligning languages in the feature space might be beneficial for finding semantic
diferences [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Words of the same type within the parallel corpus (TED dataset) are exchanged
across languages, creating new sentence pairs at the same time. Only words with a 1-1
correspondence are exchanged, as we have no explicit alignment information. Pairs with a critical
error (ERR) are created by exchanging a randomly sampled translated word that does not match
the actual translation. Non-error pairs (NOT) are generated by simply exchanging the relevant
word for its translation. Augmented data pairs would be as follows:
      </p>
      <p>Take the autonomous vehicle. − Nehmen Sie das autonome Fahrzeug.</p>
      <p>NOT: Take the autonomous Fahrzeug. − Nehmen Sie das autonome vehicle
ERR: Take the autonomous Fahrzeug. − Nehmen Sie das autonome voice</p>
      <p>We also experimented with not simply choosing a random word as the critical error, but
exchanging for a word with opposite meaning.</p>
      <p>For the imitation approach, we again exchange words to augment the data but deviate from
the multilingual pre-training approach and stick only to the approach of exchanging synonyms
and antonyms in the same language. This is due to the definition of the task, which specifies
the following as the reason for a critical error.: ”Mistranslation: critical content is translated
incorrectly into a diferent meaning, or not translated”. An exemplary pair with exchanged
synonyms and antonyms could look like this:</p>
      <p>The teacher was fascinated. − Die Lehrerin war fasziniert.</p>
      <p>NOT: The lecturer was fascinated. − Die Lehrerin war fasziniert.</p>
      <p>ERR: The pupil was fascinated. − Die Lehrerin war fasziniert.</p>
      <p>For both approaches, we substitute only selected word types (e.g. nouns, verbs, adjectives),
as there are no reasonable antonyms especially for fill and stop words. Further, we reason
that critical errors appear in connection with major word types in a sentence. As there are
sentences with and without errors in the dataset, we hope that the model can determine correct
corresponding words and shift the focus away from unimportant words.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments and Results</title>
      <p>Experimental Setup First, the CED dataset was cleaned by deleting special characters and
lowercasing all caps words. For training, we use the xlm-roberta-large checkpoint5 as a
starting point, which is an XLM-RoBERTa model pre-trained on 2.5 TB filtered CommonCrawl 6
data in 100 languages. The classification layer is composed out of two feed-forward layers
 
=  (ℎ(ℎ
0 1 +  1 ) 2 +  2 )
(1)
on top, where   ∈ ℝ2 is the model’s binary prediction, ℎ0 is the last hidden state of the special
token of XLM-RoBERTa and  1 ∈ ℝ× ,  1 ∈ ℝ ,  1 ∈ ℝ×2 and  1 ∈ ℝ2 are parameters of the
feed-forward network with  = 1024 . For both pre-training and finetuning, we make use of
the AdamW optimizer [20] with  1 = 0.9,  2 = 0.999 and  = 1 × 10 −6. Also, a linear warm-up
was used for both experiments for 10% of the total training steps, reaching a maximum value
of 5 × 10−6. Batch sizes of 68 and 16 are respectively set for pre-training and finetuning. A
weighted sampling was used during finetuning due to class imbalance. Lastly, we performed
pre-training with merely 50,000 examples for one epoch, which took around 15 minutes on an
NVIDIA V100.</p>
      <p>In order to exchange antonyms and synonyms, word types noun, adjective and verb were
extracted from the sentences and counted according to their occurrence. For each category, we
took the 500 most frequently occurring words and collected automatically up to five synonyms
and antonyms7 and randomly substituted them accordingly.</p>
      <p>Results and Evaluation In the following, we evaluate our approaches against the WMT’21
baseline and a XLM-R model without further pre-training. Special attention is kept on the
Matthews correlation coeficient (MCC), as it is the main metric used in the competition. Table 2
shows our results for the CED test set. Both approaches almost always boost the performance of
the model. We found that the best performing model incorporated both the word type adjective
and the imitation approach. In this combination, we would rank 2nd in the competition8 with
a correlation of 0.5117.9 The alignment approach is inferior to the imitation approach, which
makes it clear that pre-training on an identical task is noticeably more efective. We also did
not find performance diferences between the approaches of aligning with random words and
aligning with antonyms, and therefore decided to omit the results.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this paper, two informed approaches for pre-training with syntactic data for finding critical
errors in machine translations were proposed. We rank second in the English-German task 3,
with only 0.06% of data points used in comparison to the first placement. Therefore, we assume
5https://huggingface.co/xlm-roberta-large
6https://commoncrawl.org/
7From https://www.thesaurus.com/
8https://www.statmt.org/wmt21/quality-estimation-task_results.html#task3_results
9Because of time constraints, we did not take part in the actual challenge.</p>
      <p>WMT’21</p>
      <p>XLM-R
Alignment Simple</p>
      <p>Imitation
baseline
baseline
verb
adj.
noun
mix
verb
adj.
noun
mix</p>
      <p>0.786
0.5317
0.5810
that an informed approach using augmented data following the structure of the downstream task
i.e., training on prototypical examples as proposed by [17], can be more eficient than training
with a lot of data. Future work could include maximizing the prior information available to the
model to further reduce the amount of data needed. In addition, alignment in general seems to
be a good starting point for follow-up research in this domain. Other options for future work
include additional linguistically informed pre-training tasks, as described by [21] and [22].</p>
      <p>The insights gathered in this paper can likely be applied to other areas of text mining research
as well, e.g. natural language inference [23, 24, 25]. In the context of financial document analysis
[26, 27], our methods can be applied e.g. when comparing diferent versions of a document to
ifnd critical errors or contradictions. It can be worthwhile to design similar data augmentation
strategies that are tailored to the domain, e.g. replacing specific financial terms with their
opposite meaning.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>In parts, the authors of this work were supported by the Competence Center for Machine
Learning Rhine Ruhr (ML2R) which is funded by the Federal Ministry of Education and Research
of Germany (grant no. 01|18038 ).
(2019).
[17] L. von Rueden, S. Houben, K. Cvejoski, C. Bauckhage, N. Piatkowski, Informed pre-training
on prior knowledge, arXiv preprint arXiv:2205.11433 (2022).
[18] Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, A. Fan,
Multilingual translation with extensible multilingual pretraining and finetuning, arXiv preprint
arXiv:2008.00401 (2020).
[19] W. A. Gale, K. W. Church, et al., A program for aligning sentences in bilingual corpora,</p>
      <p>Computational linguistics 19 (1994) 75–102.
[20] I. Loshchilov, F. Hutter, Decoupled weight decay regularization, arXiv preprint
arXiv:1711.05101 (2017).
[21] J. Zhou, Z. Zhang, H. Zhao, LIMIT-BERT : Linguistic informed multi-task BERT, CoRR
abs/1910.14296 (2019).
[22] A. Wahab, R. Sifa, Dibert: Dependency injected bidirectional encoder representations
from transformers, in: Proc. of IEEE SSCI 2021, 2021.
[23] R. Sifa, M. Pielka, R. Ramamurthy, A. Ladi, L. Hillebrand, C. Bauckhage, Towards
contradiction detection in german: A translation-driven approach, in: Proc. of IEEE SSCI 2019,
2019.
[24] M. Pielka, R. Sifa, L. P. Hillebrand, D. Biesner, R. Ramamurthy, A. Ladi, C. Bauckhage,
Tackling contradiction detection in german using machine translation and end-to-end
recurrent neural networks, in: Proc. of ICPR 2020, 2021.
[25] L. Pucknat, M. Pielka, R. Sifa, Detecting contradictions in german text: A comparative
study, in: 2021 IEEE Symposium Series on Computational Intelligence (SSCI), IEEE, 2021,
pp. 01–07.
[26] R. Sifa, A. Ladi, M. Pielka, R. Ramamurthy, L. Hillebrand, B. Kirsch, D. Biesner, R.
Stenzel, T. Bell, M. Lübbering, U. Nütten, C. Bauckhage, U. Warning, B. Fürst, T.
Dilmaghani Khameneh, D. Thom, I. Huseynov, J. Kahlert, R. amd Schlums, H. Ismail, B. Kliem,
R. Loitz, Towards automated auditing with machine learning, in: Proceedings of the ACM
Symposium on Document Engineering 2019, 2019, pp. 1–4.
[27] L. Hillebrand, T. Deußer, C. Bauckhage, T. Dilmaghani, B. Kliem, R. Loitz, R. Sifa, Kpi-bert:
A joint named entity recognition and relation extraction model for financial reports, in:
Proc. ICPR (to be published), 2022.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Specia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Blain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fomicheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zerva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Martins</surname>
          </string-name>
          ,
          <article-title>Findings of the wmt 2021 shared task on quality estimation</article-title>
          ,
          <source>in: Proceedings of the Sixth Conference on Machine Translation</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>684</fpage>
          -
          <lpage>725</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Pre-training multilingual neural machine translation by leveraging alignment information</article-title>
          ,
          <source>in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2649</fpage>
          -
          <lpage>2663</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khabsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          , H. Ma,
          <article-title>Entailment as few-shot learner</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2104</volume>
          .
          <fpage>14690</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          , arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>02116</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Anantharaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <article-title>Larger-scale transformers for multilingual masked language modeling</article-title>
          ,
          <source>arXiv preprint arXiv:2105.00572</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Negri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Turchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chatterjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bertoldi</surname>
          </string-name>
          ,
          <article-title>Escape: a large-scale synthetic corpus for automatic post-editing</article-title>
          , arXiv preprint arXiv:
          <year>1803</year>
          .
          <volume>07274</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Two-phase cross-lingual language model fine-tuning for machine translation quality estimation</article-title>
          ,
          <source>in: Proceedings of the Fifth Conference on Machine Translation</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1024</fpage>
          -
          <lpage>1028</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Eo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Park</surname>
          </string-name>
          , H. Moon,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <article-title>Word-level quality estimation for korean-english neural machine translation</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>44964</fpage>
          -
          <lpage>44973</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rubino</surname>
          </string-name>
          , E. Sumita,
          <article-title>Intermediate self-supervised learning for machine translation quality estimation</article-title>
          ,
          <source>in: Proceedings of the 28th International Conference on Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4355</fpage>
          -
          <lpage>4360</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rubino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fujita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Marie</surname>
          </string-name>
          ,
          <article-title>Nict kyoto submission for the wmt'21 quality estimation task: Multimetric multilingual pretraining for critical error detection</article-title>
          ,
          <source>in: Proceedings of the Sixth Conference on Machine Translation</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>941</fpage>
          -
          <lpage>947</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Specia</surname>
          </string-name>
          ,
          <article-title>Icl's submission to the wmt21 critical error detection shared task</article-title>
          ,
          <source>in: Proceedings of the Sixth Conference on Machine Translation</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>928</fpage>
          -
          <lpage>934</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tao</surname>
          </string-name>
          , G. Jiaxin,
          <string-name>
            <given-names>W.</given-names>
            <surname>Minghan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , et al.,
          <article-title>Hw-tsc's participation at wmt 2021 quality estimation shared task</article-title>
          ,
          <source>in: Proceedings of the Sixth Conference on Machine Translation</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>890</fpage>
          -
          <lpage>896</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          <article-title>Qemind: Alibaba's submission to the wmt21 quality estimation shared task</article-title>
          ,
          <source>arXiv preprint arXiv:2112.14890</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Von Rueden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mayer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Beckh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Georgiev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Giesselbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Heese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kirsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pfrommer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramamurthy</surname>
          </string-name>
          , et al.,
          <article-title>Informed machine learning-a taxonomy and survey of integrating knowledge into learning systems</article-title>
          , arXiv preprint arXiv:
          <year>1903</year>
          .12394
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>