<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Forum for Information Retrieval Evaluation</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Translation for Indian Languages</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Surupendu Gangopadhyay</string-name>
          <email>surupendu.g@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ganesh Epili</string-name>
          <email>ganeshepili1998@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prasenjit Majumder</string-name>
          <email>prasenjit.majumder@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Baban Gain</string-name>
          <email>gainbaban@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ramakrishna Appicharla</string-name>
          <email>ramakrishnaappicharla@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asif Ekbal</string-name>
          <email>asif.ekbal@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arafat Ahsan</string-name>
          <email>arafat.ahsan@iiit.ac.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dipti Sharma</string-name>
          <email>dipti@iiit.ac.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dhirubhai Ambani Institute of Information and Communication Technology</institution>
          ,
          <addr-line>Gandhinagar</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gujarati</institution>
          ,
          <addr-line>Kannada, Odia, Punjabi, Urdu, Telugu, Kashmiri, and Sindhi. The track comprised of two tasks:</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Indian Institute of Technology Patna</institution>
          ,
          <addr-line>Patna</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>International Institute of Information Technology - Hyderabad</institution>
          ,
          <addr-line>Hyderabad</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>5</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>The objective of the MTIL track in FIRE 2023 was to encourage the development of Indian Language to Indian Language (IL-IL) Neural Machine Translation models. The languages covered included Hindi, (i) a General Translation Task and (ii) a Domain specific Translation Task with Governance and Healthcare being the chosen domains. For the listed languages, we proposed 12 diverse language directions for the general domain translation task and 8 each for healthcare and governance domains. Participants were encouraged to submit models for one or more language pairs. Consequently, we witnessed the creation of 34 distinct models spanning various language pairs and domains. Model assessments were conducted using five evaluation metrics: BLEU, CHRF, CHRF++, TER, and COMET. The submitted model outputs were ultimately ranked using the CHRF score. Neural Machine Translation, Domain specific Machine Translation, Machine Translation for Low resource Research on translation of low-resource languages opens up new challenges in the field of neural machine translation. Many Indian languages, especially, when the translation directions are IL-IL fall under a low resource scenario, hence the need for experimentation, and discovery of new techniques, that can help efectively translate between low resource language pairs. While some shared tasks previously have focused on Indic-English1 low resource language translation settings, the Indic-Indic translation directions need further exploration. This shared task, titled as Machine Translation for Indian Languages (MTIL) aims to fill this gap by proposing a number of Indic-Hindi and Hindi-Indic translation directions making test data available for a number of these pairs. Furthermore, the shared task also proposes domain-specific translation with htp:/ceur-ws.org CEUR Workshop Proceedings (CEUR-WS.org) ISN1613-073 1https://www2.statmt.org/wmt23/indic-mt-task.html</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>CEUR
Workshop
Proceedings
Governance and Healthcare being the domains in focus.</p>
      <sec id="sec-2-1">
        <title>1.1. Task 1: General Translation Task</title>
        <p>Task 1 is meant to be a general domain translation task where participants are required to build
a model to translate the language pairs shown in Table 1. They are provided pointers to existing
training data that may be mined for Hindi-Indic direction pairs. The task is unconstrained and
participants are free to adapt or leverage existing data and models to create the best models for
each translation direction.</p>
      </sec>
      <sec id="sec-2-2">
        <title>1.2. Task 2: Domain-specific Translation Task</title>
        <p>• Task 2a (Governance): In this subtask, the participants have to build a model to translate
sentences in the Governance domain.
• Task 2b (Healthcare): In this subtask, the participants have to build a model to translate
sentences in the Healthcare domain.</p>
        <p>The language pairs used in both the subtasks are shown in Table 2.</p>
        <p>We evaluate the submissions using the following evaluation metrics: BLEU, CHRF, CHRF++,
TER, and COMET. However, the final ranking is based on the CHRF scores. The evaluation
metrics are discussed in Section 3.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Dataset</title>
      <p>
        We encouraged participants to leverage existing publicly available parallel or monolingual
data for this shared task. Specifically, we encouraged the use of the Bharat Parallel Corpus
Collection BPCC[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] released by AI4Bharat. The Bharat Parallel Corpus Collection (BPCC) is
currently the largest English-Indic parallel corpus encompassing data for all 22 scheduled Indian
languages. This collection comprises of two sections, BPCC-Mined and BPCC-Human, and
contains approximately 230 million pairs of bitext. The BPCC-Mined section incorporates about
228 million pairs, with nearly 126 million pairs freshly added as part of this initiative. This
component plays a pivotal role in augmenting the available data for all 22 scheduled Indian
languages. On the other hand, BPCC-Human consists of 2.2 million gold standard English-Indic
pairs. Additionally, it includes 644K bitext pairs sourced from English Wikipedia sentences,
forming the BPCC-H-Wiki subset, and 139K sentences covering everyday use cases, forming
the BPCC-H-Daily subset. The statistics of the dataset is shown in Table 3.
      </p>
      <p>Test Data To ensure accurate evaluation of model performance, we make available a manually
translated test corpus for each language pair listed in either of the sub-tasks. The test set of Task
1 comprises of 2000 sentences, while the test set of Task 2 comprises 1000 sentences for each
language pair. Test sets are blind and only the Source is released for translation submissions.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Evaluation Metrics</title>
      <p>
        Evaluation is performed utilizing multiple metrics. Canonical string-based metrics like BLEU,
CHRF and TER are used and a pre-trained metric (COMET) is also utilized. The choice of metrics
was influenced by two factors: that the languages under evaluation exhibited considerable
morphological variation, thus the metric must not be biased against morphological complexity;
that while recently popular pre-trained metrics have shown greater correlations with human
judgements, they are yet to be proven to scale to lower resource languages, thus providing an
opportunity to test them for certain low resource languages that made up this shared task. We
use the SacreBLEU [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] library to evaluate the submissions. The evaluation metrics that we use
in this shared task are described below:
1. BLEU: The BLEU [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] score evaluates the quality of a translation based on the overlap of
n-grams between the hypothesis and reference and the length of the hypothesis w.r.t. the
reference. It uses unigrams, bigrams, trigrams, and four-grams to measure the overlap
of the n-grams. The BLEU score uses a brevity score to penalize the hypothesis that
is shorter in length than the reference. Since we are measuring the BLEU score w.r.t.
percentage, the value lies between 0-100, wherein a high BLEU indicates a better quality
of translation.
2. CHRF: The CHRF [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] score evaluates the quality of a translation based on the overlap of
character level n-gram between the hypothesis and reference. It calculates the F score
using the character level n-gram precision and recall. The CHRF score is better for
evaluating translations of morphologically rich languages. The formula of CHRF is shown
in Equation 1 wherein   is the percentage of n-grams present in the hypothesis,
which is present in the reference, and   is the percentage of n-grams present in
the reference, which is present in the hypothesis. The scores have a higher correlation
to human-level judgments as compared to BLEU. We used character level n-gram of 6
characters, and  was set to 1 to calculate the CHRF score.
      </p>
      <p>() = (1 + 
2</p>
      <p>
        .ℎ 
)  2.  +  
3. CHRF++: The CHRF++ [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is a modification over the CHRF wherein it considers the
overlap of character level and word level n-grams between the hypothesis and reference.
It takes an average of the F score of character level n-grams and the F score of word level
n-grams. We used the character level n-gram of 6 characters, word level n-gram of 1, and
 set to 1.
4. TER: The TER [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] score measures the number of edits (insert, delete, substitution, and
shift) required to match the hypothesis to the reference. The lower the TER score, the
better the performance of the model.
5. COMET: COMET [7] is an embedding-based metric to measure the similarity between
the hypothesis and reference. It uses an encoder to get the source (s), hypothesis (h),
and reference (r) sentence embeddings. The COMET score uses Equation 2 to calculate
a harmonic mean over the Euclidean distances between the reference and hypothesis
and source and hypothesis. It uses Equation 3 to calculate the similarity between the
hypothesis and reference. We use Unbabel/wmt22-comet-da as the COMET evaluation
model.
      </p>
      <p>(, ℎ,  ) =
2.(, ℎ).( , ℎ)
(, ℎ) + ( , ℎ)
(1)
(2)
 (, ℎ, ) =</p>
      <p>1
1 +  (, ℎ, )
(3)</p>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <p>We have received 34 submissions across various domains and language pairs. However, only
three of the teams submitted a paper detailing the methodology they employed. The results are
shown in Tables 4, 5, and 6.</p>
      <p>In Task 1 and 2 for the language pairs Hindi-Odia and Odia-Hindi the team IIIT-BH-MT
is the best performing team followed by BITSP. IIIT-BH-MT team used a custom Hindi-Odia
and Odia-Hindi dataset which consists of 36100 parallel sentences. The authors finetuned the
distilled version of NLLB 200 model which consists of 600M parameters for the translation
task. The BITSP team use the BPCC dataset wherein they use English as the pivot language
while translating from Odia to Hindi and vice versa. In Task 1 the authors use a combination
of IndicTrans2 and NLLB models to generate the translations from the source language and
MuRIL to generate sentence embeddings from the translations. The authors then select the
best translation based on the cosine similarity. For Task 2 the authors create a domain specific
dataset from BPCC by using BART-MNLI to assign the class labels. They finetune the NLLB
model to perform task specific translation.</p>
      <p>In case of Punjabi-Hindi and Hindi-Punjabi the CDACN-Punjabi team is the best performing
team except in Task 2a where SLPBV team is the best performing team. The CDACN-Punjabi
team use a custom dataset of Punjabi-Hindi and vice versa language pairs to train their model.
The authors finetune the NLLB-200 model for the translation task.</p>
      <p>In Hindi-Sindhi the scores for SLPBV team’s submission was unusually low. However, on
closer examination, we found that the submitted model outputs are using the Devanagri script</p>
      <sec id="sec-5-1">
        <title>Hindi-Odia</title>
      </sec>
      <sec id="sec-5-2">
        <title>Odia-Hindi</title>
      </sec>
      <sec id="sec-5-3">
        <title>Gujarati-Hindi</title>
      </sec>
      <sec id="sec-5-4">
        <title>Punjabi-Hindi</title>
      </sec>
      <sec id="sec-5-5">
        <title>Hindi-Punjabi</title>
      </sec>
      <sec id="sec-5-6">
        <title>Hindi-Sindhi</title>
      </sec>
      <sec id="sec-5-7">
        <title>BITSP</title>
      </sec>
      <sec id="sec-5-8">
        <title>SLPBV</title>
      </sec>
      <sec id="sec-5-9">
        <title>IIIT-BH-MT</title>
      </sec>
      <sec id="sec-5-10">
        <title>BITSP</title>
      </sec>
      <sec id="sec-5-11">
        <title>SLPBV</title>
      </sec>
      <sec id="sec-5-12">
        <title>IIIT-BH-MT</title>
      </sec>
      <sec id="sec-5-13">
        <title>SLPBV</title>
      </sec>
      <sec id="sec-5-14">
        <title>SLPBV</title>
      </sec>
      <sec id="sec-5-15">
        <title>CDACN-Punjabi</title>
      </sec>
      <sec id="sec-5-16">
        <title>SLPBV</title>
      </sec>
      <sec id="sec-5-17">
        <title>CDACN-Punjabi</title>
      </sec>
      <sec id="sec-5-18">
        <title>SLPBV</title>
        <p>for Sindhi, whereas our ground truth has Sindhi in the Perso-Arabic script. In Gujarati-Hindi
we had only one submission by the SLPBV team.</p>
        <p>In Task 1 from Table 4 we observe that in Hindi-Odia language pair the COMET score of
IIIT-BH-MT and BITSP teams are same. Similarly in Task 2a as well the COMET scores of the
both the teams are same and in Task 2b the COMET score of BITSP is higher than IIIT-BH-MT
by 0.017. However if we observe the other metrics then it shows a diference in performance of
both the teams. However in other language pairs we do not observe this anomalous behaviour.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Concluding Discussions</title>
      <p>Our shared task focuses on Indic-Indic language translation instead of Indic-English language
translation. From the submission we observe that NLLB model is widely used among the
participants for the translation task. Among the three teams we found one team which used
BPCC as the training dataset and English as the pivot language. While the other two teams
used custom dataset for the shared tasks. We observed some anomalous behaviour between
the chrF and COMET scores in case of Hindi-Odia language pair wherein the first and second
team had the same COMET score, and one case in the Governance domain sub-task, where the
CHRF and COMET scores were discordant. We need to investigate this further by using human
annotators to evaluate the translations. However in other language pairs we did not observe
this behaviour.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The track organizers thank all the participants for their interest in this track. We also thank
the FIRE 2023 organizers for their support in organizing the track. We also thank BHASHINI
(https://bhashini.gov.in/en) for enabling the HIMANGY consortium to create the test dataset
which helped us to conduct this track. We thank the Principal Investigator, Co-Principal
Investigators and Host Institute (IIIT Hyderabad) for providing us with this opportunity of using
the dataset in the track. We also thank Ministry of Electronics and Information Technology
(MeitY) and Ministry of Human Resource Development, Government of India for providing this
opportunity to develop the dataset and other resources.
[7] R. Rei, C. Stewart, A. C. Farinha, A. Lavie, COMET: A neural framework for MT evaluation,
in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language
Processing (EMNLP), Association for Computational Linguistics, Online, 2020, pp. 2685–2702. URL:
https://aclanthology.org/2020.emnlp-main.213. doi:10.18653/v1/2020.emnlp- main.213.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>AI4Bharat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gala</surname>
            ,
            <given-names>P. A.</given-names>
          </string-name>
          <string-name>
            <surname>Chitale</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>AK</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Doddapaneni</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Gumma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Nawale</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sujatha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Puduppully</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Khapra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Dabre</surname>
            ,
            <given-names>A. Kunchukuttan,</given-names>
          </string-name>
          <article-title>Indictrans2: Towards high-quality and accessible machine translation models for all 22 scheduled indian languages</article-title>
          ,
          <source>arXiv preprint arXiv: 2305.16307</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Post</surname>
          </string-name>
          ,
          <article-title>A call for clarity in reporting BLEU scores</article-title>
          ,
          <source>in: Proceedings of the Third Conference on Machine Translation: Research Papers</source>
          , Association for Computational Linguistics, Belgium, Brussels,
          <year>2018</year>
          , pp.
          <fpage>186</fpage>
          -
          <lpage>191</lpage>
          . URL: https://www.aclweb.org/anthology/W18-6319.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W.-J. Zhu,
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          ,
          <source>in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Popović</surname>
          </string-name>
          ,
          <article-title>chrf: character n-gram f-score for automatic mt evaluation</article-title>
          ,
          <source>in: Proceedings of the tenth workshop on statistical machine translation</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>392</fpage>
          -
          <lpage>395</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Popović</surname>
          </string-name>
          , chrf++
          <article-title>: words helping character n-grams</article-title>
          ,
          <source>in: Proceedings of the second conference on machine translation</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>612</fpage>
          -
          <lpage>618</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Snover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dorr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Micciulla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Makhoul</surname>
          </string-name>
          ,
          <article-title>A study of translation edit rate with targeted human annotation</article-title>
          ,
          <source>in: Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>