<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of Verb Phrase Translation in Machine Translation: English to Tamil and Hindi to Tamil</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vijay Sundar Ram R</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sobha Lalitha Devi</string-name>
          <email>sobha@au-kbc.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AU-KBC Research Centre, MIT Campus of Anna University</institution>
          ,
          <addr-line>Chennai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an overview of verb phrase translation in machine translation from English to Tamil and Hindi to Tamil track, where English, Hindi and Tamil belong to three different language families, namely, Indo-European, Indo-Aryan and Dravidian family respectively. Verb phrases carry syntactic information such as tense, aspect, modal, and PNG (person, number and gender) other than the main verb. The characteristics of verb phrase vary between languages, which make the task challenging. In Tamil, non-finite verbs introduce clauses, whereas in Hindi clauses are introduced by relative-correlatives and the clause has a finite verb. We have five registrations and three out of the five registered teams submitted their runs. There were three submissions for English to Tamil Verb phrase translation and two submissions for Hindi to Tamil verb phrase translation. The runs are evaluated based on the correctness of tense, aspect, modal and PNG of the verb phrase.</p>
      </abstract>
      <kwd-group>
        <kwd>Verb phrase translation</kwd>
        <kwd>English-Tamil</kwd>
        <kwd>Hindi-Tamil</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Machine translation (MT) is one of the most active research areas in Natural
Language Processing across the globe. MT in Indian languages has picked-up in the last
two decades. In developing a machine translation system, translation of the verb
phrase from source language to target language is a challenging task. The verb
phrases include finite verb, non-finite verb, auxiliary verb, main verb, verbal particles
and negation verb constructions. Verb phrases also carry information namely, tense,
aspect, modal, and PNG (person, number and gender) other than the main verb. The
characteristics of the verbs vary between languages. In languages such as Tamil,
Telugu, Kannada, Hindi, the subject and the finite verb of the sentence agree in PNG. In
languages such as English and Malayalam, there is no agreement between the subject
and finite verb. Languages vary in structure also such as SVO, SOV. These
characteristics make the translation of Verb phrases from one language to another a difficult
task</p>
      <p>
        The objective of this shared task is to boost the research in Machine translation in
Indian languages. We have narrowed down the scope of the track to translation of
Verb Phrases from English to Tamil and Hindi to Tamil. These three languages are
from different language families namely, Indo-European, Indo-Aryan and Dravidian.
Sentence structures and characteristics of the verbs vary largely across these
languages and make the task an interesting and challenging task. Sobha et al [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has
presented a work on transfer of verb phrases from Tamil to Hindi.
      </p>
      <p>The researchers were welcomed to come-up with various methodologies such as
rule-based, Machine Learning and Hybrid techniques. Participants were allowed to
use any pre-processing tools, which are in open source or developed in-house.</p>
      <p>The flow of the paper is as follows. In the following section, we have described
about the data set provided for both the training and testing. We have given the details
of the participants in the third section. In the fourth section, we have explained the
methodologies followed by each team. The metrics for evaluation is presented in fifth
section. The paper is concluded with a brief summary of the shared task.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Data set</title>
      <p>We provided the training and testing dataset. The training data had a set of three files
for each translation pair, which contains source language sentence, target language
translated sentence and verb phase mapping index. Both the source and target
language translated sentences has sentence indexing. Verb phrase (VP) mapping index
file has the sentence index and the information of the position of the VP in the source
language sentence and the position of the corresponding verb phrase in the target
language translated sentence. The structure of the VP map index fie is as follows.
&lt;vpInfo sentId='' srcLang='en/hi' tgtLang='ta' vpId='verbphrase-id' vp_src_info=''
vp_tgt_info=''&gt;
where:
sentId: is the sentence Id.
srcLang: Source language code. It can be 'en/hi'
tgtLang: Target language code. It is 'ta' as Tamil is the target language in both the
pairs.</p>
      <p>vpId: Each verb phrase is marked with an unique id.
vp_src_info: 'verb phrase start position and its length'
vp_tgt_info: 'verb phrase start position and its length'</p>
      <p>We present below a sample source sentence, target translated sentence and its verb
phrase map index from both English to Tamil and Hindi to Tamil in example 1 and 2
respectively.</p>
      <p>Ex 1:
English to Tamil translated Sentence:</p>
      <sec id="sec-2-1">
        <title>Source Sentence:</title>
        <p>&lt;Sent Id=24 lang='en'&gt; Vandiyathevan did not get up .&lt;/Sent&gt;</p>
      </sec>
      <sec id="sec-2-2">
        <title>Translated Sentence:</title>
        <p>&lt;Sent Id=24 lang='ta'&gt; வந்தியதத்தவன்
எழுந்திருகக்வில்லை
. &lt;/Sent&gt;</p>
      </sec>
      <sec id="sec-2-3">
        <title>VP Map Index:</title>
        <p>&lt;vpInfo sentId='24' srcLang='en' tgtLang='ta' vpId='36' vp_src_info='15,14'
vp_tgt_info='16,18'&gt;
The source English sentence has a verb phrase ‘did not get up’. And the translated
Tamil sentence has the equivalent verb phrase ‘எழுதந்ிருகக்வில்லை ’. vp_src_info
in VP Map Index has the starting and length information of the VP in the source
sentence, 15 and 14 respectively. Similarly vp_tgt_info has the starting and length
information of the VP in the source sentence, 16 and 18 respectively. An unique Id is
given to the verb phrase.</p>
        <p>Ex 2:
Hindi to Tamil translated Sentence:</p>
      </sec>
      <sec id="sec-2-4">
        <title>Source Sentence:</title>
        <p>पर्यटक स्थल है .&lt;/Sent&gt;
Translated Sentence
&lt;Sent Id=8 lang='ta'&gt;ஏததன்ஸ
லவக்மு</p>
        <p>அழகான ற்ைசுாத்
&lt;Sent Id=8 lang='hi'&gt;एथेंस महाद्वीप, सैलानियों को रोमांचित कर देने वाला एक आकर्षक
கூப டண்ம்
லைைானியகர்லை</p>
        <p>தெய்ைிர்கலிக
தைொக உள்ைது
.&lt;/Sent&gt;</p>
      </sec>
      <sec id="sec-2-5">
        <title>VP Map Index:</title>
        <p>&lt;vpInfo sentId='8' srcLang='hi' tgtLang='ta' vpId='12' vp_src_info='38,21'
vp_tgt_info='30,22'&gt;</p>
        <p>&lt;vpInfo sentId='8' srcLang='hi' tgtLang='ta' vpId='13' vp_src_info='82,2'
vp_tgt_info='76,6'&gt;</p>
        <p>The source sentence, Hindi sentence, has two verbs ‘कर देने वाला’ and ‘है’, whose
equivalent in Tamil are ‘லவுக்ம ’ and ‘உைள்து ’ respectively. In VP map Index,
the position of the verbs in both source and target sentences are given along with an
unique id for each verb phrase.</p>
        <p>In certain verb phrase such as verb phrase in interrogative sentence, the verbs occur
separately with the words in-between. Consider the following example Ex 3. For
these types of sentence, VP mapping index will have positional information of the
verbs separated by “;”. An example is given below.
Ex 3.</p>
        <p>Source Sentence:
&lt;Sent Id=40 lang='en'&gt;ENG:"How did that happen , Swami ?"&lt;/Sent&gt;</p>
      </sec>
      <sec id="sec-2-6">
        <title>Translated Sentence:</title>
        <p>&lt;Sent Id=40 lang='ta'&gt;TAM:"அது
எப்படி
நடந்தது
ுசவாெிகதை !"&lt;/Sent&gt;
Consider the above example, the source sentence, an integrative sentence, has the
verb phrase ‘did happen’, occurring separately with ‘that’ in the middle. vp_src_info
in VP Map Index has ‘9,3;18,6' its value. It has the starting position inform of both
‘did’ and ‘happen’ and their lengths.</p>
        <p>The verb phrases include finite verb, non-finite verb, auxiliary verb main verb,
verbal particles and negation verb constructions.</p>
        <p>The statistics of the VPs indexed in the training and testing data is given in the
table 1 below.
S.No Institution TeamId
1 Department of Computer Science, Banasthali Vidyapith, Joshi-Banasthali</p>
        <p>Banasthali - 304022, Rajasthan
2 National Institute of Technology Mizoram, Chalatlang, Aizawl Pakray-NITM
796012, Mizoram, India
3 IIIT-Delhi
4 SSN College of Engineering, Old Mahabalipuram Road,</p>
        <p>Kalavakkam, Chennai
Choudhary-IIITD
Thenmozhi-SSN</p>
        <p>Centre for Applied Linguistics and Translation Studies, Uni- Parameswari-HCU
versity of Hyderabad, Gachibowli Hyderabad
Out of five registered participants, three participants submitted their runs. The
submission details are given in the table 3.
Choudhary-IIITD has used a neural machine translation technique using
wordembedding along with Byte-Pair-Encoding (BPE) to develop an efficient translation
system, called MIDAS translator. We used OpenNMT-py a neural machine
translation toolkit for training our model. It is based on py-torch. They have used this
translation system for both of the translation pairs i.e. English-Tamil and Tamil-Hindi.</p>
        <p>Data pre-processing task included tokenization of sentences using their own
tokenizer. The tokenization available in OpenNMT-py is not used.</p>
        <p>After different trials they reported the best results using a Byte-pair-Encoded
vocabulary , 2 Layer Bi-directional encoder-decoder, Adam optimization with a learning
rate of 0.001, dropout (regularization) of 0.3, Bahdanau attention, and
wordembedding with the dimension of 500.</p>
        <p>Thenmozhi-SSN has adopted Neural Machine Translation model for this task. They
have used a deep learning approach based on Seq2Seq model for English-Tamil and
Hindi-Tamil VP translations. The network consists of an embedding layer,
encodingdecoding layer with 8-layer LSTM and a projection layer to translate the verb phrases
from English / Hindi to Tamil.</p>
        <p>They have used TensorFlow for implementing the deep neural network.</p>
        <p>They have implemented the Seq2Seq model using the bi-directional LSTM with 8
layers, dropout (regularization) of 0.2, Bahdanau attention, and the number of training
steps is 50000.</p>
        <p>Parashwari-HCU has used a rule-based approach for English to Tamil verb
translation. They have performed the translation as three step process.</p>
        <p>First they used a Stanford Dependency parser to identify the subject of the VP. The
GNP information of the subject is noted. Second, using nltk lemmatizer the verb root
is identified. Using a set of transfer rules the TAM of the target verb is generated and
the equivalent Tamil for the verb root is replaced. In the third step, they have
processed it with an in-house developed word-generator to generate Tamil verb phrase.
S.No
1
2
3</p>
        <p>A brief summary of the methodology followed by the teams is presented in the
table 4 below.
We use a scoring methodology based on the correctness of the TAM (tense, aspect
modal), Person, Number and Gender of the translated VP. The criteria for scoring the
results are given in table 5.</p>
        <sec id="sec-2-6-1">
          <title>English to Tamil Verb Phrase Translation</title>
          <p>We had submissions from three teams for English to Tamil Verb phrase translation.
We present teams against the Criteria Scores for English to Tamil in table 6. We have
given details of number of verb phrases translation scored under each of the criteria
scores for each of the teams. The precision and recall obtained by each team is
presented in table 7.</p>
        </sec>
        <sec id="sec-2-6-2">
          <title>Hindi to Tamil Verb Phrase Translation</title>
          <p>We had submissions from two teams for Hindi to Tamil Verb phrase translation. We
present teams against the Criteria Scores for Hindi to Tamil in table 8. Similar to table
6, we have given details of number of verb phrases translation scored under each of
the criteria scores for each of the teams in table 8. The precision and recall obtained
by each team is presented in table 9.</p>
          <p>It is interesting to compare the number of verb phrase translations under each
scoring criteria in table 6 and 8. We find more number of verb phrase translations
under score 4 (Completely Correct) and score 2 (Correct root and TAM partially correct)
and very less in criteria score 3 (TAM and PNG Correct). In both the tracks, all the
submission shows that the generation of correct TAM and PNG is difficult.
Identification of the root verb and translating it to the target language is found to be
good based on the statistics of criteria score 2. In both Neural network based
approach and transfer rule based approach the TAM and PNG generation to target
language needs to be improved.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>We have presented the overview of Verb phrase translation in English and Indian
languages (VPT-IL) shared task. The shared task focused on translation of verb
phrases from English to Tamil and Hindi to Tamil. As the three languages belong to
three diffident language families and vary in their verb phrase formation, the task is a
challenge. There are five registered teams, out of which three teams submitted their
runs. Out of the three teams, two teams followed neural machine translation approach
and the third linguistic rule based approach. Among the two neural machine
translation, one used neural machine translation using word-embedding and the other used
neural machine translation using Seq2Seq modeling. The third team with the
rulebased approach used dependency parser to parse the source sentence and used a set of
transfer rules to translate the verb phrases. The evaluation of the VP translation in
both the tracks clearly presents the difficulties in generating correct TAM and PNG.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Sobha</surname>
          </string-name>
          .L.,
          <string-name>
            <surname>Pralayankar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menaka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakiyavathi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ram</surname>
            ,
            <given-names>R.V.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavitha</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Verb transfer in a Tamil to Hindi machine translation system</article-title>
          .
          <source>In: Asian Language Processing (IALP)</source>
          , 2010 International Conference on. pp.
          <fpage>261</fpage>
          -
          <lpage>264</lpage>
          . IEEE (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>