<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dictionary-based Amharic-French Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Atelach Alemu Argaw and Lars Asker Department of Computer and Systems Sciences, Stockholm University/KTH</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Question answering</institution>
          ,
          <addr-line>Amharic, Cross-Language Information Retrieval</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rickard C ̈oster, Jussi Karlgren and Magnus Sahlgren Swedish Institute of Computer Science</institution>
          ,
          <addr-line>SICS</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present four approaches to the Amharic - French bilingual track at CLEF 2005. All experiments use a dictionary based approach to translate the Amharic queries into French Bags-of-words, but while one approach uses word sense discrimination on the translated side of the queries, the other one includes all senses of a translated word in the query for searching. We used two search engines: The SICS experimental engine and Lucene, hence four runs with the two approaches. Non-content bearing words were removed both before and after the dictionary lookup. TF/IDF values supplemented by a heuristic function was used to remove the stop words from the Amharic queries and two French stopwords lists were used to remove them from the French translations. In our experiments, we found that the SICS search engine performs better than Lucene and that using the word sense discriminated keywords produce a slightly better result than the full set of non discriminated keywords.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Amharic topic set
|
| 1a. Transliteration
|
2. Transliterated Amharic topic set
|
| 2a. Trigram and Bigram dictionary lookup -----|
| |
3. Remaining (non matched) Amharic topic set |
| |
| 3a. Stemming |
| |
4. Stemmed Amharic topic set |
| |
| 4a. IDF-based stop word removal |
| |
5. Reduced Amharic topic set |
| |
| 5a. Dictionary lookup |
| |
6. Topic set (in French) including all possible translations
| |
| 6a. Word sense discrimination |
| |
7. Reduced set of French terms |
| |
| 7a. Retrieval (Indexing, keyword search, ranking)
|
8. Retrieved Documents
the languages, Amharic has gained ground through out the country. Amharic is used in business,
government, and education. Newspapers are printed in Amharic as are numerous books on all
subjects [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>In this paper we describe our experiments at the CLEF 2005 Amharic - French bilingual track.</title>
    </sec>
    <sec id="sec-3">
      <title>It consists of four fully automatic approaches that differ in terms of how word sense discrimination</title>
      <p>
        is done and in terms of what search engine is used. We have experimented with two different search
engines - Lucene [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], an open source search toolbox, and Searcher, an experimental search engine
developed at SICS1. Two runs were submitted per search engine, one using all content bearing,
expanded query terms without any word sense discrimination, and the other using a smaller
’disambiguated’ set of content bearing query terms.
      </p>
    </sec>
    <sec id="sec-4">
      <title>For the dictionary lookup we used one Amharic - French machine readable dictionary (MRD)</title>
      <p>
        containing 12.000 Amharic entries with corresponding 36,000 French entries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We also used an
      </p>
    </sec>
    <sec id="sec-5">
      <title>Amharic - English machine readable dictionary with approximately 15.000 Amharic entries [2] as a complement for the cases when the Amharic terms where not found in the Amharic - French MRD.</title>
      <p>1The Swedish Institute of Computer Science</p>
      <sec id="sec-5-1">
        <title>Method</title>
        <p>The English topic set was initially translated into Amharic by human translators. Amharic uses
its own and unique alphabet (Fidel) and there exist a number of fonts for this, but to date there is
no standard for the language. The Amharic topic set was originally represented using an Ethiopic
font but for ease of use and compatibility reasons we transliterated it into an ASCII representation
using SERA2. The transliterated Amharic topic set was then used as the input to the following
steps.
Before any stemming was done on the Amharic topic set, the sentences from each topic was used
to generate all possible trigrams and bigrams. These trigrams and bigrams where then matched
against the entries in the two dictionaries. First the full (unstemmed) trigrams where matched
against the Amharic - French and then the Amharic - English dictionaries. Secondly, prefixes were
removed from the first word of each trigram and suffixes were removed from the last word of the
same trigram and then what remained was matched against the two dictionaries. In this way, one
trigram was matched and translated for the full Amharic topic set, using the Amharic - French
dictionary.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Next, all bigrams where matched against the Amharic - French and the Amharic - English</title>
      <p>dictionaries. Including the prefix suffix removal, this resulted in the match and translation of 15
unique bigrams. Six were found only in the Amharic - French dictionary, another six were found
in both dictionaries, and three were found only in the Amharic - English dictionary. For the six
bigrams that were found in both dictionaries, the French translation was used.
2.3</p>
      <sec id="sec-6-1">
        <title>Stop word removal</title>
        <p>In these experiments, stop words were removed both before and after the dictionary lookup. First
the number of Amharic words in the queries was reduced by using a stopword list that had been
generated from a 2 million word Amharic news corpus using IDF measures. After the dictionary
lookup further stop words removal was conducted on the French side separately for the two sets
of experiments using the SICS engine and Lucene. For the SICS engine, this was done by using a
separete French stop words list. For the Lucene experiments, we used the French Analyszer from
the Apache Lucene Sandbox which supplements the query analyzer with its own list of French
stop words and removes them before searching for a specific keywords list.
2.4</p>
      </sec>
      <sec id="sec-6-2">
        <title>Amharic stemming and dictionary lookup</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>The remaining Amharic words where then stemmed and matched against the entries in the two</title>
      <p>dictionaries. The Amharic - French dictionary was always preferred over the Amharic - English
one. Only in cases when a term had not been matched in the French dictionary was it matched
against the English one. In a similar way, trigrams were matched before bigrams, bigrams before
unigrams, unstemmed terms before stemmed terms, unchanged root forms were matched before
modified root forms, longer matches in the dictionary were preferred before shorter etc.</p>
    </sec>
    <sec id="sec-8">
      <title>The terms for which matches were found only in the Amharic-English MRD where first translated into English and then further translated from English into French using an online electronic dictionary from WordReference (www.wordreference.com).</title>
      <p>2SERA stands for System for Ethiopic Representation in ASCII,
http://www.abyssiniacybergateway.net/fidel/serafaq.html</p>
      <p>Words and phrases that where not found in any of the dictionaries (mostly proper names or
inherited words) were not translated and instead handled by an edit-distance based similarity
matching algorithm. Frequency counts in a 2.4 million words Amharic news corpus was used to
determine whether an out of dictionary word would qualify as a candidate for a proper name
or not. The assumption here is that if a word that is not included in any dictionary appears
quite often in an Amharic text collection, then it is likely that the word is a term in the language
although not found in the dictionary. On the other hand, if a term rarely occurs in the news corpus
(in our case we used a threshold of nine times or less, but this of course depends on the size of the
corpus), the word has a higher probability of being a proper name or an inherited word. Although
this is a crude assumption and inherited words may occur frequently in a language, those words
tend to be mostly domain specific. In a news corpus such as the one we used, the occurrence of
almost all inherited words which could not be matched in the MRDs was very limited.
2.5</p>
      <sec id="sec-8-1">
        <title>Word sense discrimination</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>For the word sense discrimination we made use of two MRDs to get all the different senses of a</title>
      <p>term (word or phrase) - as given by the MRD, and a statistical collocation measure of mutual
information using the target language corpus to assign each term to the appropriate sense.</p>
      <p>
        In our experiments we used the bag of words approach where context is considered as words
in some window surrounding the target word, regarded as a group without consideration for their
relationships to the target in terms of distance, grammatical relations, etc. There is a big difference
between the two languages under consideration (Amharic and French) in terms of word ordering,
morphology, syntax etc, and hence limiting the context to a few number of words surrounding the
target word was intuitively undesirable. A sentence could have been taken as a context window,
but following the “one sense per discourse” constraint [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] in discriminating amongst word senses,
a context window of a whole article was implemented. This constraint states that the sense of a
word is highly consistent within any given document, in our case a French news article. The words
to be sense discriminated are the query keywords, which are mostly composed of nouns rather than
verbs, or adjectives. Noun sense discrimination is reported to be aided by word collocations that
have a context window of hundreds of words, while verb and adjective senses tend to fall off rapidly
with distance from the target word. After going through the list of translated content bearing
keywords, we noticed that the majority of these words are nouns, and hence the selection of the
document context window.
      </p>
    </sec>
    <sec id="sec-10">
      <title>In these experiments the Mutual Information between word pairs in the target language text</title>
      <p>collection is used to discriminate word senses. (Pointwise) mutual information compares the
probability of observing two events x and y together (the joint probability) with the probabilities
of observing x and y independently (chance). If two (words), x and y, have probabilities P(x) and</p>
    </sec>
    <sec id="sec-11">
      <title>P(y), then their mutual information, I(x,y), is defined to be:</title>
      <p>I(x, y) = log2 P (x).P (y) = log2 PP((xx/)y))</p>
      <p>P (x,y)</p>
    </sec>
    <sec id="sec-12">
      <title>If there is a genuine association between x and y, P(x,y) will be much larger than chance P(x)*</title>
    </sec>
    <sec id="sec-13">
      <title>P(y), thus I(x,y) will be greater than 0. If there is no interesting relationship between x and y,</title>
    </sec>
    <sec id="sec-14">
      <title>P(x,y) will be approximately equal to P(x)* P(y), and thus, I(x,y) will be close to 0. And if x and y are in complementary distribution, P(x,y) will be much less than P(x)* P(y), and I(x,y) will be less than O.</title>
    </sec>
    <sec id="sec-15">
      <title>Although very widely used by researchers for different applications, MI has also been criticized by many as to its ability to capture the similarity between two events especially when there is data scarcity [6]. Since we had access to a large amount of text collection in the target language, and because of its wide implementation, we chose to use MI.</title>
      <p>The translated French query terms were put in a bag of words, and the mutual information for
each of the possible word pairs was calculated. When we put the expanded words we treat both
synonyms and translations with a distinct sense as given in the MRD equally. Another way of
handling this situation is to group synonyms before the discrimination. We chose the first approach
with two assumptions: one is that even though words may be synonymous, it doesn’t necessarily
mean that they are all equally used in a certain context, and the other being even though a word
may have distinct senses defined in the MRD, those distinctions may not necessarily be applicable
in the context the term is currently used. This approach is believed to ensure that words with
inappropriate senses and synonyms with less contextual usage will be removed while at the same
time the query is being expanded with appropriate terms.</p>
    </sec>
    <sec id="sec-16">
      <title>We used a subset of the CLEF French document collection consisting of 14,000 news articles</title>
      <p>with 4.5 million words in calculating the MI values. Both the French keywords and the document
collection were lemmatized (by SICS using tools from connexor, http://www.connexor.com/) in
order to cater for the different forms of each word under consideration.</p>
      <p>Following the idea that ambiguous words can be used in a variety of contexts but collectively
they indicate a single context and particular meanings, we relied on the number of association as
given by MI values that a certain word has in order to determine whether the word should be
removed from the query or not. Given the bag of words for each query, we calculated the mutual
information for each unique pair. The next step was to see for each unique word how many positive
associations it has with the rest of the words in the bag. We experimented with different levels
of combining precision and recall values depending on which one of these two measures we want
to give more importance to. To contrast the approach of using the maximum recall of words (no
discrimination) we decided that precision should be given much more priority over recall (beta
value of 0.15), and we set an empirical threshold value of 0.4. i.e. a word is kept in the query if it
shows positive associations with 40% of the words in the list, otherwise it is removed. Here, note
that the mutual information values are converted to a binary 0, and 1. 0 being assigned to words
that have less than or equal to 0 MI values (independent term pairs), and 1 to those with positive</p>
    </sec>
    <sec id="sec-17">
      <title>MI values (dependent term pairs). We are simply taking all positive MI values as indicators of</title>
      <p>association without any consideration as to how strong the association is. This is done to input as
much association between all the words in the query as possible rather than putting the focus on
individual pairwise association values. Results of the experiments are given in the next section.</p>
    </sec>
    <sec id="sec-18">
      <title>The amount of words in each query (both in the English and corresponding translated Amharic)</title>
      <p>differed substantially from one query to another. After the dictionary lookup and stop word
removal, there were queries with French words that ranged from 2 to 71. This is due to a large
difference in the number of words and in the number of stop words in each query as well as the
number of senses and synonyms that are given in the dictionary for each word.</p>
      <p>When there were less than or equal to 8 words in the expanded query, there was no word sense
discrimination done for those queries. This is an arbitrary number, and the idea here is that if
the number of terms is as small as that, then it is much better to keep all words. We believe that
erroneously removing appropriate words in short queries has a lot more disadvantage than keeping
one with an inappropriate sense.
2.6
2.6.1</p>
      <sec id="sec-18-1">
        <title>Retrieval</title>
        <p>Retrieval using Lucene</p>
      </sec>
    </sec>
    <sec id="sec-19">
      <title>Apache Lucene is an open source high-performance, full-featured text search engine library written in Java [9]. It is a technology deemed suitable for applications that require full-text search, especially in a cross-platform.</title>
      <p>2.6.2</p>
      <p>Retrieval using Searcher</p>
    </sec>
    <sec id="sec-20">
      <title>The underlying retrieval engine is an experimental system developed at SICS. For retrieval, we</title>
      <p>
        use Pivoted Unique Normalization [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], where the score for a document d given a query with m
query terms is defined as
      </p>
      <p>1+log (tfi,d)</p>
      <p>
        Pim=1 1+log (average tfd)
(1 − slope) × pivot + slope × # of unique terms
where tfi,d is the term frequency of query term i in document d, and average tfd is the average
term frequency in document d. The slope was set to 0.3, and the pivot to the average number of
unique terms in a document, as suggested in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
3
      </p>
      <sec id="sec-20-1">
        <title>Results</title>
      </sec>
    </sec>
    <sec id="sec-21">
      <title>We have submitted four parallel Amharic-French runs at the CLEF 2005 ad-hoc bilingual track.</title>
      <p>
        We have used two search engines - Lucene [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], an open source search toolbox, and an experimental
search engine developed at SICS (Searcher). The aim of using these two search engines is to
compare the performance of the systems as well as to investigate the impact of performing word
sense discrimination. Two runs were submitted that use the same search engine, with one of them
searching for all content bearing, expanded query terms without any word sense discrimination
while the other one searches for the ’disambiguated’ set of content bearing query terms. The four
runs are:
      </p>
    </sec>
    <sec id="sec-22">
      <title>1. Lucene with word sense discrimination (am-fr-da-l)</title>
    </sec>
    <sec id="sec-23">
      <title>2. Lucene without word sense discrimination (am-fr-nonda-l)</title>
    </sec>
    <sec id="sec-24">
      <title>3. Searcher with word sense discrimination (am-fr-da-s)</title>
    </sec>
    <sec id="sec-25">
      <title>4. Searcher without word sense discrimination (am-fr-nonda-s)</title>
    </sec>
    <sec id="sec-26">
      <title>We have demonstrated the feasability of doing cross language information retrieval between</title>
      <p>Amharic and French. Although there is still much room for improvement of the results, we
are pleased to have been able to use a fully automatic approach. The work on this project and
the performed experiments have highlighted some of the more crucial steps on the road to better
information access and retrieval between the two languages. The lack of electronic resources such
as morphological analysers and large machine readable dictionaries have forced us to spend
considerable time on getting access to, or developing these resources ourselves. We also believe that,
in the absense of larger electronic dictionaries, one of the more important obstacles on this road
is how to handle out-of-dictionary words. The approach that we tested in our experiments, to use
fuzzy string matching in the retrieval step, seems to be only partially successful, mainly due to the
large differences between the two languages. We have also been able to compare the performance
between different search engines and to test different approaches to word sense discrimination.</p>
      <sec id="sec-26-1">
        <title>Acknowledgements</title>
      </sec>
    </sec>
    <sec id="sec-27">
      <title>The copyright to the two volumes of the French-Amharic and Amharic-French dictionary (”Dictio</title>
      <p>nnaire Francais-Amharique” and ”Dictionnaire Amharique-Francais”) by Dr Berhanou Abebe and
loi Fiquet is owned by the French Ministry of Foreign Affairs. We would like to thank the authors
and the French embassy in Addis Ababafor allowing us to use the dictionary in this research.</p>
    </sec>
    <sec id="sec-28">
      <title>The content of the “English - Amharic Dictionary” is the intellectual property of Dr Amsalu</title>
    </sec>
    <sec id="sec-29">
      <title>Aklilu. We would like to thank Dr Amsalu as well as Daniel Yacob of the Geez frontier foundation for making it possible for us to use the dictionary and other resources in this work.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Berhanou</given-names>
            <surname>Abebe. Dictionnaire</surname>
          </string-name>
          Amharique-Francais.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Amsalu</given-names>
            <surname>Aklilu</surname>
          </string-name>
          . Amharic English Dictionary.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Bender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Head</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Cowley</surname>
          </string-name>
          .
          <article-title>The ethiopian writing system</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>William</given-names>
            <surname>Gale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Church</surname>
          </string-name>
          , and David Yarowsky.
          <article-title>One sense per discourse</article-title>
          .
          <source>In the 4th DARPA Speech and Language Workshop</source>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>W.</given-names>
            <surname>Leslau</surname>
          </string-name>
          . Amharic Textbook. Berkeley University, Berkeley, California,
          <year>1968</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          and
          <article-title>Hinrich Schu¨tze. Foundations of Statistical Natural Language Processing</article-title>
          . MIT Press, Cambridge, MA,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Amit</given-names>
            <surname>Singhal</surname>
          </string-name>
          , Chris Buckley, and
          <string-name>
            <given-names>M</given-names>
            <surname>Mitra</surname>
          </string-name>
          .
          <article-title>Pivoted document length normalization</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] URL. http://www.ethnologue.org/,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9] URL. http://lucene.apache.org/java/docs/index.html,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>