<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>2. Related Works</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dictionary-based Thai CLIR: Experimental Survey of Thai CLIR Jaruskulchai Chuleerat Department of Computer Science Faculty of Science Kasetsart University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2001</year>
      </pub-date>
      <abstract>
        <p>This paper describes our work, which participated in the Cross-Language Information Retrieval (CLIR) at the Cross-Language Evaluation Forum. Our objectives for this experiment have three folds. Firstly, the coverage of the Thai-bilingual dictionary was evaluated when translating queries. Secondly, whether the segmentation process has effected the CLIR. Lastly, this research investigates the query formations techniques. Since this is the first international experimental in CLIR, our approach used dictionary-based technique to translate Thai queries into English queries. Four runs are submitted to the CLEF: (a) single mapping translation with manual segmentation, (b) multiple mapping translation with manual segmentation, ( c ) single mapping translation with automatic segmentation and (d) Single mapping with query enhancing with the Thai thesaurus words. The retrieval effectiveness is worse than our expected. The simple dictionary mapping technique is unable to achieve the retrieval effectiveness, although the dictionary lookup gave very good high percentage of mapping word. The words from the dictionary lookup are not specific terms but each is mapped to a definition or meaning of that term. Furthermore, Thai stopword, stemmed word and word separation have effected in Thai CLIR.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Most of the CLIR research community believes that
CLIR would be useful for people who do not speak a
foreign language well. Unfortunately, some of the
Thai CLIR hasn’t evaluated their results with proper
data. Thus, we participate in the Cross-Language
Evaluation Forum (CLEF) as an opportunity for us to
better understanding the issues in the research of the
Cross-Language Information Retrieval (CLIR). We
performed four Thai-English cross language retrieval
runs. Our approach to the CLIR was to translate the
Thai topics into English by using dictionary
mapping. The bilingual dictionaries are LEXiTRON
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and Seasite [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. These two dictionaries were
compiled by the Software and Language Engineering
Laboratory, National Electronics and Computer
Technology Center (NECTEC) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and Northern
Illinois university [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>The objectives of these four runs are follows: to
survey the available of Machine Readable Dictionary
(MRD) and the coverage of the vocabulary, to
investigate the query formation techniques, to
explore the possibility of automatic translation, and
to enhance query by using Thai Thesaurus.</p>
      <p>According to the Thai CLIR’s objective, the four
official runs are the single dictionary mapping,
multiple dictionary mapping, manual and automatic
segmentation, query expansion using Thai
thesauruses.</p>
      <p>The shareable or public MRDs are LEXiTRON from
Software and Language Engineering Laboratory,
National Electronics and Computer Technology
Center and Seasite from Northern Illinois University
were used in our experiment.</p>
      <p>The rest of this paper is as fellows. Section 2 briefly
reported the related fields in the Thai text
information retrieval. Summaries of related Thai
CLIR resources are given. Thai CLIR experimental
design is described in section 4. Experimental results
are presented in the last section.
In this section, the Thai computer processing and the
Thai Natural Language process is briefly discussed
for understanding the current technology, which play
an importance in the CLIR.</p>
    </sec>
    <sec id="sec-2">
      <title>2.1 Computer Processing of Thai</title>
      <p>
        Attempts to work with the Thai language on the
computer started when computers were frist
introduced into the country more than four decades.
There are no any special characters to separate words
from phrase and sentences in the Thai writing
system. To overcome this problem, artificial
intelligent, natural language processing, and
computational linguistic are exhaustively studies.
The accomplishment of these studies established of
machine translation project by the National
Electronics and Computer Technology Center
(NECTEC) in 1980 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Additionally, a number of
research output has been commercially promoted, for
example, hand-held electronic dictionaries and
translators from English to Thai (Pasit) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the Thai
spelling checking software and the word
segmentation programs.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2.2 Thai Text Retrieval</title>
      <p>
        Most of Thai Text Retrieval system is always
coordinated with the segmentation algorithms.
Automatic extraction keyword from the documents is
nontrivial task. Trie Structure along with dictionary
based word segmentation are proposed in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to
solve the unknown words. However, only the
indexing process is presented, there is no report on
the retrieval effectiveness. The work done in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] was
more contributed in the information retrieval method.
The paper presented a number of comparison in
segmentation process, the indexing techniques and
term weighting system for Thai text retrieval. Three
methods of indexing are proposed, ngram-based,
word-based and rule-based. When applying term
weight system, the segmentation process does not
much effect the retrieval performances. All the
performance metric is tested on the Thai news. The
collection size is about 8 MB and 4800 documents.
Additionally, the environment for testing the
hypothesis used SMART text retrieval system from
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The other indexing technique, the signature file,
has been proposed for indexing for Thai Text [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
This paper studied the number of bit for representing
the each document signature and the test collection
was from Thai Holy Bible.
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.3 Works in Thai CLIR</title>
      <p>
        There are some Thai research papers [
        <xref ref-type="bibr" rid="ref11 ref12 ref3">3, 11-12</xref>
        ],
which presented their work in the area of CLIR. All
of their techniques are based on the transliterated
words. The research paper in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] presented
transliterated word encoding algorithm and creating
5000 Thai English personal names. Then, the
retrieval process is against with this database. This
paper claimed that the CLIR effectiveness is 69 and
73% in precision and recall. The second paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is
also from the same research lab to achieve a better
precision and recall over 80% in the CLIR. Their
CLIR model retrieved document containing either the
English or Thai transliterated words using phonetic
codes for keywords and the phonetic coding is based
on Soundex coding of Odell and Russell. Their result
of experiment is compared with the Thai-English
transliterated words which are collected from Royal
Academy in transliteration Guideline, Science
Dictionary, mathematics Dictionary, Chemistry
Boook1: High School Level. Most of those words are
proper nouns, and technical terms. The last paper
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] also presented the transliteration from Thai to
English for solving the loan words. This paper are
more emphatic solving loan word problems such as
non-native accent, information losing and
orthographic translation. There are two processes to
identify load word. First, the explicit unknown words
are recognized by mapping with the Thai dictionary.
Secondly, the hidden unknown words, which are
composed of one or more known words, are identify
by frequency checking. However, it is unclear how
these algorithms are applied to work with CLIR.
In Asian CLIR research, the dictionary-based method
is the well-known method and the query translation
strategy is employed. The work done in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], also
employed the dictionary-based method for
Indonesian-English Cross-Language Text Retrieval.
The local-feedback techniques are applied to expand
the queries terms for improving the retrieval
effectiveness. Their research is conducted on TREC’s
data. Chen and his colleges worked on the Japanese
English cross language. They stated the segmentation
problem of Japanese language, which contain a
number of technical terms. To increase the number
vocabulary, the parallel corpus is employed. They
stated that the retrieval effectiveness of CLIR is
effect by the coverage of term in the dictionary
      </p>
    </sec>
    <sec id="sec-5">
      <title>3. Resource available for Thai CLIR</title>
      <p>
        The most important resource for the CLIR is
bilingual dictionary. In our survey of the bilingual
electronic dictionary, a number of Thai-English
bilingual electronic dictionaries are found, for
examples:- the Thai internet education project [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], an
Online Thai Dictionary (Seasite) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and
LEXiTRON [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Only the last two dictionaries are
able to get the whole electronic form. However, the
Seasite dictionary need to be reencoded since the
original encoding system is different from the current
system. The total number of words in each system is
16,060 and 11,188 words from LEXiTRON and
Seasite respectively. The electronic format of Thai
thesaurus is not available to share for public. We
prepared our own Thai thesaurus from [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Around
20,000 Thai thesaurus words are collected and used
in this research.
      </p>
      <p>
        Another important resource is the Thai segmentation
program. Processing of Thai language has been
working for more than 3 decades. The free resource
for breaking phrase or sentence into words is the
wordbreak from NECTEC [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and from University of
Massachusetts [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], wordbreak, which is from
NECTEC, gave the best the effectiveness of
segmentation process.
      </p>
      <p>Therefore, in our initiative CLIR research, we
deployed the LEXiTRON, and Seasite for checking
the coverage of the number of vocabulary used in the
automatic query translation. Additionally, to be able
to automatically translation, the segmentation process
needs to be verified and the wordbreak (Swath) from
the NECTEC is deployed in Thai CLIR research.</p>
    </sec>
    <sec id="sec-6">
      <title>4. Experimental Design</title>
      <p>The Thai CLEF is aimed at the bilingual task.
English documents are retrieved from the Thai
topics.</p>
      <p>Since Thai language is not an official language in the
CLEF, no topic is provided by CLEF. Thus, the
CLEF’s English topics were chosen and translated by
manual into two types of Thai queries. One is
segmented Thai queries by human and another is like
normal Thai writing system. The Swath’s NECTEC
is used to break phrase or sentence of the
unsegmented Thai queries into words. The
disadvantage of manual translation is that it relies on
human judgment and may be bias. Then we apply the
dictionary mapping techniques to translate the Thai
queries back to English queries. In the dictionary
lookup process, if any words are able to lookup, the
process will leave that word from the topic. Thus, the
concept terms or relevance terms may not include the
topics. The four official runs, which are rely on the
query formation, are as follows.</p>
      <p>(a) Single Mapping: The bilingual Thai-English
definition trend to give several senses or
meanings. Thus, English queries are
translated by using single dictionary
mapping, and only the first map is selected
for translation.
(b) Multiple Mapping: Since the first map was
not always to give the right translation. This
second run, English queries are replaced
with all meaning found in the dictionary.
(c) Single Mapping and Segmentation: The
unsegmented Thai queries are segmented
using the NECTEC wordbreak program and
single mapping is applied for the query
translation.
(d) Query Expansion: Thai thesaurus words are
added to the single mapping queries. The
process of expanded query terms is done
before translation.</p>
      <p>
        For testing the coverage of the number of terms in
electronic dictionary, SEASITE dictionary is used to
translate English queries. Figure 1 shows our
experimental design. The SMART system from
Cornell University [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] is used to measure the
retrieval effectiveness of our Thai CLIR. In all runs,
stop words and stemming were applied to query and
text collection. The term weight was applied to the
document collection.
      </p>
      <p>Documents collection, the Los Angeles Time of
1994, is indexed using the SMART vector model.
English query is indexed based on the long query
format, or on the descriptions, &lt;DESC&gt; marked tag.
Although SMART is based on the vector model, we
do not modify the original topics. When a query was
sent to the system, the 1,000 highest-ranked records
are returned.</p>
      <p>The dictionary terms of the dictionary mapping
algorithm are loaded into MySQL database. The
mapping algorithm is deployed using Java
technology and running on Linux Environment.</p>
      <p>Handed
Segmented
Thai Query</p>
      <p>CLEF</p>
      <p>Topics
Translated
by manual</p>
      <p>Unsegmented</p>
      <p>Thai Query
Segmented by</p>
      <p>Swath
Dictionary Mapping using
bilingual Machine Readale</p>
      <p>Dictionary
Los Angle</p>
      <p>Time</p>
      <p>English
Queries</p>
      <p>Qrel
Evaluating the result
using SMART system</p>
    </sec>
    <sec id="sec-7">
      <title>4. Results</title>
      <p>
        We have learned from [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], the retrieval
effectiveness of CLIR which is based on the
dictionary mapping, will drop about half. For Thai
CLIR, the retrieval results are worse when the
ThaiCLIR was tested with CLEF’s text collection. This
principle of Thai CLIR has been experimented with
ZIFF’s TREC collection. The retrieval results
dropped about 40% [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This section reports the
Thai CLIR results.
(Result in the paper is slightly changed from what we
had submitted to CLEF. We submitted the wrong
data of the first runs, single mapping)
      </p>
      <sec id="sec-7-1">
        <title>4.1 The Coverage of the Vocabulary of the</title>
      </sec>
      <sec id="sec-7-2">
        <title>Dictionary</title>
        <p>
          As mention in section 3, the number of term in
LEXiTRON dictionary is more cover than SEASITE.
However, the effectiveness of both dictionaries is
almost the same and over 90% of words are found.
Figure 2 shows the characteristic of Thai queries.
When applying the single mapping techniques, it
turns out that SEASITE can retrieval a little bit better
than LEXiTRON, which is opposite to our
experiment in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. However, this number is not
significant achievement.
        </p>
      </sec>
      <sec id="sec-7-3">
        <title>4.2 The Effect of Segmentation Algorithm in</title>
      </sec>
      <sec id="sec-7-4">
        <title>CLIR</title>
        <p>
          As mentioned in section 3 for automatically query
translation, the effectiveness of the current
technology for segmenting phrase or sentences into
words needed to verify. In our experiment, it is not
clear that the segmentation has effected in the CLIR.
Though it has been reported in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] that the
segmentation has effected in the CLIR, we are unable
to prove in the experiment. In our discussion, the
different is of the Topic translation techniques from
English to Thai. Our first experiment, the researcher
is lean on dictionary to translate the English to Thai
and try to break words according to the mapping of
word in dictionary. Additionally, some unofficial
report stated that terms in the Thai-bilingual
electronic dictionary are the smallest term with
meaning. Thus, some manual segmented words
cannot found in dictionary. However, the percentage
of number of word found is quite high. Therefore, we
compared the original query and the translation
query, we found that only 15 percent of matching
words. It means that the words, which are found in
the dictionary, are not relevance to the search terms.
        </p>
      </sec>
      <sec id="sec-7-5">
        <title>4.3 The Effectiveness of CLIR</title>
        <p>
          Comparing with our previous results [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], in which
the retrieval effectiveness is around 40% of the
monolingual, there is many different in the design
process and can be summarized as fellows:
1. The manual translation techniques from English
to Thai: As mention in section 4.2, our previous
translation technique is based on the vocabulary
in the dictionary terms.
2. Query length: Our previous query topics are
translated from the &lt;TITLE&gt; tag. The Thai
keywords from the &lt;TITLE&gt; tag are more
relevance to the retrieval system since the
translation process was biased. In this
experiment, the query topics are translated from
&lt;DESC&gt; tag and avoiding consult the
dictionary. Although, there are more terms,
most of the terms are not specific to query or
more general. As we learned from expanding
the query topics with Thai thesaurus, it does not
increase the retrieval effectiveness.
3. Thai stopword and Thai stemming: Relatively
few intensive studies in Thai stopword and
stemming have been reported. Some Thai
stopwords are reported in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. It is not clear
whether Thai language has stemming property.
The words ‘การ:kan:when prefix to a noun, it
indicate action’ and ‘ความ:kwam:prefix to an
adjunctive indicate stats, condition’ are two
Thai prefix. Removing or not removing has
effected in Retrieval effectiveness and is
required language knowledge judgment. Since
removing the Thai prefix, the meaning of the
stem word may not relate to the original
meaning. Therefore adding these stem words
will degrade the retrieval effectiveness.
        </p>
        <p>Thus, to prove the above issues, the Thai topics are
modified by human judgment, some Thai stopwords
are removed from the topics, and choosing the search
terms which can be found in the dictionary. At this
time the retrieval effectiveness is improved 1.5 of the
unremoved Thai stopword query (see Table 2). The
number of relevance retrieval increase to 40% of
monolingual Though, the retrieval effectiveness still
cannot achieve as of other CLIR, which deployed
dictionary mapping techniques.</p>
        <p>Results in Table 2 brought back our confident. It
showed that the number of terms in LEXiTRON is
more coverage than SEASITE. Segmentation process
still is the critical issue in CLIR. However, adding
the Thai thesaurus terms still cannot improve the
retrieval effectiveness. There has some changed in
average precision for individual query.</p>
      </sec>
      <sec id="sec-7-6">
        <title>4.4 Implementation of Thai CLIR</title>
        <p>The algorithm of Thai CLIR has implemented and
opened for publicly try and the web site is
http://www.cs.sci.ku.ac.th/~ThaiIr/CLIR/demo. The
demonstrated web site receives Thai keywords from
users and then translate using single dictionary
mapping. The result of translation is sent back and
allow user for selecting the English keywords. Then,
query is sent to Google or Altravista for searching
English web pages. Furthermore, the results of cross
language retrieval may be translated from English to
Thai by Pasit. This part of the demonstration program
was supported by the NECTEC.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>5. Discussion</title>
      <p>The Thai CLIR faced the same problems as other
MRD CLIR based. The fundamental problems of the
MRD CLIR based are as follows: phrase translation,
polysemy translation, and the coverage of dictionary.
The phrase translation is very critical for Thai CLIR.
Some of Thai words may be classified into sentence
or phrase. Therefore, the phrase translation will be
dependent on the segmentation process. The
researching of segmented algorithm in Thailand can
be classified into two types. The first preferable
segmentation is based on the longest matching. The
second research group will segment text into the
smallest word. Theoretically, these smallest words
can be formed a new word. However, electronic
dictionary is collected based on the first approach.
We also have learned that doing research in the area
of CLIR only knowledge from the information
retrieval but also requires knowledge and resource
from Machine Learning. Although the Machine
Learning project has been activated more than 4
decades, the resources from the MRD still very
limited. The limiting of resources is regarded from
the uncompatible or not ready to disseminate to
public use. Thus, there exists a need to accelerate the
research area. Especially, it needs to set up data
format for the electronic dictionary for reusing the
dictionary.</p>
      <p>We have learned from our demonstration the
ThaiCLIR, users quite satisfy the Thai-CLIR system.
Unfortunately, in research experimental, not all query
translation techniques can achieve. This raises
awareness in the Thai-CLIR area. The basic
infrastructure of Thai-CLIR needs to be stimulated
and urgently needed to further develop.</p>
    </sec>
    <sec id="sec-9">
      <title>6. Future Work</title>
      <p>In this initiative CLIR research, the fundamental of
CLIR research has been established. A number of
research techniques to enhance CLIR performance is
of solving disambiguate terms, detecting the
transliterated word, local feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] ____, Thai Wordbreak Insertion Services, National Electronics and Computer Technology Center, URL:http://ntl.nectec.or.th/services/wordbreak/ (download in June,
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Jaruskulchai</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <article-title>An Automatic Indexing for Thai Text Retrieval</article-title>
          ,
          <source>Ph.D. Thesis</source>
          , George Washington University, U.S.A.,
          <year>Aug 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Suwanvisat</surname>
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Prasijutrakul</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thai-English Cross-Language Transliterated Word Retrieval Soundex Technique</surname>
          </string-name>
          ,
          <fpage>NCSEC2000</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Pirkola</given-names>
            <surname>Ari</surname>
          </string-name>
          ,
          <article-title>The Effects of Query Structure and Dictionary Setups in Dictionary-Based Crosslanguage Information Retrieval</article-title>
          , SIGIR'
          <fpage>98</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Mirna</given-names>
            <surname>Adriani</surname>
          </string-name>
          ,
          <article-title>Dictionary-based CLIR for the CLEF Multilingual Track</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] ____, Online Thai English Dictionary, Northern Illinois University, www.seasite.niu.edu/Thai/home_page/online_thai _dictionaries.
          <source>htm (download in June</source>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] ____,
          <source>The Thai Internet Education Project</source>
          , http://www.cyberc.com/crcl/ehelp/base.htm
          <article-title>(doug@crcl</article-title>
          .chula.edu: Contract person, download in June,
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] ____, LEXiTRON, Thai&lt;-&gt;
          <string-name>
            <surname>English</surname>
            <given-names>Dictionary</given-names>
          </string-name>
          , Software and Language Engineering Laboratory, National Electronics and Computer Technology Center, http://www.links.nectec.or.th/lexit/lex_t.html (download in June,
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9] ____, Parsit, Information Research and Development Division, National Electronics and Computer Technology Center, http://www.links.nectec.or.th/services/parsit/inde x2.
          <source>html (download in June</source>
          ,
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Sophonpanich</surname>
            <given-names>Kalaya</given-names>
          </string-name>
          ,
          <string-name>
            <surname>The</surname>
            <given-names>R</given-names>
          </string-name>
          &amp;
          <article-title>D Activities of MT in Thailand, The National Electronics</article-title>
          and Computer Technology Center, Bangkok, Thailand.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Suwanvisat</given-names>
            <surname>Prayut</surname>
          </string-name>
          and
          <string-name>
            <given-names>Prasitjutrakul</given-names>
            <surname>Somchai</surname>
          </string-name>
          ,
          <article-title>Transliterated Word Encoding and Retrieval Algorithms for Thai-English Cross-Language Retrieval</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kawtrakul</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deemagarn</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thumkanon</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khantonthong</surname>
            <given-names>N</given-names>
          </string-name>
          and
          <string-name>
            <given-names>McFetridge</given-names>
            <surname>Paul</surname>
          </string-name>
          ., Backward Transliteration for Thai Document Retrieval,
          <source>Natural Language Processing and Intelligent Information System Technology</source>
          , Research Laboratory, Dept. of Computer Engineering, Kasetsart University, Bangkok, Thailand.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <article-title>Yuen Phuwarawan and team, Thai Thesaurus</article-title>
          , in Thai, Ed publisher.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Adriani</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>and Croft</given-names>
            <surname>Bruce</surname>
          </string-name>
          ,
          <article-title>The Effectiveness of a Dictionary-Based Technique for Indonesian-English Cross-Language Text Retrieval, Center for Intelligent Information Retrieval</article-title>
          , Computer Science Department, University of Massachusetts, USA.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Kanlayanawat</surname>
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Prasitjutrakul</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Automatic Indexing for Thai Text with Unknown Words using Trie Structure</article-title>
          , Department of Computer Engineering, Chulalongkorn University.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>[16] SMART, ftp.cs.cornell.edu/pub/smart/smart.11.0.tar.z</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Charoenkitkarn</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Udomporntawee</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Optimal Text Signature Length for Word Searching on Thai Holy Bible(in Thai)</article-title>
          .
          <source>Proceeding of Electrical Engineering Conference</source>
          , KMUTT, Bangkok,
          <year>November 1998</year>
          ,
          <fpage>549</fpage>
          -
          <lpage>552</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Sripimonwan</surname>
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Jaruskulchai</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <article-title>CrossLanguage Retrieval from Thai to English (in Thai), to be summated to The Fifth National Computer Science</article-title>
          and Engineering Conference, Thailand.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>