<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Multi-Aspect Comparison and Evaluation on Thai Word Segmentation Programs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chaluemwut Noyunsan</string-name>
          <email>chaluemwut@kkumail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Choochart Haruechaiyasak</string-name>
          <email>choochart.haruechaiyasak@nectec.or.th</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Seksan Poltree</string-name>
          <email>seksan@morange.co.th</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kanda Runapongsa Saikeaw</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Engineering, Faculty of Engineering, Khon Kaen University 123 Mittrapap</institution>
          ,
          <addr-line>Mueang, Khon Kaen 40002</addr-line>
          ,
          <country country="TH">Thailand</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Human Language Technology Laboratory (HLT) National Electronics and Computer Technology Center (NECTEC)</institution>
          ,
          <addr-line>Thailand Science Park, Klong Luang, Pathumthani 12120</addr-line>
          ,
          <country country="TH">Thailand</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Morange Solution Company Limited</institution>
          ,
          <addr-line>456/1 Klang Mueang, Mueang, Khon Kaen 40000</addr-line>
          ,
          <country country="TH">Thailand</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Word segmentation is an important task in natural language processing, especially for languages without word boundaries, such as Thai language. Many Thai word segmentation programs have been developed. Researchers and developers in Thai documents usually spend a tremendous amount of time in studying and trying di erent Thai word segmentation programs. This paper presents the performance of six Thai word segmentation programs which include Libthai, Swath, Wordcut, CRF++, Thaisemantics, and Tlexs. Based on experimental results, we compare these programs in terms of usage, response time, time outs, and relevance.</p>
      </abstract>
      <kwd-group>
        <kwd>Word segmentation</kwd>
        <kwd>term tokenization</kwd>
        <kwd>software tools</kwd>
        <kwd>natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Natural Language Processing (NLP) enables computers to understand human
languages. It consists of many processes such as word segmentation,
part-ofspeech tagging, automatic summarization, and speech synthesis. Most NLP
applications require input text to be segmented into words before being processed
?? Corresponding author
further. For example, in sentences similarity application, text must rst be
tokenized into a series of terms before being analyzed grammatically and
semantically. Word segmentation is an essential part for Asian languages such as Thai,
Chinese, Japanese, and Korean. This is because these languages are written
without delimiter spaces for words in the same sentence.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Thai Word Segmentation Programs</title>
      <p>
        Many researches and several programs have been developed for word
segmentation. We chose six Thai word segmentation programs which were Libthai, Swath,
Wordcut, CRF++, Thaisemantics, and Tlexs. They were selected because they
were actively maintained and widely used. Libthai [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a set of C
programming function to support Thai word segmentation. Swath [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and Wordcut [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
are command-line programs. CRF++ is a tool for supporting Condiction
Random Field (CRF) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Thaisemantics has been developed by using Restful web
service [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Tlexs uses CRF to train models for segmentation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experimental Analysis</title>
      <p>Usage Table 1 summarizes the features about usage, o ine support, and whether
installation is needed.
Response Time We ran the experiments and measured response times on
the computer with Intel Dual 1.73GHz and 3 GB of RAM using Ubuntu 64bit
as an operating system. The program execution times are shown in Table 2.
Libthai performed the best with the response time 0.022 seconds while Swath
and CRF++ had response times as 0.024 seconds. On the other hand, the
responses times of Thaisemantics, Wordcut, and Tlexs were large. This is because
Thaisemantics and Tlexs are programs that are used over the internet thus their
response times depend on the internet bandwidth and tra c time. Wordcut is
implemented using JavaScript language which may increase the response time
of the program.
Number of Time Outs Tlexs and Thaisemantics had several time outs
because they are called over the internet. Tlexs had average 2,070 time outs while
Thaisemantics had 5 time outs. Tlexs caused a large number of occurrences of
time outs might be due to server settings or errors.</p>
      <p>Relevance Messages from the BEST were sent to each program and the word
segmentation output was kept in a list. After that, we compared this log list
with a correct list from BEST by using the percentage of precision, recall and
F-measure. We ran this test by using 5-fold cross-validation, and then computed
the average value as shown in Table 3. Both Tlexs and CRF++ have the best
F-measure because they use CRF.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>This paper presents the comparison of six Thai word segment programs in terms
of usage, response time, time outs, and relevance. Swath, Libthai, and CRF++
programs provide the smallest response times because they are native programs.
Thaisemantics yields the largest response time because Thaisemantics is called
over the internet and uses a dictionary. Although Tlexs is also called over the
internet, it has better response time because it uses CRF. Both Tlexs and CRF++
give the highest F-measure because they employ CRF.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>T.</given-names>
            <surname>Karoonboonyanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Silpa-Anan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kiatisevi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Veerathanabutr</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ampornaramveth</surname>
          </string-name>
          , Libthai Library,
          <source>retrieved on Jul 1</source>
          ,
          <year>2014</year>
          from http://linux.thai. net/projects/libthai.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kudo</surname>
          </string-name>
          , CRTF++:
          <source>Yet Another CFT toolkit retrieved on July 2</source>
          ,
          <year>2014</year>
          from http: //crfpp.googlecode.com/svn/trunk/doc/index.html?source=navbar.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>National</given-names>
            <surname>Electronics</surname>
          </string-name>
          and
          <article-title>Computer Technology Center (NECTEC), BEST: Bencharmk for Enhancing the Standard for Thai Language Processing retrived on Jul 5, 2014 from http</article-title>
          ://thailang.nectec.or.th/best/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>National</given-names>
            <surname>Electronics</surname>
          </string-name>
          and Computer Technology Center (NECTEC),
          <source>TLexs: Thai Lexeme Analyser, retrieved on Jul 3</source>
          ,
          <year>2014</year>
          from http://sansarn.com/tlex/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Poltree</surname>
          </string-name>
          ,
          <article-title>Thaisemantics: Free Thai Language Resources</article-title>
          and Services,
          <source>retrieved on Jul 2</source>
          ,
          <year>2014</year>
          from http://www.thaisemantics.org/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>V.</given-names>
            <surname>Satayamas</surname>
          </string-name>
          ,
          <source>wordcut program retrieved on Jul 3</source>
          ,
          <year>2014</year>
          from https://github.com/ veer66/wordcut.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>P.</given-names>
            <surname>Charoenpornsawat</surname>
          </string-name>
          , SWATH - Thai
          <source>Word Segmentation, retrieved on Jul 1</source>
          ,
          <year>2014</year>
          from http://www.cs.cmu.edu/~paisarn/software.html.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>