<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ukrainian and their</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Acquisition of medical terminology for Ukrainian from parallel corpora and Wikipedia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thierry Hamon</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>LIMSI-CNRS</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Orsay</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sorbonne Paris Cité</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K?QM!HBbX</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>U Lille</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <volume>34</volume>
      <issue>267</issue>
      <fpage>71</fpage>
      <lpage>80</lpage>
      <abstract>
        <p>The increasing availability of parallel bilingual corpora and of automatic methods and tools for their processing makes it possible to build linguistic and terminological resources for low-resourced languages.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        with French and English terms.
The acquisition of terminology has gone through a
very active period and provides nowadays several
automatic tools and methods
        <xref ref-type="bibr" rid="ref16 ref7">(Kageura and Umino,
1996; Cabré et al., 2001; Pazienza et al., 2005)</xref>
        for
several European languages and Japanese.
Nevertheless, other languages remain low-resourced
and require specific Natural Language Processing
(NLP) developments.
      </p>
      <p>Our main objective is to create terminological
resources for Ukrainian, for which very little
digitized or electronic resources are available. Yet, the
terminology extraction tools usually require the
morpho-syntactic tagging of texts, which can be
problematic if the corresponding automatic tools
are not available for a given language. For
in</p>
    </sec>
    <sec id="sec-2">
      <title>Natalia Grabar</title>
      <p>UMR8163 STL</p>
      <p>France
does not perform the syntactic and morphological
disambiguation of the tags. Hence, it becomes
impossible to use it for the pre-processing of corpora
before the traditional terminology acquisition
process.</p>
      <p>In this situation, we propose first to compile
terminological resources for Ukrainian in order to
build the basis for the observation of the
specificities of terminological units in this language. Such
observations will allow to develop and parameter
the terminology extraction tool for Ukrainian.</p>
      <p>The motivation of our work is double. We want</p>
      <sec id="sec-2-1">
        <title>1. automatically build terminologies for</title>
      </sec>
      <sec id="sec-2-2">
        <title>Ukrainian,</title>
        <p>2. design specific methods for the acquisition of
such terminological resources.</p>
        <p>The work is carried out with medical data, and in
three languages (Ukrainian, French, and English).</p>
        <p>The work we present starts from the exploitation
of two kinds of corpora (Section 2.1): Wikipedia
in Ukrainian which provides several useful kinds
of information (such as term labels and their codes)
with a high level of quality, and the parallel corpus
MedlinePlus. The term detection and extraction
can be either manual or automatic. Since, there
is no appropriate POS-tagging and term extraction
tools for Ukrainian, we propose to use such tools in
French and English, and to take advantage of these
to transfer English and French extracted terms on
the Ukrainian corpus.</p>
        <p>
          Indeed, the transfer methodology can be
considstance, the UGtag Part-of-Speech (POS) tagger
ered as suitable for such objectives. Suppose we
have parallel and aligned corpora with two
languages L1 and L2, and we have several types of
syntactic or semantic annotations and information
associated to L1. The transfer approach permits
to transpose these annotations or information from
L1 to L2, and to obtain in this way the
corresponding annotations and information in the L2
text. From this point of view, L1 is considered as
the source language while L2 is considered as the
target language. This kind of approach is
particularly interesting when working with low-resourced
languages for which less tools and semantic
resources are available. An increasing
availability of parallel bilingual corpora, and of automatic
methods and tools for their processing makes it
possible to build linguistic and terminological
resources using the transfer methodology
          <xref ref-type="bibr" rid="ref10 ref22">(Yarowsky
et al., 2001; Lopez et al., 2002)</xref>
          . Very few works
have been done in this direction, and we assume
they open novel and efficient ways for the
processing of multilingual texts in particular from
lowresourced languages
          <xref ref-type="bibr" rid="ref11 ref23">(Zeman and Resnik, 2008;
McDonald et al., 2011)</xref>
          . Notice that the modeling
of cross-language features aims at using
languageindependent features to create various types of
annotations. Among such features, we can mention
part-of-speech, semantic categories or even
acoustic and prosodic features.
        </p>
        <p>We propose to apply this method for the
acquisition of bilingual or trilingual terminologies
involving Ukrainian. In our work, each corpus is
exploited through dedicated methods. The
MedlinePlus corpus provides the basis for the building of
the terminology, while the Wikipedia corpus
permits to enrich this information and helps the
wordlevel alignment of the MedlinePlus corpus.</p>
        <p>
          Terminology-related research on Ukrainian is
an active area, although the main
terminological work shows mainly theoretical and
linguistic orientation
          <xref ref-type="bibr" rid="ref14 ref15 ref25 ref4 ref6">(Коссак, 2000; Dmytruk, 2009;
Рожанківський and Кузан, 2000; Ivashchenko,
2013; Oliinyk, 2013)</xref>
          . Very few works are
oriented on the use of terminologies and their
automatic processing, such as the software localization
          <xref ref-type="bibr" rid="ref18">(Shyshkina et al., 2010)</xref>
          .
        </p>
        <p>In the following of this paper, we first present
the material used for the acquisition of bilingual
terminology (section 2), and the methods designed
for achieving this objective (section 3). We then
discuss the results we obtain (section 4), and
conclude with directions for the future work
(section 5).
2
2.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Material</title>
      <sec id="sec-3-1">
        <title>Corpora</title>
        <sec id="sec-3-1-1">
          <title>We use two kinds of corpora:</title>
          <p>• MedlinePlus: parallel medical corpus from
MedlinePlus. These data are built by
MedlinePlus from the National Library of
Medicine1. They contain patient-oriented
1www.nlm.nih.gov/medlineplus/healthtopics.html
Corpus
Wikipedia/UKmed
MedlinePlus/UK
MedlinePlus/FR
MedlinePlus/EN</p>
          <p>Size (occ of words)
246,368,411
43,184
53,067
46,544
brochures on several medical topics (body
systems, disorders and conditions, diagnosis
and therapy, demographic groups, health and
wellness). These brochures have been
created in English and then translated in several
other languages, among which French and
Ukrainian;
• Wikipedia: medicine-related articles from
Wikipedia. This corpus is extracted from
the Ukrainian part of the Wikipedia
using medicine-related categories, such as
Медицина (medicine) or Захворювання
(disorders). The corpus potentially covers a wide
range of medical notions. In Figure 1, we
indicate an example of the source pages which
propose the navigation frame on the left, the
text with explanations and the infobox with
illustration and coding on the right.</p>
          <p>In Table 1, we indicate the size of the corpora. Not
surprisingly, the Wikipedia corpus is much larger
although only part of its information is exploited,
as we will see in the next section.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>UMLS: Unified Medical Language</title>
      </sec>
      <sec id="sec-3-3">
        <title>System</title>
        <p>
          The UMLS (Unified Medical Language System)
          <xref ref-type="bibr" rid="ref9">(Lindberg et al., 1993)</xref>
          merges several (over
100) biomedical terminologies, such as
inter
          <xref ref-type="bibr" rid="ref13">national terminologies MeSH (NLM, 2001</xref>
          ) and ICD
(Brämer, 1988). Such international
terminologies may exist in several languages. For instance,
French and English versions of MeSH are included
in the UMLS. No terminologies in Ukrainian are
part of the UMLS. Each UMLS term is provided
with unique identifiers, which allows to find the
corresponding terms in other terminologies or
languages.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <p>The methods we propose for the extraction of
bilingual terminology are adapted to each kind of
corpora and of data they contain: the
MedlinePlus corpus (section 3.1) and the Wikipedia
corpus (section 3.2). We then present their
crossfertilization (Section 3.3), and the evaluation of the
results (Section 3.4).
3.1
Prior to the exploitation of the MedlinePlus data,
the documents are first transformed in a suitable
format:
• the source PDF documents are converted in
the text format;
• in each language, the documents are
segmented in paragraphs;
• the alignments French/Ukrainian and
English/Ukrainian are generated, in which nth
paragraph from one language is associated
with the nth paragraph from the other
language;
• the alignment between the two pairs of
languages is then verified manually.</p>
      <p>In Figure 2, we present an excerpt from the
English/Ukrainian aligned corpus.</p>
      <p>
        Then in French and English, we can use the
existing terminology extraction tools which results
bootstrap the acquisition of bilingual terminology.
Hence, we use the YATEAterm extractor
        <xref ref-type="bibr" rid="ref1">(Aubin and
Hamon, 2006)</xref>
        , that is applied to documents
POStagged. The extracted terms are then projected on
the French and English corpora. In Figure 2,
candidate terms are marked in bold.
      </p>
      <p>The exploitation of the MedlinePlus parallel and
aligned corpus is performed in several ways
(Figure 3).</p>
      <p>Transfer 1 First, the simplest situation is when
the two aligned lines contain term candidates in
either language: these terms are recorded as
candidates for the alignments. For instance, in Figure
2, the pairs {Tiredness, Втома} and {Pain, Біль}
are issued from this kind of alignment.</p>
      <p>
        Transfer 2 Secondly, when the paragraphs
contain complex expressions or sentences, the
processing is done as follows (Figure 4):
1. the paragraph-aligned corpora are aligned at
the word level using GIZA++
        <xref ref-type="bibr" rid="ref14 ref25">(Och and Ney,
2000)</xref>
        ,
English
Cancer cells grow and divide more quickly than
healthy cells. Cancer treatments are made to
work on these fast growing cells.
- Tiredness
- Nausea or vomiting
- Pain
- Hair loss called alopecia
Ukrainian
Ракові клітини ростуть і діляться швидше,
ніж здорові клітини. При лікуванні раку
здійснюється вплив на ці клітини, що
швидко ростуть.
- Втома
- Нудота або блювота
- Біль
- Втрата волосся, що називається алопецією
3. the alignments extracted are recorded as
candidates for building the bilingual
terminology.
      </p>
      <p>For instance, in Figure 2, the term Cancer cells
is automatically extracted from the English
corpus. GIZA++ proposes that Cancer cells is aligned
with Ракові клітини. Thus, through the
wordaligned text, we can propose that Cancer cells is
the translation of Ракові клітини. This
processing is performed on the two pairs of languages
(French/Ukrainian and English/Ukrainian).</p>
      <p>As indicated in Table 1, the size of our corpora is
rather small for the statistical alignment performed
by GIZA++. For this reason, we provide GIZA++
with a bilingual dictionary in order to help the
alignment at the word level (see Section 3.3).
Besides, in preliminary experiments, we also observe
that word level alignment errors lead to the
extraction of Ukrainian stopwords as term candidates (на
(on), або (or), etc.). To remove such obvious
errors, we filter out such candidates if they occur in a
list of 385 stop-word forms issued from an existing
resource dedicated to the localization of graphical
interfaces2.
3.2</p>
      <sec id="sec-4-1">
        <title>Extraction of bilingual terminology from the Wikipedia corpus</title>
        <p>
          The Wikipedia corpus is used to complete and to
help the method applied to the MedlinePlus
corpus. The content we propose to exploit is included
in infoboxes (on the right in Figure 1) and is
reachable through the MediaWiki source code of the
Wikipedia. This provides the label of the
medical terms in Ukrainian and their MeSH codes. The
process is the following (Figure 4):
1. the infobox content is extracted and parsed3
in order to obtain the term label and its MeSH
code,
glish terms,
ogy.
2. the MeSH code is used to query the UMLS,
and to get the corresponding French and
En3. the term pairs French/Ukrainian and
English/Ukrainian are then built and provide
good candidates for the bilingual
terminolThis part of the method exploits specific and
intentionally created content for a given medical
notion in Ukrainian: term for a given medical
notion and its MeSH code. This information is
reliable. For instance, in Figure 1, the term нанізм
is extracted, as well as its MeSH code D004392.
Through the UMLS, the corresponding English
terms are dwarfism and nanism, while the
corresponding French term is nanisme. Notice that
similar method has been used for the building of
medical terminology in the Arabic language
          <xref ref-type="bibr" rid="ref21">(Vivaldi
and Rodríguez, 2014)</xref>
          .
bKi2`flFBMQTr/Xt
(?iT,fb2`+XMQ;$x#
2hti@J/BrF6Q`K
2
i?Tb,f;Bm#X+QK7HtM
        </p>
        <sec id="sec-4-1-1">
          <title>3We use the Perl module h2ti,J/BrF6Q`K</title>
          <p>)</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Ukrainian Wikipedia medical part</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>Processing of the InfoBoxes</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Medical terms with MeSH codes</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>Querying UMLS</title>
        </sec>
        <sec id="sec-4-1-6">
          <title>UMLS</title>
        </sec>
        <sec id="sec-4-1-7">
          <title>Pairs of medical terms (UK/FR and UK/EN)</title>
          <p>• the single-word terms extracted by other
approaches can be provided to GIZA++, as
an additional bilingual dictionary, in order
to help the alignment of MedlinePlus at the
word level.</p>
          <p>During preliminary experiments, we test several
combinations of parameters for the pre-processing
and the alignments.</p>
        </sec>
        <sec id="sec-4-1-8">
          <title>While pre-processing the</title>
          <p>
            French corpus, the Part-of-Speech is performed by
TreeTagger
            <xref ref-type="bibr" rid="ref17">(Schmid, 1994)</xref>
            and can be improved
by the morphological analyzer Flemm
            <xref ref-type="bibr" rid="ref12">(Namer,
2000)</xref>
            . We also experiment with the use of
GeniaTagger
            <xref ref-type="bibr" rid="ref20">(Tsuruoka et al., 2005)</xref>
            on the English
corpus. We also experiment with the use of the terms
extracted from Wikipedia, or by the MedlinePLus
method Transfer 1, or both, for guiding the Giza++
alignment.
          </p>
          <p>Thus, based on the results of the preliminary
experiments, we choose to pre-process the
English corpus with TreeTagger and the French
corpus with TreeTagger and Flemm.</p>
        </sec>
        <sec id="sec-4-1-9">
          <title>Single-word terms extracted from Wikipedia and by the method Transfer 1 are used as bilingual dictionary to help</title>
          <p>the Giza++ word level alignment. We only present
the results obtained with this configuration in the
following.
3.4</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation</title>
        <p>The evaluation is performed manually in order to
check whether the candidates extracted for
building the bilingual terminologies are correct. It has
been performed by an Ukrainian native speaker
having knowledge in medical informatics. Terms
are validated independently in each language, but
we also evaluation the bilingual and trilingual
relations between the Ukrainian, English and French
terms. With this kind of evaluation, precision of
the results can be computed, i.e. the ratio between
the correct answers and all the answers.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>Table 2 presents the results and the precision for
the extracted terms by the three methods. Table 3
presents the results and the precision concerning
the pairs and triples of terms.
4.1</p>
      <sec id="sec-5-1">
        <title>Extraction of bilingual terminology from the Wikipedia corpus</title>
        <p>The exploitation of the Wikipedia infobox allow
to collect 357 Ukrainian medical terms among
which 177 are single-word terms. By querying
UMLS with the MeSH codes, those terms are
associated with 1428 French terms (among them,
339 single-word terms) and 3625 English terms
(among them, 448 single-word terms). The
number of French and English terms compared to the
number of Ukrainian terms are due to the
synonyms proposed by MeSH. As for the bilingual
pairs of terms, we obtain 1,515 Ukrainian/French
term pairs and 3,789 Ukrainian/English term pairs,
including, respectively, 270 and 405 pairs between
single-word terms. Since each Ukrainian term is
associated with at least one French and English
terms, this allows to build 28,840 triples. We
consider that the precision of this terminology is 1
because the collecting manner.
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Extraction of bilingual terminology from the MedlinePlus corpus</title>
        <p>The use of the first method of transfer (Transfer 1)
allows to extract 436 Ukrainian terms with a high
precision unsurprisingly (0.966). These terms
are associated with 316 French terms and 354
English terms in 282 triples between Ukrainian,
French and English terms, 63 pairs only
between Ukrainian and French terms and 115 pairs
only between Ukrainian and English terms, with
0.954, 0.937 and 0.965 precision, respectively.
Thus, the Transfer 1 method allows to collect
334 Ukrainian/French term pairs (among them
108 pairs between single-word terms) and 380
Ukrainian/English term pairs (among them 135
pairs between single-word terms). We observe
that these relations can involve synonyms in either
language: {фаллопієва труба, trompes de
fallope/trompe utérine} (fallopian tube), {втрата
слуху/втрачається слух, hearing loss}, {втома,
fatigue/tiredness}. Besides, in Ukrainian, several
case forms can be associated to a same English
of French form: {вагітність, pregnancy} and
{вагітності, pregnancy}.</p>
        <p>As the precision values suggest, this first
transfer method leads to few errors. Their analysis
shows that they mainly concern partial match
between one language and another involved by the
translation: {появу виразок у роті, mouth sores}
-- lit. (appearance of) mouth sores, {ви можете
спати, dormir/sleep} -- lit. you can sleep. The
silence of the method can be explained by two
reasons. First, again the variation due to the
translation prevents the transfer 1 method to extract term
in French or English. For instance, since the title
Soins in the French corpus is the English
translation of Your care, the French term matches with
the line, contrary to the English term. The
Transfer 2 method will solve this problem. However, the
main reason of the silence is the incapacity of the
term extractor to identify French or English terms
because its extraction strategy or errors in the POS
tagging.</p>
        <p>As for the second transfer method (transfer 2),
we present the results obtained when the pairs of
single-words terms issued from the MedlinePlus
corpus and from Wikipedia are used to help the
GIZA++ alignment. In that context, the
transfer 2 method allows to extract 9,040 Ukrainian
terms with 0.454 precision (exact match).
Precision of the French and English terms is higher:
0.674 and 0.761 respectively (exact match).
Moreover, the number of French and English terms
is dramatically lower (about -45% and -40%)
than in Ukrainian: the rich morphology of the
Ukrainian language provides several inflected
Prec.
1
0.966
0.998
0.454
0.84
0.481
Source
Wikipedia
MedlinePlusT ransfer1</p>
        <p>inexact match
MedlinePlusT ransfer2</p>
        <p>inexact match</p>
        <sec id="sec-5-2-1">
          <title>Total Total of correct terms</title>
          <p>forms for a given term ({напад, нападу} --
attack, {припадків, припадки} -- seizure, {костей,
кістки} -- bones). Besides, the method allows
also to extract synonymous terms ({приступам,
припадків} -- attacks/seizures, {биття, удару}
-- beats). The precision values with the inexact
match (the correct term is included or includes the
term candidates) are much higher and gain 0.40
points for the Ukrainian terms and 0.05 for the
French and English terms. We assume this
difference on Ukrainian candidate terms is mainly due to
the alignment quality. As for the interlingual
relations, the Transfer 2 method collects 3,724 pairs
of Ukrainian/French terms with 0.309 precision,
4,745 pairs of Ukrainian/English terms with 0.401
precision and 4,724 triples with 0.419 precision.</p>
          <p>An analysis of the results shows that most of
the errors are due to the alignment problems.
Indeed, we observe that when the alignment is
correct, the Ukrainian terms are correctly extracted by
the transfer. Otherwise, the errors occur.</p>
          <p>Moreover, even if the documents
(patientoriented brochures) are not highly specialized,
most of the extracted terms are specific to the
medical domain ({трахеотомією, tracheostomy}),
{фактори ризику, risk factors}, {шприца,
syringe}, {холестерину, cholesterol}). Other terms
also refer to close and approximating notions
which reflects this type of documents: {діти,
children}, {здорову їжу, healthy diet}, {серцевий
напад, heart attack}, {склянок рідини, glasses of
liquid}. An interesting observation is that some
French and English terms correspond to
propositions in Ukrainian: {не до кінця приготовлену
їжу, undercooked foods} (lit. food which is not
fully cooked), {При цьому обстеженні Ви не
відчуєте жодного болю, indolore (painless)} (lit.
With this exam you will feel no pain).</p>
          <p>Finally, all the methods combined allow to
build a terminological resource containing 4,588
Ukrainian medical terms and their 34,267 relations
with French and English terms.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>In this work, we propose to exploit two kinds of
freely available multilingual corpora in French,
English and Ukrainian. Each corpus is exploited
with appropriate methods which allows to
extract the term candidates and to create term pairs
Ukrainian/French and Ukrainian/English. In
particularly, French and English corpora are
processed with NLP and term extraction tools. Then,
thanks to the transfer methods these terms are
transposed on the Ukrainian language. We also
propose to use existing terminologies and to
exploit simple terms for improving the alignment
performed at the word level with GIZA++.</p>
      <p>
        Our future work will address the enrichment of
the created resource with terms from other
corpora. Besides, in the Wikipedia corpus, we can use
other codes, such as those from МКХ-10 (ICD10)
or MedlinePlus. This will also augment the
coverage of the term pairs extracted in the current work.
Another perspective of this work is the
improvement of the bilingual alignment of documents at
the word level. In that respect, we plan to
investigate the use of other alignment algorithms, such as
Fast-Align
        <xref ref-type="bibr" rid="ref5">(Dyer et al., 2013)</xref>
        or the Lingua::Align
toolbox
        <xref ref-type="bibr" rid="ref19 ref8">(Tiedemann and Kotzé, 2009)</xref>
        . Other
curators will be involved. Further improvements of
the proposed transfer method can be obtained with
statistical and morphological cues.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work is funded by the LIMSI-CNRS AI
project Outiller l'Ukranien. We are thankful to the
reviewers for their useful comments which
permitted to improve the quality of the paper.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>S</given-names>
            <surname>Aubin</surname>
          </string-name>
          and
          <string-name>
            <given-names>T</given-names>
            <surname>Hamon</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Improving term extraction with terminological resources</article-title>
          .
          <source>In FinTAL</source>
          <year>2006</year>
          ,
          <article-title>number</article-title>
          4139 in LNAI, pages
          <fpage>380</fpage>
          --
          <lpage>387</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>GR</given-names>
            <surname>Brämer</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>International statistical classification of diseases and related health problems. tenth revision</article-title>
          .
          <source>World Health Stat Q</source>
          ,
          <volume>41</volume>
          (
          <issue>1</issue>
          ):
          <fpage>32</fpage>
          --
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>MT</given-names>
            <surname>Cabré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Estopà</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J</given-names>
            <surname>Vivaldi</surname>
          </string-name>
          ,
          <year>2001</year>
          .
          <article-title>Automatic term detection: a review of current systems</article-title>
          , pages
          <fpage>53</fpage>
          --
          <lpage>88</lpage>
          . John Benjamins.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Veronica</given-names>
            <surname>Dmytruk</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Typological features of word-formation in computing, the internet and programming in the first decade оf the XXI century</article-title>
          .
          <source>In УДК</source>
          , pages
          <fpage>1</fpage>
          --
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Chris</given-names>
            <surname>Dyer</surname>
          </string-name>
          , Victor Chahuneau, and
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A simple, fast, and effective reparameterization of ibm model 2</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>NAACL</given-names>
          </string-name>
          /HLT, pages
          <fpage>644</fpage>
          --
          <lpage>648</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>VL</given-names>
            <surname>Ivashchenko</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Historiography of terminology: metalanguage and structural units</article-title>
          .
          <source>In UDC</source>
          , pages
          <fpage>1</fpage>
          --
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>K</given-names>
            <surname>Kageura and B Umino</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Methods of automatic term recognition</article-title>
          .
          <source>In National Center for Science Information Systems</source>
          , pages
          <fpage>1</fpage>
          --
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Natalia</given-names>
            <surname>Kotsyba</surname>
          </string-name>
          , Andriy Mykulyak, and
          <string-name>
            <surname>Ihor</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Shevchenko</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Ugtag: morphological analyzer and tagger for the ukrainian language</article-title>
          .
          <source>In Proceedings of the international conference Practical Applications in Language and Computers (PALC</source>
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>DA</given-names>
            <surname>Lindberg</surname>
          </string-name>
          ,
          <source>BL Humphreys, and AT McCray</source>
          .
          <year>1993</year>
          .
          <article-title>The unified medical language system</article-title>
          .
          <source>Methods Inf Med</source>
          ,
          <volume>32</volume>
          (
          <issue>4</issue>
          ):
          <fpage>281</fpage>
          --
          <lpage>291</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Lopez</surname>
          </string-name>
          , Mike Nossal, Rebecca Hwa, and
          <string-name>
            <given-names>Philip</given-names>
            <surname>Resnik</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Word-level alignment for multilingual resource acquisition</article-title>
          .
          <source>In LREC Workshop on Linguistic Knowledge Acquisition and Representation: Bootstrapping Annotated Data</source>
          , Las Palmas, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Keith</given-names>
            <surname>Hall</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Multi-source transfer of delexicalized dependency parsers</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP '11</source>
          , pages
          <fpage>62</fpage>
          --
          <lpage>72</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>F</given-names>
            <surname>Namer</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>FLEMM : un analyseur flexionnel du français à base de règles</article-title>
          .
          <source>Traitement automatique des langues (TAL)</source>
          ,
          <volume>41</volume>
          (
          <issue>2</issue>
          ):
          <fpage>523</fpage>
          --
          <lpage>547</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>National Library of Medicine, Bethesda, Maryland</source>
          ,
          <year>2001</year>
          . Medical Subject Headings. www.nlm.nih.gov/mesh/meshhome.html.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>FJ</given-names>
            <surname>Och and H Ney</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Improved statistical alignment models</article-title>
          .
          <source>In ACL</source>
          , pages
          <fpage>440</fpage>
          --
          <lpage>447</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>OY</given-names>
            <surname>Oliinyk</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Terminology for description of linguistic landscape in native and foreign linguistics</article-title>
          .
          <source>Terminolohichnyi visnyk</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          --
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Pazienza</surname>
          </string-name>
          , Marco Pennacchiotti, and Fabio Massimo Zanzotto.
          <year>2005</year>
          .
          <article-title>Terminology extraction: An analysis of linguistic and statistical approaches</article-title>
          . In Spiros Sirmakessis, editor,
          <source>Knowledge Mining</source>
          , volume
          <volume>185</volume>
          <source>of Studies in Fuzziness and Soft Computing</source>
          , pages
          <fpage>255</fpage>
          --
          <lpage>279</lpage>
          . Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>H</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>44</fpage>
          --
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Nataliia</given-names>
            <surname>Shyshkina</surname>
          </string-name>
          , Galina Zorko, and
          <string-name>
            <given-names>Larisa</given-names>
            <surname>Lesko</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Terminology work and software localization in Ukraine</article-title>
          . In Problems of Cybernetics and Informatics, pages
          <fpage>17</fpage>
          --
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Jörg</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Gideon</given-names>
            <surname>Kotzé</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A discriminative approach to tree alignment</article-title>
          . In Iustina Ilisei, Viktor Pekar, and Silvia Bernardini, editors,
          <source>Proceedings of the International Workshop on Natural Language Processing Methods</source>
          and
          <article-title>Corpora in Translation, Lexicography and Language Learning (in connection with</article-title>
          <source>RANLP'09)</source>
          , pages
          <fpage>33</fpage>
          --
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Yoshimasa</given-names>
            <surname>Tsuruoka</surname>
          </string-name>
          , Yuka Tateishi,
          <string-name>
            <surname>Jin-Dong</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Tomoko Ohta,
          <string-name>
            <surname>John McNaught</surname>
            ,
            <given-names>Sophia</given-names>
          </string-name>
          <string-name>
            <surname>Ananiadou</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jun'ichi Tsujii</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Developing a robust part-of-speech tagger for biomedical text</article-title>
          .
          <source>LNCS</source>
          ,
          <volume>3746</volume>
          :
          <fpage>382</fpage>
          --
          <lpage>392</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>J</given-names>
            <surname>Vivaldi</surname>
          </string-name>
          and
          <string-name>
            <given-names>H</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Arabic medical term compilation from Wikipedia</article-title>
          .
          <source>In Proc of CIST</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Yarowsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Grace</given-names>
            <surname>Ngai</surname>
          </string-name>
          , and Richard Wicentowski.
          <year>2001</year>
          .
          <article-title>Inducing multilingual text analysis tools via robust projection across aligned corpora</article-title>
          .
          <source>In HLT.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>D</given-names>
            <surname>Zeman</surname>
          </string-name>
          and
          <string-name>
            <given-names>P</given-names>
            <surname>Resnik</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Cross-language parser adaptation between related languages</article-title>
          .
          <source>In NLP for Less Privileged Languages.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Орест</given-names>
            <surname>Коссак</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Українська комп'ютерна термінологія</article-title>
          .
          <source>In Сучасні проблеми в комп'ютерних науках</source>
          , pages
          <fpage>39</fpage>
          --
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>Р Рожанківський and М Кузан</source>
          .
          <year>2000</year>
          .
          <article-title>Комп'ютерні проблеми стандартизації термінології. In Сучасні проблеми в комп'ютерних науках (CCU'</article-title>
          <year>2000</year>
          ), pages
          <fpage>42</fpage>
          --
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>