<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extending Czech thesauri using word-formation network</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Karolína Horˇenˇovská</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University, Faculty of Mathematics and Physics</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we attempt to extend existing Czech thesauri by using a word-formation network, DeriNet. Thesauri are an important resource for synonym retrieval / substitution generation but their lexical sparsity is an issue in Czech. We discuss the properties of existing thesauri and DeriNet and propose several ways of using DeriNet to extend the thesauri, such as deriving a synonym of an adverb from a synonym of corresponding adjective. We also evaluate some of our proposals.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        A lot of effort has been invested in creating large
thesauri, of which the best known example is probably
WordNet [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], followed by others such as FrameNet [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While
these thesauri address English, there are many thesauri for
other languages as well (there are e.g. WordNet versions
for Arabic [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Swedish [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], or Czech [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). We wish to
emphasize Czech WordNet since Czech is the language we
currently deal with.
      </p>
      <p>
        However, those thesauri are heavily incomplete for
some languages, including the above-mentioned Czech
language. This incompletness presents a problem for
various NLP tasks, e.g. substitution generation as part of
lexical simplification (see [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] or [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] for more detail).
      </p>
      <p>On the other hand, for some languages (including
Czech), a rich word-formation network is available. We
propose using such network to extend existing thesauri,
i.e. to discover synonymy relations between new pairs of
words. Please note that while we target synonymy, as it
is the only relation covered by all existing Czech thesauri,
the approach would hold for any relation.</p>
      <p>The rest of the paper is organized as follows: we briefly
describe existing related work (section 2), present existing
Czech thesauri (section 3) and describe the Czech
wordformation network DeriNet (section 4). We then
introduce several ways of combining DeriNet with thesauri to
produce new relations (section 5) and evaluate the most
promising of them (section 6).</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>Since thesauri are generally incomplete, there have been
lots of attempts at extending them in an automated way.</p>
      <p>
        These attempts have included aligning multilingual
resources (e.g. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]), mining the Wikipedia ([
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]) or the
web in general ([
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]), translating English WordNet (which
has been tried especially for Czech [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], even though the
extension itself, to the best of our knowledge, is not publicly
available), making use of word embeddings ([
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) as
well as by employing derivational morphology ([
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>
        This paper is in its nature similar to a previous attempt
of extending Czech WordNet with derivational relations
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] which the authors claimed was successful. Unlike
them, we use a publicly available source of derivations and
we do not limit ourselves to WordNet – we try using
various thesauri and compare the outcome obtained with each
of them and with their combination. We also share a more
thorough evaluation of the resulting pairs.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Existing thesauri</title>
      <p>
        We are aware of five notable Czech thesauri:
the most recent version of Czech WordNet [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], and
a slightly divergent version of Czech WordNet [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
which lacks some synsets but contains some
others which were created to enable the lexico-semantic
annotation of Prague Dependency Treebank ([
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]),
which we refer to as WordNet (PDT);
thesaurus formerly distributed as a part of office
software LibreOffice,
      </p>
      <sec id="sec-3-1">
        <title>Czech Wiktionary, and</title>
        <p>ÚFAL thesaurus, a thesaurus developed at our
department.</p>
        <p>Both WordNet versions explicitly utilize synsets, each
synset represents a meaning and lists literals (words or
phrases) which can be used to express the meaning.
Synsets might include a definition of the meaning but few
have it filled.</p>
        <p>The last three thesauri employ synsets implicitly, either
by assigning a word with a set of sets of synonyms (as
done in LibreOffice thesaurus and Wiktionary) or by
listing sets of synonyms and including some words in more
such sets (as done in our department thesaurus).</p>
        <p>We perform our experiments both using each of the
thesauri individually and using a concatenated thesaurus, i.e.
an artificial thesaurus created by concatenating all synsets
from each of the real thesaurus.</p>
        <p>In our work, we do not make use of synsets. For each
word, we merge its synonyms from all synsets and
produce a set of its synonyms (despite the context, i.e. words
which share the meaning at least in some contexts). This is
partially to simplify the proof of concept, partially because
senses in both WordNets are much more fine-grained than
senses in other resources.</p>
        <p>However, this step is in no way crucial. One could keep
the synsets, and whenever we refer to retrieving synonym,
they could first retrieve the synsets and only then retrieve
the words (either from specific or all synsets). We actually
expect to do this in our future work.</p>
        <p>Some further statistics about the thesauri are provided in
table 1. The concatenated line corresponds to
concatenating all synsets. Please note that we only work with
singleword expressions (as opposed to multi-word expressions).
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>DeriNet</title>
      <p>
        DeriNet [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] is a Czech word-formation network. Its
nodes are Czech lexemes, i.e. lemmata, and the nodes do
not have to cover all sensesl. The authors report to have
decided to take a rather minimalistic approach to polysemy,
and only represent a lemma with more nodes if at least
one of two conditions is met: it was coincidentally derived
from two different words (could be demonstrated by verb
proudit, which is represented as a base word, though it is
likely related to noun proud ’flow’, and also as a verb
derived from udit ’to smoke’, when proudit refers to smoking
something thoroughly), or the senses lead to different sets
of derived words (i.e. verb stát ’to stand, to melt away’).
      </p>
      <p>
        The directed edges then represent the fact that one word
is derived from the other one. The edges should be taken as
implicative, some derivations might not be captured in
DeriNet (yet). They are discovered using a variety of
methods, including manual deduction, rule-based automated
processes and machine learning; many of them were also
taken from the MorfFlex CZ morphological dictionary [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
All discovered edges are manually confirmed before being
added to the network.
      </p>
      <p>By the authors’ design decision, no word is allowed to
have more than one parent, which simplifies the structure
and could be justified by low occurence of compounds in
Czech. Even though only one parent is allowed, recent
versions of DeriNet allow for an indication of being a
compound in part of speech specification.</p>
      <p>
        The then current version of the network (1.7) contains
1; 027; 655 nodes, though only some of the nodes are
supported by corpus evidence (when compared to SYN v4
version of Czech National Corpus [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we found out that as
many as 591; 486 nodes (i.e. more than a half) do not
occur in the corpus). For the first version, only words which
occured at least twice in a SYN subcorpus of Czech
National Corpus (and fullfiled a few other conditions) were
inserted in the network; this condition does not hold for
lemmata inserted from MorfFlex CZ dictionary.
      </p>
      <p>Of all nodes, 104; 563 (approx. 10 %) are isolated, i.e.
they are not connected with any other node.</p>
      <p>Except for the parent and part of speech, there is no
further annotation, i.e. one cannot learn for example that the
derived noun is agent noun of the base verb. DeriNet
format is therefore farily simple: it gives node ID, its lemma
and technical lemma (which contains some additional
details such as sense disambiguation), its part of speech
(perhaps with the above-mentioned indication of being a
compound) and its parent’s ID (if the node has a parent).
5</p>
    </sec>
    <sec id="sec-5">
      <title>Proposed thesauri extensions</title>
      <p>We propose the following principle of discovering new
word relations:
1. Find a non-root node A (i.e. a node which has a
parent).
2. Get A’s parent, B.
3. Retrieve B’s synonyms using the existing thesauri.
4. Find all nodes C which correspond to the retrieved
synonyms.
5. For each C, check if it has a child D which shares
requested features with A.</p>
      <sec id="sec-5-1">
        <title>6. Declare A and D a related word pair.</title>
        <p>This outline does not specify how to deal with the
situation when more than one D exist (share given features
with A) for single C. In our experiments, we opted for
choosing neither (i.e. skipping the whole C subtree) but
one could also develop strategies to select the best D or
generate more pairs for A from single C.</p>
        <p>It should be noted that due to this decision, discovering
synonymous word pairs is not symmetric, that is, a word
pair might be discovered when starting from one word, but
not when starting with the other one.</p>
        <p>
          We actually suggest further constraining all of A, B, C
and D to improve the reliability of the discovered relations,
i.e. by constraining their part of speech. While part of
speech is the only feature available in DeriNet itself, we
can use e.g. MorfFlex CZ dictionary or MorphoDiTa tool
for morphological analysis [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] to enable more features.
        </p>
        <p>One could be tempted to only search for those
nonroot nodes A which are not covered by any thesauri, the
reasoning being that such nodes already have their
synonyms in the thesauri. However, thesauri entries for
individual words are often incomplete and the outlined
process could still find new synonyms for node A, even if
node A is present in a thesaurus. Furthemore, considering
only nodes A which are not covered by thesauri could lead
to a decrease in number of retrieved pairs after adding a
new thesaurus as some nodes could be newly skipped. We
therefore do not constrain node A on its presence/absence
in the thesauri.</p>
        <p>While we describe the process as deriving synonyms by
using the synonymy relation of parent nodes, the
parentchild relation is not crucial. One could reword the
process e.g. with finding A’s child B and C’s parent D, or
with finding A’s grand-parent B (and C’s grand-child D).
However, by the nature of thesauri creation, we expect
them to contain the base word rather than the derived word
(though the direction of the derivation is sometimes
ambigous). Longer distance relations, on the other hand, are
more likely to introduce noise.</p>
        <p>Having introduced the basic principle, we suggest
specific approaches to relation derivation. We assume the
access to richer morphological annotation, e.g. by using the
MorphoDiTa tool.
5.1</p>
        <sec id="sec-5-1-1">
          <title>Deriving adverb synonyms using adjectives</title>
          <p>Adverbs are often derived from adjectives and while
adverb ratio in thesauri is close to corpus ratio (approx.
2:6 %-3:2 % of content words as measured using Czech
National Corpus syn v4), we often ran into issues with
them in our text simplification experiments.</p>
          <p>We suggest constraining nodes A and D to be adverbs
and nodes B and C to be adjectives.
5.2</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Deriving feminine forms from masculine</title>
          <p>In Czech, some words come in different forms for men
(or males generally) and women. This in particular holds
for roles in relationships and for agent nouns, e.g. there is
ucˇitel ’teacher (man)’ and ucˇitelka ’teacher (woman)’.</p>
          <p>Thesauri usually only cover the masculine variants, both
because they are usually the default and because native
speakers can infer the feminine variant (still, some
language knowledge is required, e.g. there is ucˇitelka to ucˇitel
but ministryneˇ to ministr ’minister’, not *ministrka).</p>
          <p>We suggest constraining nodes A and D to be nouns
having feminine gender and nodes B and C to be nouns
having masculine gender.
5.3</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>Deriving possesive adjectives from nouns</title>
          <p>Similarly to omitting feminine forms, thesauri generally
do not cover possesive adjectives since they can be easilly
inferred from the corresponding noun.</p>
          <p>We suggest constraining nodes A and D to be possesive
adjectives and nodes B and C to be nouns.
5.4</p>
        </sec>
        <sec id="sec-5-1-4">
          <title>Deriving verbs using verbs of opposite aspect</title>
          <p>Thesauri differ in treating verb aspects, and often thesauri
are not consistent even internally. Sometimes the verb of
opposite aspect is listed as synonym, sometimes both
aspects form their own synset, sometimes the other aspect is
completely missing.</p>
          <p>We suggest constraining nodes A and D to have
matching aspect, nodes B and C to have matching aspect, and
nodes A and B to have opposite aspect.</p>
          <p>
            The issue with this suggestion is that aspect
information is neither present within MorfFlex CZ dictionary nor
provided by MorpohDiTa. We believe, however, that
annotation from Czech National Corpus (e.g. the
beforementioned version syn 4) [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ], which is enriched with
aspect annotation, could be used.
6
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Evaluation</title>
      <p>We tried generating synonym pairs from adverbs using
adjectives, feminine forms using masculine forms and
possesive adjectives using nouns (see section 5 for more
detail).</p>
      <p>The generation procedure was carried out for each of
the thesauri individually and also for the result of
concatenating all the thesauri together.</p>
      <p>Our results are reported in table 2. We report the
number of obtained pairs for each of the thesauri as well as
for the concatenated thesaurus. The reported numbers are
after symmetrization, i.e. after expanding any pair A-D
into both A-D and D-A. The actual numbers of discovered
pairs are usually 1.6-1.9 times greater as most pairs (but
not all of them) are discovered in both directions. These
ratios seem to slightly correlate with the selected strategy
(symmetrization is of greatest help when finding feminine
variants).</p>
      <p>When evaluating a specific thesaurus, we can discover
a synonym pair which is actually present in some other
thesaurus. Whenever this happens, we consider such pair</p>
      <sec id="sec-6-1">
        <title>Concatenated</title>
        <p>a confirmed one. We do not evaluate it further and expect
that the pair is correctly derived and synonymous.</p>
        <p>For each strategy, we sampled 100 non-confirmed pairs
and asked 4 annotators to annotate them as either
synonym, antonym or unrelated. The annotators were of
varying gender and age, though all of them have obtained a
university degree during their life. Asking annotators to
distinguish antonyms from unrelated pairs was done based
on our informal result analysis, which revealed antonym
pairs do occur.</p>
        <p>In 5 cases, the annotators admitted they did not know,
in a few other, they noted they were not really sure. In
all cases, a very infrequent word was involved. We treat
I don’t know as unrelated when reporting precision and
inter-annotator agreement. We treat answers marked with
not sure in the same way as unmarked.</p>
        <p>The inter-annotator agreement (Fleiss’ kappa) was 0.47
and it slightly varied over the strategies (0.47, 0.42 and
0.50, respectively). These numbers might seem low but it
is important to keep in mind that most answers were
synonym, hence this answer had a great probability, and
therefore any disagreement on other answers had a big impact.</p>
        <p>Examples of correctly discovered pairs (pairs annotated
as synonyms by all four annotators) are given in table 3.
There were 9 pairs marked as antonyms by all
annotators (out of 28 marked as antonyms by at least one
annotator), they are all listed in table 4. In all 9 cases, the
pair was derived using a synonymy relation from
LibreOffice thesaurus, when either two antonyms were suggested
as synonyms to the same word, e.g. both tlouštík ’fatty’
and hubenˇour ’thin man’ to tlust’och’ ’fatty’, or when an
antonym was suggested directly as e.g. inkompatibilní
’incompatible’ to kompatibilní ’compatible’. While these
pairs are not synonymous, their existence should not be
used to decline the principle. On the contrary, should
the thesaurus pairs be correctly marked as antonyms,
we would correctly derive antonymous pairs using our
method.</p>
        <p>Finally, there were 14 pairs annotated as unrelated by
vodpoveˇdneˇ</p>
        <p>hanebneˇ
vyzveˇdacˇka
cˇarodeˇjnice
surovcu˚v
maršálku˚v</p>
        <p>spolehliveˇ
bezcharakterneˇ
špehounka
divotvorkyneˇ
krut’asu˚v
maršálu˚v</p>
        <p>dependably
unscrupulously
she-spy
witch
bully’s
marshal’s
all annotators. They are listed in table 5. In some cases
(1, 4, 5), there is some evidence that the words can share
a meaning but at least one of the words is associated with
another meaning so strongly that the annotators probably
did not realize the meaning could be the same.</p>
        <p>Some pairs (2, 3, 10, 11, 13) come from a synset in
thesauri, even though we could not find any other evidence
that these pairs really could share the meaning.</p>
        <p>Other pairs (6, 14) occur because of insufficiencies in
the derivational process. While the base words are
synonyms, the derived words are of distinct genders. These
pairs could be prevented by constraining the suggested
pairs more carefully.</p>
        <p>Case 9 is quite similar. Both words stru˚jce ’creator’ and
otec ’father’ could refer to a creator (author) and stru˚jkyneˇ
’she-creator’ is a feminine variant of stru˚jce. However,
while the word otcˇina ’fatherland’ is directly derived from
otec and is feminine, it does not in any way refer to
shefather. This could be prevented with more detailed
annotations in the word-formation network.</p>
        <p>There are cases (7, 8) when, despite the principle
proving good, derived words are not really perceived
synonymous, even though the base words could be. For example,
both words ku˚nˇ ’horse’ and osel ’donkey’ could be used to
refer to a dumb person but their feminine variants are not
used in that way (even though in theory they could be).</p>
        <p>The last case, 12, is special in many ways. The word
zárovenˇ ’at the same time, simultaneously’ is reported to
be derived from word rovný ’straight’, which might seem
surprising. The pair is further derived from thesauri pair
pokrˇiveneˇ
povšechneˇ
ru˚zneˇ
mladice
tlust’oška
živelneˇ
jemneˇ
inkompatibilneˇ
bezcitneˇ
distorted-ly
in general
diversely, differently
young woman</p>
        <p>fat woman
elementally, unrestrainedly
softly, lightly
incompatibly
heartlessly</p>
        <p>rovno
konkrétneˇ
identicky
starˇice
hubenˇourka
organizovaneˇ</p>
        <p>pikantneˇ
kompatibilneˇ
vrˇele</p>
        <p>straight
specifically
identically
old woman
thin woman
organized-ly
spicy, zesty
compatibly
heartily
rovný ’straight’ – zakroucený ’tortuous’, which is rather
antonymous.</p>
        <p>We do not provide detailed report on pairs annotated
differently by different annotators, though we have examined
them too. In most cases, some evidence of shared meaning
exist but some of the annotators did not consider the words
synonymous.</p>
        <p>Following from the above analysis, more than half of
unrelated pairs is not less related that their base word
counterparts. These pairs do not contradict our method, they
only evidence the necessity of both checking thesauri
quality and being careful about the synonymy itself as it is
perceived differently by different people.</p>
        <p>There are cases when our method fails to filter out
nonsynonymous derived pairs. This could be improved both
by better filtering during the inference process and by
having better annotation in the word-formation network.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>We have presented a method of deriving new synonym
pairs using existing thesauri and word-formation network,
we have suggested several strategies to do the actual
derivation and we have evaluated some of them.</p>
      <p>Our evaluation revealed that about half of derived
synonym pairs are really perceived synonymous by all of our
human annotators and around 80 % are perceived
synonymous by at least two of them. The erroneous word pairs
are caused by two distinct factors. First, there are errors
in the thesauri synsets: unrelated, or even antonymous,
words are occasionally marked as synonymous. Second,
there are limitations of our method, where the derived
words are not synonymous, despite being derived from
synonymous base words.</p>
      <p>Some of the limitations could be overcome by better
filtering within our method or by more detailed
annotations in the word-formation network. The latter has
become available soon after we carried out our experiments,
as DeriNet 2.0 has been released. This version has a more
detailed annotation of both nodes (e.g. noun gender) and
edges (the purpose of derivation is annotated, e.g.
diminutivization), and we expect this version to be helpful in
future experiments.</p>
      <p>
        We also plan to try using Derivancze [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] (which also
includes derivation annotations) instead of DeriNet as the
word-formation network and see if it helps to improve our
results.
      </p>
      <p>Overall, we consider our results good because they
suggest that thesauri authors can focus on capturing the
relations between the base words and NLP applications can
still make good use of those thesauri even for derived
words.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgement</title>
      <p>This work has been supported by the grant No. 1704218
of the Grant Agency of Charles University. It has been
using language resources and tools stored and distributed
by the LINDAT/CLARIN project of the Ministry of
Education, Youth and Sports of the Czech Republic (project
LM2015071). The research was also partially supported
by SVV project number 260 453.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Musa</given-names>
            <surname>Alkhalifa</surname>
          </string-name>
          and
          <string-name>
            <given-names>Horacio</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          .
          <article-title>Automatically extending NE coverage of Arabic WordNet using Wikipedia</article-title>
          .
          <source>In Proc. Of the 3rd International Conference on Arabic Language Processing CITALA2009</source>
          , Rabat, Morocco,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Collin</surname>
            <given-names>F Baker</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Charles J Fillmore</surname>
            ,
            <given-names>and John B Lowe.</given-names>
          </string-name>
          <article-title>The Berkeley FrameNet project</article-title>
          .
          <source>In Proceedings of the 17th international conference on Computational linguisticsVolume 1</source>
          , pages
          <fpage>86</fpage>
          -
          <lpage>90</lpage>
          . Association for Computational Linguistics,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Eduard</surname>
            <given-names>Bejcˇek</given-names>
          </string-name>
          , Petra Hoffmannová, Martin Holub,
          <string-name>
            <surname>Marie</surname>
            <given-names>Hucˇínová</given-names>
          </string-name>
          , Pavel Pecina,
          <string-name>
            <surname>Pavel</surname>
            <given-names>Stranˇák</given-names>
          </string-name>
          , Pavel Šidák, and Jan Hajicˇ.
          <article-title>Lexico-semantic annotation of PDT using Czech WordNet,</article-title>
          <year>2011</year>
          .
          <article-title>LINDAT/CLARIN digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>William</given-names>
            <surname>Black</surname>
          </string-name>
          , Sabri Elkateb, and
          <string-name>
            <given-names>Piek</given-names>
            <surname>Vossen</surname>
          </string-name>
          .
          <article-title>Introducing the Arabic WordNet project</article-title>
          .
          <source>In In Proceedings of the third International WordNet Conference (GWC-06. Citeseer</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Marek</given-names>
            <surname>Blahuš</surname>
          </string-name>
          .
          <article-title>Extending Czech WordNet using a bilingual dictionary</article-title>
          .
          <source>Master's thesis</source>
          , Faculty of Informatics, Masaryk University,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jan</given-names>
            <surname>Hajicˇ and Jaroslava Hlavácˇová. MorfFlex</surname>
          </string-name>
          <string-name>
            <surname>CZ</surname>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>LINDAT/CLARIN digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Jugal</given-names>
            <surname>Kalita</surname>
          </string-name>
          et al.
          <article-title>Enhancing automatic WordNet construction using word embeddings</article-title>
          .
          <source>In Proceedings of the Workshop on Multilingual and Cross-lingual Methods in NLP</source>
          , pages
          <fpage>30</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Svetla</given-names>
            <surname>Koeva</surname>
          </string-name>
          , Cvetana Krstev, and
          <string-name>
            <given-names>Duško</given-names>
            <surname>Vitas</surname>
          </string-name>
          .
          <article-title>Morphosemantic relations in WordNet-a case study for two Slavic languages</article-title>
          .
          <source>In Global wordnet conference</source>
          , pages
          <fpage>239</fpage>
          -
          <lpage>253</lpage>
          . University of Szeged, Department of Informatics,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Michal</surname>
            <given-names>Krˇen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Václav</surname>
            <given-names>Cvrcˇek</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomáš</surname>
            <given-names>Cˇapka</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anna</surname>
            <given-names>Cˇermáková</given-names>
          </string-name>
          , Milena Hnátková, Lucie Chlumská, Tomáš Jelínek,
          <string-name>
            <surname>Dominika</surname>
            <given-names>Kovárˇíková</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vladimír</surname>
            <given-names>Petkevicˇ</given-names>
          </string-name>
          , Pavel Procházka, Hana Skoumalová, Michal Škrabal,
          <string-name>
            <surname>Petr</surname>
            <given-names>Trunecˇek</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavel</surname>
            <given-names>Vondrˇicˇka</given-names>
          </string-name>
          , and Adrian Zasina.
          <source>SYN v4: large corpus of written Czech</source>
          ,
          <year>2016</year>
          .
          <article-title>LINDAT/CLARIN digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Robert</surname>
            <given-names>Meusel</given-names>
          </string-name>
          , Mathias Niepert, Kai Eckert, and
          <string-name>
            <given-names>Heiner</given-names>
            <surname>Stuckenschmidt</surname>
          </string-name>
          .
          <article-title>Thesaurus extension using web search engines</article-title>
          .
          <source>In International Conference on Asian Digital Libraries</source>
          , pages
          <fpage>198</fpage>
          -
          <lpage>207</lpage>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>George</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Wordnet: a lexical database for English</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Verginica</surname>
            <given-names>Barbu</given-names>
          </string-name>
          <string-name>
            <surname>Mititelu</surname>
          </string-name>
          .
          <article-title>Adding morpho-semantic relations to the Romanian WordNet</article-title>
          . In LREC, pages
          <fpage>2596</fpage>
          -
          <lpage>2601</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Karel</surname>
            <given-names>Pala</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomáš</surname>
            <given-names>Cˇapek</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barbora</surname>
            <given-names>Zajícˇková</given-names>
          </string-name>
          , Dita Bartu˚šková, Katerˇina Kulková, Petra Hoffmannová,
          <string-name>
            <surname>Eduard</surname>
            <given-names>Bejcˇek</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavel</surname>
            <given-names>Stranˇák</given-names>
          </string-name>
          , and Jan Hajicˇ.
          <source>Czech WordNet 1.9 PDT</source>
          ,
          <year>2011</year>
          .
          <article-title>LINDAT/CLARIN digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Karel</given-names>
            <surname>Pala</surname>
          </string-name>
          and
          <article-title>Dana Hlavácˇková</article-title>
          .
          <article-title>Derivational relations in Czech WordNet</article-title>
          .
          <source>In Proceedings of the workshop on baltoslavonic natural language processing: Information extraction and enabling technologies</source>
          , pages
          <fpage>75</fpage>
          -
          <lpage>81</lpage>
          . Association for Computational Linguistics,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Karel</given-names>
            <surname>Pala</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Šmerk</surname>
          </string-name>
          .
          <article-title>Derivancze-derivational analyzer of Czech</article-title>
          . In International conference on text, speech, and dialogue, pages
          <fpage>515</fpage>
          -
          <lpage>523</lpage>
          . Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Karel</given-names>
            <surname>Pala</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Smrž</surname>
          </string-name>
          .
          <article-title>Building Czech WordNet</article-title>
          .
          <source>Romanian Journal of Information Science and Technology</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          -2):
          <fpage>79</fpage>
          -
          <lpage>88</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Heidi</surname>
            <given-names>Sand</given-names>
          </string-name>
          , Erik Velldal, and Lilja Øvrelid.
          <article-title>WordNet extension via word embeddings: Experiments on the Norwegian WordNet</article-title>
          .
          <source>In Proceedings of the 21st Nordic Conference on Computational Linguistics</source>
          , pages
          <fpage>298</fpage>
          -
          <lpage>302</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Magda</given-names>
            <surname>Ševcˇíková and Zdeneˇk Žabokrtský</surname>
          </string-name>
          .
          <article-title>Wordformation network for Czech</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC-2014)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Shardlow</surname>
          </string-name>
          .
          <article-title>A survey of automated text simplification</article-title>
          .
          <source>International Journal of Advanced Computer Science and Applications</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <fpage>58</fpage>
          -
          <lpage>70</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Lucia</surname>
            <given-names>Specia</given-names>
          </string-name>
          , Sujay Kumar Jauhar, and
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          .
          <article-title>Semeval-2012 task 1: English lexical simplification</article-title>
          .
          <source>In Proceedings of the First Joint Conference on Lexical and Computational Semantics-Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>347</fpage>
          -
          <lpage>355</lpage>
          . Association for Computational Linguistics,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Jana</surname>
            <given-names>Straková</given-names>
          </string-name>
          , Milan Straka, and Jan Hajicˇ.
          <article-title>Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition</article-title>
          .
          <source>In Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          , pages
          <fpage>13</fpage>
          -
          <lpage>18</lpage>
          , Baltimore, Maryland,
          <year>June 2014</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Lonneke</surname>
            <given-names>Van der Plas and Jörg</given-names>
          </string-name>
          <string-name>
            <surname>Tiedemann</surname>
          </string-name>
          .
          <article-title>Finding synonyms using automatic word alignment and measures of distributional similarity</article-title>
          .
          <source>In Proceedings of the COLING/ACL on Main conference poster sessions</source>
          , pages
          <fpage>866</fpage>
          -
          <lpage>873</lpage>
          . Association for Computational Linguistics,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Ake</surname>
            <given-names>Viberg</given-names>
          </string-name>
          , Kerstin Lindmark, Ann Lindvall, and
          <string-name>
            <given-names>Ingmarie</given-names>
            <surname>Mellenius</surname>
          </string-name>
          .
          <article-title>The Swedish WordNet project</article-title>
          .
          <source>In Proceedings of the Tenth EURALEX International Congress</source>
          ,
          <string-name>
            <surname>EURALEX</surname>
          </string-name>
          <year>2002</year>
          : Copenhagen, Denmark,
          <source>August 13-17</source>
          ,
          <year>2002</year>
          , pages
          <fpage>407</fpage>
          -
          <lpage>412</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Ichiro</surname>
            <given-names>Yamada</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jong-Hoon</surname>
            <given-names>Oh</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Chikara</given-names>
            <surname>Hashimoto</surname>
          </string-name>
          , Kentaro Torisawa,
          <string-name>
            <surname>Jun'ichi Kazama</surname>
            , Stijn De Saeger, and
            <given-names>Takuya</given-names>
          </string-name>
          <string-name>
            <surname>Kawada</surname>
          </string-name>
          .
          <article-title>Extending wordnet with hypernyms and siblings acquired from Wikipedia</article-title>
          .
          <source>In Proceedings of 5th International Joint Conference on Natural Language Processing</source>
          , pages
          <fpage>874</fpage>
          -
          <lpage>882</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Zdeneˇk</given-names>
            <surname>Žabokrtský</surname>
          </string-name>
          , Magda Ševcˇíková, Milan Straka, Jonáš Vidra, and
          <string-name>
            <given-names>Adéla</given-names>
            <surname>Limburská</surname>
          </string-name>
          .
          <article-title>Merging data resources for inflectional and derivational morphology in Czech</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), pages
          <fpage>1307</fpage>
          -
          <lpage>1314</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>