<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview and Future of Czech Wordnet</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adam Rambousek</string-name>
          <email>rambousek@fi.muni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karel Pala</string-name>
          <email>pala@fi.muni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Tukacova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Natural Language Processing Centre Faculty of Informatics, Masaryk University Botanick 68a</institution>
          ,
          <addr-line>602 00 Brno</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <abstract>
        <p>Czech Wordnet represents one of the national wordnets created during the EuroWordNet and Balkanet projects. However, the data contains various issues that a ects the use of Czech Wordnet in NLP applications. Due to lack of resources, it was not possible to update Czech Wordnet thoroughly since the publication of the rst version. In 2017, we have started a project to evaluate and update Czech Wordnet, followed by the connection to Collaborative Interlingual Index. This paper provides overview of various updates and extensions of the Czech Wordnet data, and presents the roadmap to publish revised version of Czech Wordnet under open license.</p>
      </abstract>
      <kwd-group>
        <kwd>EuroWordnet</kwd>
        <kwd>Balkanet</kwd>
        <kwd>wordnet</kwd>
        <kwd>Czech Wordnet</kwd>
        <kwd>DEBVisDic</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        After its publication, Princeton WordNet [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] proved its usability as a lexical
resource, both for users and various NLP tasks. Wordnet also inspired many
projects aiming either to create semantic networks in other languages, or extend
wordnet with new features. The rst attempt to build localized wordnets was the
EuroWordNet [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] project in 1996, coordinated by Piek Vossen at the University of
Amsterdam. In the rst phase, EuroWordNet included Dutch, Italian, Spanish,
and English wordnets. In the next phase, German, French, Estonian, and Czech
wordnets were added.
      </p>
      <p>EuroWordNet introduced two new features that were necessary for language
compatibility. With the aim to build semantic networks in several languages
that share the same language core, list of Base Concepts was described. The
list includes 1310 synsets shared amongst all EuroWordNet languages and
represents the part of wordnet that should be encoded rst. Another purpose of
Base Concepts were linguistic studies of language di erences.</p>
      <p>Because the national wordnets re ect various languages with speci c
hierarchy, EuroWordNet project established Interlingual Index (ILI). The index
contained language independent ontology. Each wordnet connects synsets to ILI,
thus enabling multi-lingual links. The features and processes developed during
the EuroWordNet project were later re-used during building of other national
wordnets.</p>
      <p>
        One of such projects was the Balkanet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] project in 2001-2004, aiming to
expand the number of national wordnets for European languages. Balkanet project
covered Bulgarian, Greek, Romanian, Serbian, and Turkish wordnets. Together
with newly developed wordnets, verb synsets in Czech wordnet were extended.
      </p>
      <p>As mentioned, Czech wordnet was created in EuroWordNet and Balkanet
projects by the Natural Language Processing Centre at the Faculty of
Informatics, Masaryk University (NLPC). However, it was published through ELRA
under closed and paid license. Although it is possible to get research license,
Czech Wordnet data are still not available in an open form, and that issue
hampers many attempts to include synset data into any 3rd party tool.</p>
      <p>Since 2004, no thorough work was possible to extend, x, or update Czech
Wordnet data. In 2017, we have started a project to evaluate and update Czech
Wordnet. Followed by connection to Collaborative Interlingual Index and open
publication of the wordnet.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Available versions of Czech Wordnet</title>
      <sec id="sec-2-1">
        <title>Original Czech Wordnet</title>
        <p>
          The original version of the Czech Wordnet [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is available for licensing from
ELRA. This is the version created during EuroWordNet and Balkanet projects,
and contains 28,201 synsets with 43,958 literals. All the synsets are linked to
their counterpart in Princeton Wordnet 2.0. Part of verb synsets (824) were also
enriched with verb frames.
        </p>
        <p>
          Primary method for the wordnet creation was the top-down approach
(proposed in the EuroWordNet project). Lexicographers consulted several resources,
available at the time in electronic form { Czech explanatory dictionary [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
English-Czech dictionary, Czech synonymy dictionary, and the DESAM corpus.
Although the explanatory dictionary contained information about hypernym for
some headwords, this information was not entered systematically. This led to the
solution that most of the hyperonymical relations were directly transferred from
the Princeton Wordnet. Information on Czech synonyms was more extensive,
however not covering all concepts needed. As a result, many synsets were exact
translations of synsets from Princeton Wordnet.
        </p>
        <p>This approach caused various issues with the data. Most notable example are
the synsets containing words that are not exactly synonyms, or only rare in the
Czech language, but present in the Czech Wordnet because of the translation
from English. For example, English synset cabriolet:1, cab:2 has the
equivalent Czech synset kabriolet:2, dvoukolovy jednosprezn povoz:1, konska drozka:1
(cabriolet, two-wheeled one horse cart, horse-drawn carriage). Although the
translation is correct, this sense of kabriolet in Czech is very archaic, in
current language the only sense used in spoken language is the convertible car.
Another problem is the inclusion of multiword expressions in the synset, which
may be justi ed in some cases, these are not xed lexical units in the Czech
language.
2.2
To deal with some of the issues mentioned above, core synsets of the Czech
Wordnet were edited by lexicographers in 2009. In total, 2,400 synsets from
the Base Concept set were edited. Updates included synonyms revision and
de nition editing. Total number of synsets is the same (28,201). This version of
Czech wordnet was not published publicly, but is available for research.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Extended with Bilingual dictionary</title>
        <p>
          To increase coverage of the Czech Wordnet, semi-automatic method was
proposed in 2011 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. We acquired machine-readable data from the largest
onevolume English-Czech dictionary ever published. It contains more than 100,000
headwords and sub-headwords, more than 200,000 words and phrases and roughly
400,000 equivalents. We used the following algorithm to add new words and
synsets:
{ Extract translation pairs from the dictionary.
{ Keep only pairs in which English literals are monosemous.
{ If desired, keep only pairs with unique source literals (one-to-one
translations).
{ Match English literals with monosemous PWN literals.
        </p>
        <p>The extended version of the Czech Wordnet contains 83,769 literals (growth
by 76 %) organized into 40,621 synsets (growth by 43 %). Out of the synsets,
27,658 are noun synsets (increase by 6640, or 31.6 %), 5852 are verb synsets
(increase by 690, or 13.3 %), 5651 are adjective synsets (increase by 3522, or
165.4 %) and 1457 are adverb synsets (increase by 1291, or 877.7 %).</p>
        <p>Because of unsupervised nature of the extension, the newly produced Czech
Wordnet data need to be inspected manually. We have checked a sample of 600
synsets, with the results that 30 % of the synsets contain wrong or unwanted
synonyms, and 20 % of the newly created synsets are connected to an incorrect
hypernym. For this reason, extended Czech Wordnet will not published publicly
before the thorough editing, but is available for research.
2.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Connection to Verbalex</title>
        <p>
          VerbaLex [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is a large lexical database of Czech verb valency frames and has
been under development at NLPC since 2005. The organization of lexical data
in VerbaLex is derived from the wordnet structure and entries follows the form
of synsets. The current version of VerbaLex contains 6,360 synsets, 21,193 verb
senses, 10,482 verb lemmata and 19,556 valency frames. When possible, the
synset from VerbaLex is linked to its equivalent in Princeton Wordnet. Out of
the total number, 3,725 synsets have English equivalent, remaining 2,635 are
verbs speci c for the Czech language.
2.5
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Added de nitions</title>
        <p>Because a lot of synsets in the Czech Wordnet is missing de nitions, students of
the linguistics course at the Faculty of Arts were asked to update the missing
parts. Czech de nitions were written for 5,676 synsets from the Base Concepts
set, consulting both Princeton Wordnet de nitions and Czech explanatory
dictionaries. These revisions are currently only saved in text les and were not
inserted into Czech Wordnet.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>DEBVisDic integration with Open</title>
    </sec>
    <sec id="sec-4">
      <title>Wordnet</title>
    </sec>
    <sec id="sec-5">
      <title>Multilingual</title>
      <p>
        Since the Balkanet project, NLPC is developing browser and editor for
wordnetlike lexical databases { VisDic, later reimplemented as DEBVisDic [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The
editor is storing wordnet data in the XML format, thus making the wordnet-like
databases more standard and exchangeable. Current DEBVisDic version is based
on the DEB platform, general lexicographic platform, based on client-server
architecture and adaptable for wide range of dictionary projects.
      </p>
      <p>
        DEBVisDic is available as a web application and o ers various features for
wordnet browsing and editing. Users may work with several wordnets at once,
utilizing linking and referencing between dictionaries. The application allows any
user to create a new wordnet, without any complicated set-up, and start editing
in a few minutes [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. To promote wordnet sharing, DEBVisDic supports export
to the Wordnet-LMF [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] format.
      </p>
      <p>
        As the part of preparation of new version of Czech Wordnet, DEBVisDic
editor will be updated to integrate better with Open Multilingual Wordnet [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
repository. Users will be able to easily connect synsets to the Collaborative
Interlingual Index [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and upload data to OMW repository directly from the
DEBVisDic.
4
      </p>
    </sec>
    <sec id="sec-6">
      <title>Open Czech Wordnet</title>
      <p>Main impulse to speed up the creation of new version of Czech Wordnet was the
proposal of integrating all available wordnets in the Global WordNet Association
repository with Collaborative Interlingual Index. However, current Czech
Wordnet is not published under open license. Another important motivation is the
need to x various linguistic issues that make it harder to use Czech Wordnet
data in NLP applications.</p>
      <p>We have decided to evaluate and combine all the available updates and
extensions to Czech Wordnet. NLPC team has compiled the following roadmap
that will lead to the publication of Open Czech Wordnet:
{ Start with 2009 Edited version and combine it with de nitions created for</p>
      <p>Base Concepts.
{ Check synonyms present in synsets, remove unnecessary synonyms and add
missing words.
{ Revise or create de nitions where missing. Join or split synsets to follow
word senses used in Czech language, where necessary.
{ Verify all types of relations between synsets semi-automatically and x
broken relations.
{ Link Czech synsets to their equivalents in Princeton Wordnet 3.1 and to
Collaborative Interlingual Index.</p>
      <p>We plan to include extensions from the semi-automatically translated Czech
Wordnet, but the data have to be evaluated by lexicographers rst. Evaluation
is planned at the end of year 2017.</p>
      <p>It was not yet decided, in which way to include VerbaLex data. However, the
best option for the wordnet composition is to create new synsets based on the
VerbaLex entries, including only the synonyms and de nition to the wordnet
data and linking to the VerbaLex for full verb valency information. VerbaLex
does not contain relations between synsets, thus hyperonymy and troponymy
relations have to be set in the wordnet.</p>
      <p>We are building the tool to allow any user of Open Czech Wordnet to
submit suggestions and comments regarding wordnet data. All suggestions will be
reviewed by linguists and if approved, the data will be automatically updated.
We believe this tool will help to improve the quality of Open Czech Wordnet
data.
5</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and Future Work</title>
      <p>Editing work is already under way with the plan to nish the rst phase in
summer 2017. We plan to release the version of Czech Wordnet linked to
Collaborative Interlingual Index in 2018 under open license and then continue with the
evaluation of translated data. Depending on the funding and resources available,
we plan to expand Czech Wordnet and reach the coverage of Princeton Wordnet.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work has been partly supported by the Grant Agency of CR within the
project 15-13277S by the Ministry of Education of CR within the national
COSTCZ project LD15066.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C., ed.:
          <source>WordNet: An Electronic Lexical Database</source>
          . MIT Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Vossen</surname>
          </string-name>
          , P., ed.:
          <article-title>EuroWordNet: a multilingual database with lexical semantic networks for European Languages</article-title>
          . Kluwer (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Christodoulakis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <source>Balkanet Final Report</source>
          , University of Patras, DBLAB (
          <year>2004</year>
          )
          <article-title>No</article-title>
          . IST-2000-29388.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pala</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smrz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Building Czech Wordnet.
          <source>Romanian Journal of Information Science and Technology</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          -2) (
          <year>2004</year>
          )
          <volume>79</volume>
          {
          <fpage>88</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Filipec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Slovnk spisovn etiny (SS). 1st edn</article-title>
          . Academia,
          <string-name>
            <surname>Praha</surname>
          </string-name>
          (
          <year>1995</year>
          )
          <article-title>elektronick verze</article-title>
          , LEDA, Praha.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Blahus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pala</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Extending Czech WordNet using a bilingual dictionary</article-title>
          . In Fellbaum,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Vossen</surname>
          </string-name>
          , P., eds.: 6th International Global Wordnet Conference Proceedings, Matsue, Japan, Toyohashi University of Technology (
          <year>2012</year>
          )
          <volume>50</volume>
          {
          <fpage>55</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hlavackova</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>VerbaLex { New Comprehensive Lexicon of Verb Valencies for Czech</article-title>
          .
          <source>In: Proceedings of the Slovko Conference</source>
          , Bratislava, Slovakia (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rambousek</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruso</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Web Application for Semantic Network Editing</article-title>
          .
          <source>In: RASLAN 2013: Seventh Workshop on Recent Advances in Slavonic Natural Language Processing</source>
          , Brno, Czech Republic,
          <string-name>
            <surname>Tribun</surname>
            <given-names>EU</given-names>
          </string-name>
          (
          <year>2013</year>
          )
          <volume>13</volume>
          {
          <fpage>19</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Rambousek</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>DEBVisDic: Instant Wordnet Building</article-title>
          . In Barbu Mititelu,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Forascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Fellbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Vossen</surname>
          </string-name>
          , P., eds.
          <source>: Proceedings of the Eighth Global WordNet Conference</source>
          , Bucharest, Romania, Romanian
          <string-name>
            <surname>Academy</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <volume>317</volume>
          {
          <fpage>321</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Soria</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monachini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vossen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Wordnet-LMF: eshing out a standardized format for wordnet interoperability</article-title>
          .
          <source>In: Proceedings of IWIC2009</source>
          , New York, ACM Press (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bond</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
          </string-name>
          , R.:
          <article-title>Linking and extending an open multilingual wordnet</article-title>
          .
          <source>In: ACL (1)</source>
          , The Association for Computer Linguistics (
          <year>2013</year>
          )
          <volume>1352</volume>
          {
          <fpage>1362</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bond</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vossen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>CILI: the Collaborative Interlingual Index</article-title>
          . In Barbu Mititelu,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Forascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Fellbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Vossen</surname>
          </string-name>
          , P., eds.
          <source>: Proceedings of the Eighth Global WordNet Conference</source>
          , Romanian
          <string-name>
            <surname>Academy</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <volume>50</volume>
          {
          <fpage>57</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>