<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Representing Arabic Lexicons in Lemon - a Preliminary Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>John P. McCrae</string-name>
          <email>john.mccrae@insight-centre.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hamzeh Amayreh Birzeit University</institution>
          ,
          <addr-line>Palestine</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mustafa Jarrar Birzeit University</institution>
          ,
          <addr-line>Palestine</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National University of Ireland Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Birzeit University; licensed under Creative Commons License CC-BY LDK 2019 - Posters Track. Editors: Thierry Declerck and John P. McCrae</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present our progress in representing 150 Arabic multilingual lexicons using Lemon, which we have been digitizing from scratch. These lexicons are available through a lexicographic search engine (https://ontology.birzeit.edu) that allows searching for translations, synonyms, and definitions. Representing these lexicons in Lemon will enable them to be used by ontologies and NLP applications, as well as to be interlinked with the Open Linguistic Data Cloud. 2012 ACM Subject Classification Information systems → Resource Description Framework (RDF); Computing methodologies → Language resources New trends in lexical semantics are demanding lexicons not only to be digitized and well-structured but to also be published and interlinked with other resources. This was realized by the Linguistic Linked Open Data paradigm [13], which is a large collaborative community project to interlink the lexical entries of many different linguistic data sources. The W3C's Lemon RDF model [2], developed by the OntoLex Community Group, aims to enable lexicons to be used by ontologies and NLP applications [4]. Lemon can be used to describe lexical entries and their syntactic and semantic information, encouraging not only the reuse of existing lexicographic data inside modern IT applications, but also the interlinking with other lexicographic resources. Unlike many languages, there is only a limited number of structured Arabic lexicons available in digital format. Earlier attempts to digitize and represent Arabic lexicons using the ISO LMF standard [3] can be found in [14] for Arabic morphological data, [12] for Dutch-Arabic linguistic data, [11] for Al-Madar lexicon, and [15] for classical Hadith lexicons. A recent attempt to digitize Al-Qamus Almuhit lexicon and represent it using the W3C's Lemon can be found in [10]. However, none of these attempts provided access to their lexicons or interlinked it with other resources. In this paper, we report on our progress on representing 150 Arabic mono/multilingual lexicons using the W3C's Lemon model. These lexicons were digitized over 9 years, during which we had to obtain copyright permissions, digitize most of them by hand, then clean, restructure, normalize and store them in a database - forming the largest Arabic lexicographic database (see [7] and [1]). The database currently contains about 1.1M lexical concepts, 2.4M multilingual lexical entries, 1.5M translation pairs, 0.7M glosses, and 0.5M semantic relations. The database also contains the Arabic Ontology, which is a formal Arabic wordnet that we have built on the basis of a carefully designed ontology [6][5]. It consists currently of about 1.3K concepts, in addition to 11K concepts being validated. The Arabic ontology, which is mapped to WordNet, BFO, and DOLCE, is currently being used</p>
      </abstract>
      <kwd-group>
        <kwd>and phrases Arabic</kwd>
        <kwd>Lexicographic Database</kwd>
        <kwd>Lexicographic search engine</kwd>
        <kwd>Lemon Model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Work</title>
      <p>to reference lexical concepts in all lexicons, as will be explained later. A lexicographic search engine
[7] was built atop this database (see Figure 1), allowing people to search for translations, synonyms,
definitions, morphology, and other information. The results are retrieved from the ontology and the
150 lexicons. As will be explained later, an RDF icon is shown beside each retrieved result (i.e., a
lexical concept), which allows accessing the Lemon representation of this concept. The search
engine also allows applications to query the database directly, through a set of RESTfull web services,
and retrieve the results in JSON format.</p>
      <p>ontology.birzeit.edu
&lt; &gt;
country</p>
      <p>Translations</p>
      <p>Synonyms</p>
      <p>Definitions
Ontology Dictionaries</p>
      <p>Morphology</p>
      <p>Share 146 results (0.05 secs)</p>
      <p>country دﻠﺑ | ﺔﻟود
لﻛّﺷُﯾو بﻌْﺷَ ﮫﻟ ﺎﮭﯾﻠﻋ قﻔﺗﻣُﻟا ﺔﯾﺳﺎﯾﺳﻟا هدوُدﺣُﺑ ف رَﻌُﯾ يرﺎﺑﺗﻋا دوﺟوﻣ
.ﺔﻣَظﱠَﻧﻣُ تﺎﺳَﺳﱠؤﻣُو ﺔﻣَوﻛُﺣُ تاذ لﻘﺗﺳﻣ ﺔﻣوظﻧﻣ</p>
      <p>BZU Thesaurus ©</p>
      <p>ARABIC ONTOLOGY
country – 1 results (0.04 secs)</p>
      <p>ع| En</p>
      <p>About
country | land نطَ وَ | دﻠََﺑ | ﺔﻟَ وَْد
A geopolitical area with fiat borders, occupied by
nation(s) governed by a state.
تﺎﺳَﺳﱠؤﻣُو ﺔﻣَوُﻛﺣُ ﺎﮭﯾﻓو بﻌْﺷَو ﺔﻓورﻌﻣ دوُدﺣُ ﺎﮭﻟ ﺔﯾﺳﺎﯾﺳوﯾﺟ ﺔﻘطﻧﻣ</p>
      <p>ﺎﮭﺑ فرﻌﺗ ﺎﯾﻟود ﺔﻓورﻌﻣ ﺔﯾرﺎﺑﺗﻋا ﺔﯾﺻﺧﺷ ﺎﮭﻟو ،ﺔﻣَظﱠﻧَﻣُ
293121 TypeOf: {geopolitical area} Instances
state | land | country ﺔَﻟ وَْد | دَﻠﺑَ | ض رْأ
the territory occupied by a nation; ''he returned to the land of
"his birth''; ''he visited several European countries</p>
      <p>Arabic WordNet ©
nation | land | country ﺔﻣأ
the people who live in a nation or country; ''a statement that
sums up the nation's mood’’</p>
      <p>Arabic WordNet ©
developing country ﺔﯾﻣﺎﻧﻟا ﺔﻟودﻟا
ﺔﯾﻣﻧﺗﻠﻟ ﺞﻣارﺑ قﯾرط نﻣ ،يدﺎﺻﺗﻗﻻا وﻣﻧﻟا ﻰﻟإ ﻊﻠطﺗﺗ ﻲﺗﻟا ،ﺔﻔﻠﺧﺗﻣﻟا ﺔﻟودﻠﻟ رﺧآ مﺳا
.. دﯾزﻣﻠﻟ .لﺟﻷا ﺔﻠﯾوط ﺔﯾدﺎﺻﺗﻗﻻا</p>
      <p>Economics Glossary ©
country | country side فﯾر</p>
      <p>The Unified Dictionary of Tourism Terms ©
Representing Arabic multilingual lexicons using Lemon is non-trivial, because of some specificities
of Arabic, and because there are different types of lexicons with different structures. Before
discussing these challenges, it is important to understand the types of lexicons according to their internal
structure and content type, which we relatively classify as: (i) Dictionary: a list of terms, each with
some bi/trilingual translations. (ii) Thesaurus: sets of synonymous lexical entries, in one or multiple
languages, and might contain relations between these sets. (iii) Glossary: a set of entries each with
a domain-specific short gloss. Advanced glossaries may also provide synonyms, translation(s), and
references to other entries, e.g. equivalent, or related. (iv) Linguistic Lexicon: a set of headwords,
each with its sense(s) and features (e.g., root, POS, and inflections). A headword may have several
meanings, which some lexicons designate into separate senses, while in others, senses need to be
designated and extracted. (v) Semantic Variations Lexicon: a set of pairs of semantically close lexical
entries and the differences between them, (e.g. like – love, pain – ache).</p>
      <p>In what follows, we present how the content of such types of lexicons is represented in Lemon
(illustrated in Figure 2), focusing on Lemon’s core semantic features:</p>
      <p>Lexical entry: Each translation term in a dictionary, a synonym in a thesaurus, a term in a
glossary, or a headword in a linguistic lexicon, is represented as a Lemon’s lexical entry.
Lexical concept: Each meaning of an entry (a gloss in a glossary, a set of synonyms in a thesaurus,
or a translations set in a dictionary) is represented as a Lemon’s lexical concept. For linguistic
lexicons, the different senses of a lexical entry, each is designated and mapped into a separate
lexical concept.</p>
      <p>Ontology concepts: Each entity in the Arabic Ontology is considered a Lemon’s ontology entity,
and is linked with lexical concepts in other lexicons using the Concept/isConceptOf properties,
Relations: If references to other senses are provided in a lexicon (i.e., semantic relations like
related, border/narrower, etc), we represent them as conceptRel in Lemon.</p>
      <p>Linguistic features: Glosses and sense definitions are represented using the skos:definition .
Features like POS, root, and inflections are specified using other properties in Lemon.</p>
      <p>As illustrated in Figure 1, an RDF icon is displayed beside each of the retrieved results. When
this icon is clicked, its lemon representation is generated and shown in a separate page. Figure 2
illustrates the Lemon representation of a lexical concept from the BZU Thesaurus, and its mapping
to the concept 293121 in the Arabic Ontology using the Concept property.
...
@prefix aot: &lt;http://ontology.birzeit.edu/term/&gt;.
@prefix ao:&lt;http://ontology.birzeit.edu/concept/&gt;.
@prefix aoc: &lt;http://ontology.birzeit.edu/lexicalconcept/&gt;. &lt;aot:lex-country&gt; a ontolex:LexicalEntry, ontolex:Word;
@prefix aor: &lt;http://ontology.birzeit.edu/lexicon/&gt;. ontolex:canonicalForm [ontolex:writtenRep ”country"@en];
&lt;aoc:1623&gt; a ontolex:LexicalConcept; skos:inScheme &lt;aor:BZU_Thesaurus_43&gt;.
ontolex:isEvokedBy &lt;aot:Lex-country&gt;; &lt;aot:lex-ﺔﻟود&gt; a ontolex:LexicalEntry, ontolex:Word;
ontolex:isEvokedBy &lt;aot:Lex-ﺔﻟود&gt;; ontolex:canonicalForm [ontolex:writtenRep "ﺔﻟود"@ar];
ontolex:isEvokedBy &lt;aot:Lex-دﻠﺑ&gt;; skos:inScheme &lt;aor:BZU_Thesaurus_43&gt;.
skos:definition "...ﻞﻜّﺸُﯾوﺐﻌْﺷَﮫﻟﺎﮭﯿﻠﻋﻖﻔﺘﻤُﻟاﺔﯿﺳﺎﯿﺴﻟا هدوُﺪﺤُﺑف ﺮَﻌُﯾ يرﺎﺒﺘﻋادﻮﺟﻮﻣ"@ar; &lt;aot:lex-دﻠﺑ&gt; a ontolex:LexicalEntry, ontolex:Word;
skos:inScheme &lt;aor:BZU_Thesaurus_43&gt;; ontolex:canonicalForm [ontolex:writtenRep "ﺪﻠﺑ"@ar];
ontolex:Concept &lt;ao:293121&gt;. skos:inScheme &lt;aor:BZU_Thesaurus_43&gt;.</p>
      <p>This representation is tentative. Each lexical entry in each lexicon is currently considered a
canonical form (i.e., lemma), but in fact it is not always the case. Unlike most English lexicons where
a lexical entry is often a lemma, Arabic entries are less often lemmas [9], for two main reasons:</p>
      <p>First – many Arabic lexicons do not strictly follow lemmatization conventions. Lemmas are
typically used as headwords in lexicons – representing a class of inflectionally related words with
the same meanings. In Arabic, the convention for a noun lemma is to be the singular masculine form,
and the third person singular perfective form for a verb lemma [8][9]. However, many lexicons are
less likely to follow this convention in practice. For example, inflected words like عرِ ﺎ ﺷَ (road) and
عرِ اﻮَ ﺷَ (roads), كرﺪﯾ (realizes) and كاردإ (realizing), or كاردإ (realization) and كاردﻹا (the realization)
might be used as separate headwords within the same or across lexicons. This means that, although
such lexical entries are used as separate headwords, they are not necessarily different lemmas. Thus,
they should not be considered separate canonical forms when representing them in Lemon. Such
cases are more likely to occur in case of dictionaries, glossaries and thesauri. Furthermore, some
linguistic lexicons use the same headword for different lemmas. For example, the same word ﺖﯿْﺑَ
could mean (house) with the plural تﻮﯿﺑ, and could mean (verse, a piece of poetry) with another
plural تﺎﯿﺑأ. Hence, such ambiguous words should be considered two headwords, or should be given
different lemma codes, like (ﺖﯿْﺑَ 1, ﺖﯿْﺑَ 2).</p>
      <p>Second – lexical entries in Arabic lexicons might be partially or not at all diacritized. Words in
Arabic consist of letters and diacritics, thus two words with different diacritics are not necessarily
the same word. The problem is that words are typically written without diacritics in practice. This is
acceptable by humans who can read and disambiguate words from their contexts, but is more
challenging for machines. Headwords in Arabic lexicons might be none or partially diacritized, which
makes it difficult to disambiguate them since they have no context, see our experiment in reducing
such disambiguation in [9]. For the Lemon representation, considering each headword in a lexicon
as a canonical form is not always correct, since headwords might not be fully diacritized; and thus,
two none or partially diacritized words might not be the same word within or across lexicons.</p>
      <p>To correctly represent Arabic lexical entries in Lemon, each Arabic lexical entry needs to be
carefully lemmatized first, which is a challenging task. That is, for each lexical entry, in each of
the 150 lexicons, its lemma should be specified. This would enable lexicons then to be interlinked
based on their lemmas. Although this is a challenging task as it cannot be fully automated [9], we
believe it cannot be avoided specially if lexicons need to be interlinked with external resources as
the Linguistic Linked Open Data Cloud.</p>
      <p>Furthermore, the Lemon morph module need to be extended to represent some Arabic-specific
linguistic and morphological features, such as imperfect and imperative verbs, verbal nouns,
intensive participle, place nouns, time nouns, instrumental nouns, and others.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we have presented a tentative representation of 150 Arabic multilingual lexicons using
the W3C’s Lemon model which can be accessed online. We have discussed the major challenges
related to representing lexical entries as canonical forms, especially lemmas and missing diacritics. We
plan to conduct a full lemmatization of all lexical entries and disambiguate them in case of missing
diacritics. We also plan to extend the Lemon morph module to represent Arabic-specific
morphological features.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgment</title>
      <p>This work is part of the Arabic Lemon project funded by the research committee at Birzeit university.
Dr. McCrae is supported in part by a research grant from Science Foundation Ireland (SFI)
under Grant Number SFI/12/RC/2289, co-funded by the European Regional Development Fund, and
the European Union’s Horizon 2020 research and innovation programme under grant agreement No
731015, ELEXIS - European Lexical Infrastructure and grant agreement No 825182, Prêt-à-LLOD.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1 2
          <string-name>
            <given-names>Hamzeh</given-names>
            <surname>Amayreh</surname>
          </string-name>
          , Mohammad Dwaikat, and
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          .
          <article-title>Lexicon Digitization -A Framework for Structuring, Normalizing</article-title>
          and Cleaning Lexical Entries,
          <year>2018</year>
          . URL: https://ontology.birzeit.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>edu/TR2018</source>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <surname>John P. McCrae</surname>
            , and
            <given-names>Paul</given-names>
          </string-name>
          <string-name>
            <surname>Buitelaar</surname>
          </string-name>
          .
          <source>Lexicon Model for Ontologies: Community Report</source>
          ,
          <year>2016</year>
          . URL: https://www.w3.org/
          <year>2016</year>
          /05/ontolex/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Gil</given-names>
            <surname>Francopoulo</surname>
          </string-name>
          , Nuria Bel, and et al.
          <article-title>Lexical Markup Framework (LMF) for NLP Multilingual Resources</article-title>
          .
          <source>In Workshop on Multilingual Language Resources and Interoperability. ACL</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          .
          <article-title>Towards the notion of gloss, and the adoption of linguistic resources in formal ontology engineering</article-title>
          .
          <source>In The 15th international conference on World Wide Web. ACM Press</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          .
          <article-title>Building a Formal Arabic Ontology (Invited Paper)</article-title>
          .
          <source>In Proceedings of the Experts Meeting on Arabic Ontologies and Semantic Networks. ALECSO</source>
          ,
          <string-name>
            <surname>Arab</surname>
            <given-names>League</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          .
          <source>The Arabic Ontology - An Arabic Wordnet with Ontologically Clean Content. Applied Ontology Journal</source>
          ,
          <year>2019</year>
          [Forthcoming].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hamzeh</given-names>
            <surname>Amayreh</surname>
          </string-name>
          .
          <article-title>An Arabic-Multilingual Database with a Lexicographic Search Engine</article-title>
          .
          <source>Proceedings of the 24th International Conference on Applications of Natural Language to Information Systems (NLDB)</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          , Nizar Habash, Faeq Alrimawi, Diyam Akra, and
          <string-name>
            <given-names>Nasser</given-names>
            <surname>Zalmout</surname>
          </string-name>
          .
          <article-title>Curras: An Annotated Corpus for the Palestinian Arabic Dialect</article-title>
          .
          <source>Journal Language Resources and Evaluation</source>
          ,
          <volume>51</volume>
          (
          <issue>3</issue>
          ):
          <fpage>745</fpage>
          -
          <lpage>775</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Mustafa</given-names>
            <surname>Jarrar</surname>
          </string-name>
          , Fadi Zaraket, Rami Asia, and
          <string-name>
            <given-names>Hamzeh</given-names>
            <surname>Amayreh</surname>
          </string-name>
          .
          <source>Diacritic-Based Matching of Arabic Words. ACM Asian and Low-Resource Language Information Processing</source>
          ,
          <volume>18</volume>
          (
          <issue>2</issue>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Khalfi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Nahli</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zarghili. Classical Dictionary</surname>
          </string-name>
          Al-Qamus in Lemon.
          <source>In 2016 4th IEEE International Colloquium on Information Science and Technology (CiSt)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Aïda</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          , Bilel Gargouri,
          <article-title>Abdelmajid Ben Hamadou, and Gil Francopoulo. ISO Standard Modeling of a Large Arabic Dictionary</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>22</volume>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Isa</given-names>
            <surname>Maks</surname>
          </string-name>
          , Carole Tiberius, and Remco van Veenendaal.
          <article-title>Standardising Bilingual Lexical Resources According to the Lexicon Markup Framework</article-title>
          .
          <volume>01</volume>
          <fpage>2008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>John</surname>
            <given-names>McCrae</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Chiarcos</surname>
          </string-name>
          , Francis Bond, Philipp Cimiano, Thierry Declerck, Gerard de Melo, Jorge Gracia,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann</surname>
          </string-name>
          , Bettina Klimek, Steven Moran, Petya Osenova, Antonio Pareja-Lora, and
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Pool</surname>
          </string-name>
          . The Open Linguistics Working Group:
          <article-title>Developing the Linguistic Linked Open Data Cloud</article-title>
          .
          <volume>05</volume>
          <fpage>2016</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Susanne</given-names>
            <surname>Salmon-Alt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Amine</given-names>
            <surname>Akrout</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Romary</surname>
          </string-name>
          .
          <article-title>Proposals for a Normalized Representation of Standard Arabic Full Form Lexica</article-title>
          .
          <source>In International Conference on Machine Intelligence</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>In International Conference on Islamic Applications in Computer Science and Technologies</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>