<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Is lemon Su cient for Building Multilingual Ontologies for Bantu Languages?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Catherine Chavula</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. Maria Keet</string-name>
          <email>mkeetg@cs.uct.ac.za</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Cape Town</institution>
          ,
          <country country="ZA">South Africa</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The current enormous amount of data on the Semantic Web and its increasing uptake raises the question of how this data can be accessed in several languages. OWL provides limited support for multilingualism through the use of an annotation property. However, it is known that more expressive models are required for linguistically demanding applications. Among the possible solutions, Lexicon Model for Ontologies (lemon) enables associating linguistic information with ontology elements by separating the lexical from the ontological layer. This paper investigates whether lemon is su cient for specifying multilingual ontologies for Bantu languages. Speci cally, the paper: (i) identi es the requirements for building lexica in lemon format for Bantu languages; (ii) describes the results in overcoming some of the challenges, notably concerning noun classes; and (iii) presents some open issues that will have to be addressed to increase usability of lemon.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Multilingual ontologies are required to provide access to ontology-based
information in the languages of the users. However, most ontologies are available
in English, i.e., ontology elements are named with English terms, which, at
least, brings afore the requirement to localise these ontologies to other
natural languages. For example, vocabularies for the Semantic Web such as
FriendOf-A-Friend (FOAF) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and GoodRelations [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] have annotations in English
only but are widely used to annotate resources on the Web. OWL provides
support for multilingualism using annotation properties such as rdfs:label and
rdfs:comment; e.g., a lexicalisation of the class foaf:Person can be given in English,
Chichewa, and isiZulu through adding annotations as shown in Fig. 1. However,
the amount of linguistic annotation that can be included in this labelling system
is limited and most multilingual applications require more data such as Part of
Speech (POS) and grammatical features, among others. Moreover, ontologies are
for representing knowledge, and such linguistic data need not to be included in
an ontology. Several approaches that separate the ontological layer from the
terminological layer have been proposed [
        <xref ref-type="bibr" rid="ref24 ref6 ref9">6, 9, 24</xref>
        ] and the W3C Community Group
ontolex-lemon submission is under way1. Notably, the LExicon Model for
ONtologies (lemon) [
        <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
        ] separates the ontological layer and linguistic layer, and
1 http://www.w3.org/community/ontolex/wiki/Final_Model_Specification
&lt;rdfs:Class rdf:about="http://xmlns.com/foaf/0.1/Person"&gt;
&lt;rdfs:label xml:lang = "en"&gt; person &lt;/rdfs:label&gt;
&lt;rdfs:label xml:lang = "ny"&gt; munthu &lt;/rdfs:label&gt;
&lt;rdfs:label xml:lang = "zu"&gt; umuntu &lt;/rdfs:label&gt;
&lt;/rdfs:Class&gt;
is gaining momentum in adoption for multilingual ontologies. In lemon, each
ontology element is associated with an entry in a separate lexical resource, which
in turn is annotated with linguistic data. In this manner, the ontology provides
the semantics of terms in a lexical resource while the entries provide the
lexicalisation of the ontology elements. This looks like a promising solution to problems
identi ed for Indigenous Knowledge Management Systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as well as possible
ontology-driven applications in the ICT4D area and ontology verbalisation [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        Bantu languages are characterised by a comprehensive noun class and
concordial agreement system among terms. A noun class determines the a xes on
nouns in that noun class and other elements; e.g. umfula (`river') is in noun class
3, where -fula is the stem and um- the pre x for that noun class. Each noun class
has its own concords for the noun and lexical categories such as adjectives and
verbs. This is emblematic for all Bantu languages that have between 10 and 23
noun classes, depending on the language. This paper investigates whether lemon
can be used to model these characteristics to create multilingual ontologies with
Bantu languages terms. The general issue on representing lexical information
is addressed, which requires a new ontology module that has noun class as a
grammatical category|the new noun class system ontology|and the use of the
lemon morphology module is described with elements of that ontology. The
development of a lexicon for FOAF [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and GoodRelations [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] in Chichewa are
used to examine implementability of the requirements for multilingual
ontologies to accommodate Bantu languages. To ensure the evaluation and proposed
solution is not tted to Chichewa only, isiZulu is also considered, which is in a
di erent sub-family of Bantu languages.
      </p>
      <p>The remainder of this paper is structured as follows. Section 2 describes
requirements for ontology-based applications with Bantu languages and Section 3
discusses related work. Section 4 describes the process of enriching a domain
ontology with lexical information using the lemon model. Section 5 discusses the
challenges in the process of developing the resources and Section 6 concludes
and presents ideas for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Linguistic Requirements for Ontologies in Bantu</title>
    </sec>
    <sec id="sec-3">
      <title>Languages</title>
      <p>
        Bantu languages are characterised by complex morphosyntactic features due to
a Noun Class System (NCS) and a system of concordial agreement: each noun
belongs to a noun class (nc) and each class has its collection of a xes, which then
also determines the agreement markers (grammatical concord) on related lexical
categories such as adjectives and verbs. For instance, ubuntu is in nc:14
(ubu+-ntu) and umuntu in nc:1 (umu-+-ntu) in isiZulu, and munthu (mu-+-nthu)
in nc:1 in Chichewa. While each noun class typically has some semantics (e.g.,
nc:1 for humans and other animates), the semantics of the noun classi cation in
di erent Bantu languages is still a topic of investigation [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Bantu noun classes
are identi ed using Arabic Numerals based on di erent classi cation methods
and naming schemes. Meinhof's scheme of 1948 consists of a generic table for
Bantu languages commonly used in comparative studies [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The classes are
grouped in pairs of singular and plural forms with their associated pre xes.
The number of classes varies among languages, with some languages exhibiting
similarities in the pre xes; see Table 1 for examples. A linguistic requirement
for multilingual ontologies of Bantu languages is thus to have a way to annotate
the nc of the term whose sense is denoted by the OWL class.
      </p>
      <p>
        In addition, an ontology-based task such as ontology verbalisation requires
information about a class's nc to determine what combination of pre xes to
add to the verb in the name of the object/data property [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. For example, the
foaf:knows property has as domain and range foaf:Person, and in order to verbalise
a fact using these vocabulary elements, the nc for person as well as its associated
pre xes such as tense need to be available. Lexicalising the fact, `John knows Jim'
in Chichewa would be John amadziwa Jim (agreement marker underscored), but
if the domain was of a di erent class (not a person), then the agreement marker
for that other nc has to be used. For instance, eats with verb stem -dla: when
a gira e (in nc:9) eats something it is idla and for a person (nc:1) it is udla.
That is, an object or data property is not simply named with a verb in third
person singular, as is deemed good practice in ontology engineering, but the
term depends on the domain and range and its use in an axiom.
      </p>
      <p>
        As all peculiarities of Bantu languages cannot be covered here, this study uses
Chichewa and isiZulu. Chichewa, a dialect of the Nyanja language, is spoken by
over 12 Million people in Malawi [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], and isiZulu is among the Nguni languages
of South Africa spoken as L1 language by over 10 million people. According to
the Guthrie classi cation [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] of Bantu languages into zones based on language
characteristics, Chichewa is in zone N, unit N31, and isiZulu in zone S, unit S42.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Related Work</title>
      <p>
        Research into multilingual ontologies investigates models for representing a
lexical/terminological layer for ontologies and so far three architectures have been
proposed [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], namely: (i) using multilingual OWL annotation property, such
as rdfs:label and rdfs:comment; (ii) mapping ontology elements designed for
di erent cultures and languages, e.g., EuroWordNet2 and BabelNet [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]; and
(iii) using external lexical resources to linguistically enrich ontologies [7, 6, 9, 10,
      </p>
      <sec id="sec-4-1">
        <title>2 www.illc.uva.nl/EuroWordNet/</title>
        <p>
          27]. The last approach provides a means of modelling morphosyntactic features
required by more linguistically demanding tasks such as ontology-based
information access, and is the focus of this paper. Regarding models for building lexical
resources, there have been di erent approaches for publishing lexical data, and
WordNet is the most popular English lexical database [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], which is organised
around semantic relationships between words called synsets (i.e., synonym sets).
The WordNet model has also been applied to other languages and e orts to
create a global WordNet are underway3. The WordNet model, however, is limited
in working with environments of diverse requirements [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] and Bantu languages
have a complex morphosyntactic structure which cannot be modelled in the
WordNet structure. The Lexical Markup Framework (LMF) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], an ISO
stan
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>3 http://globalwordnet.org/wordnets-in-the-world/</title>
        <p>
          dard for representing lexica in XML/UML, provides a format for data exchange
and interoperability. However, it is di cult to fully exploit lexica in LMF
format since the semantics of the model are not formalised. Models that associate
lexical information with ontological elements such as LexInfo [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], LexOnto [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ],
Linguistic Information Repository (LIR) [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], LingInfo [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and lemon [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] have
been proposed. Lemon is informed by LMF, LIR and LexInfo, and provides a
means for separating the lexical and semantic layer of the ontology. A lemon
lexicon de nes how ontological elements are realised in a particular natural
language. The lemon design is based on a premise that a word sense relates to the
conceptualisation of the world which can be speci ed in an ontology. It provides a
rich model for modelling semantic multilingual knowledge thanks to its modular
design and extensibility. It consists of a core module and ve task-speci c
modules, namely: Morphology, Syntax and Mapping, Phrase structure, Variation and
Linguistic Description [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. lemon has been used with many languages, notablys
English [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] and German [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], and a collection of 50 languages in BabelNet 2.0
[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. To promote interoperability among the linguistic resources, agreed-upon
grammatical categories are used. For example, LMF and lemon advocate using
general language ontologies such as GOLD ontology [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and the ISOcat data
category registry [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>
          Zooming in on lexical data models for Bantu languages, Bosch et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
proposed to represent South African Bantu languages' complex structure using
sub-entries. A typical entry is organised into head and body tags, with head
containing the stem or root and body containing syntactic and morphological
information about the stem. From the entries in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], it can be seen that repetition
is used to model morphological processes such as in ection. However, repetition
of NCS data in the lexicon can make the size of the lexicon to grow very big
as the same information is repeated within each entry. The lemon morphology
module provides an economical way of modelling highly in ectional languages.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Building Lemon Lexica in Bantu Languages</title>
      <p>Before immediate usage, several practical modelling and design choices have to
be made, which are described rst. This is followed by an assessment of the
practical feasibility of using lemon with Bantu languages, by lexicalising the
FOAF and GoodRelations ontologies in Chichewa.
4.1</p>
      <sec id="sec-5-1">
        <title>General Modelling Aspects</title>
        <p>
          Ontology-based applications for Bantu languages require morphosyntactic data,
notably: (i) de ning a NCS with associated pre xes and associating the NCS
with the lexical entries; (ii) de ning rules for verbs and adjectives to ensure
agreement with the noun class; and (iii) writing rules for agglutination process.
Modelling the NCS. The rst requirement on handling the NCS can be
accommodated by lemon's extensible approach that promotes the use of externally
de ned properties to annotate lexica. lemon advocates the usage of linguistic
ontologies such as GOLD [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and linguistic data repositories such as ISOcat [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and
user de ned properties. However, the properties de ned in GOLD and ISOcat do
not fully meet the requirement. GOLD has a concept gold:ArabicNumeralGender,
re ecting that the NCS has been proposed by some Bantu linguists [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] as a
grammatical gender system. The ISOcat so-called \data categories" has an
isocat:otherGender that is intended to be used in the same manner analogous to
gender that has genders as instances. However, these properties do not have
their analogous instances for the 23 Bantu noun classes and thus do not
provide a means of labelling the entries. As lemon's LexInfo ontology for linguistic
annotation imports ISOcat, it does not provide for this either. Moreover, it is
arguable whether the NCS is the same as gender as is commonly understood
(masculine etc.; see also [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]), as, e.g., isiZulu nc:9 is for animals and nc:14 for
abstract nouns, i.e., semantic groups that have nothing to do with that
interpretation of gender. Di erent schemes have been developed to specify the Bantu
NCS for individual Bantu languages. Instead of encoding each scheme, we note
that the Meinhof (1948) one is used among linguists as standard for de ning
noun classes; hence, this is used here as well, for it facilitates cross-language
comparisons and use.
        </p>
        <p>Because of the shortcomings of the extant linguistic resources, an ontology
based on the Bantu languages' noun class system was developed to allow the
annotation of ontology elements with noun class information. The goal of the
ontology was to specify the conceptualisation of nouns in the sense of Bantu
languages structure. In order to achieve this, the ncs:partOfSpeech class was
introduced, which subsumes ncs:Noun, and a ncs:Property class was added that
subsumes ncs:MorphoSyntacticProperty to capture the characteristics of nouns in this
domain. The ncs:Gender class is made a sibling class of ncs:NounClass, which are
subsumed by ncs:MorphoSyntacticProperty. In this regard, other classes for
classifying nouns that do not relate to the NCS and nominal morphology can be
added without interfering with the conceptualisation of the NCS. The Bantu
languages NCS concepts based on the Meinhof noun classi cation scheme were
speci ed with the assumption that all speci c noun class classi cations of Bantu
languages can be aligned to this scheme. The Meinhof classi cation has 23 classes
labelled mostly with arabic numerals only. However, classes recognised later after
the Meinhof scheme are tabulated as subdivisions of Meinhof classes to freeze
the number of classes. Therefore, the numbering scheme of Arabic numerals is
sometimes augmented with an alphabetical letter, such as class 1a, 2b and 3a
to signify that the classes had elements of class 1, 2, and 3, respectively (but
is disjoint from it). In order to capture all these complexities, classes labelled
with Arabic Numeral classes as well as their derivatives were added as separate
disjoint classes.</p>
        <p>
          Object properties were de ned in the ncs ontology to specify the
relationships between the classes in the ontology as conceptualised in the domain of
Bantu Languages. The properties ncs:hasNounclass, ncs:hasNumber, ncs:hasPlural
and ncs:hasSingular were introduced. Additionally, restrictions were added to
specify the constraints on the relationship as well as relationship characteristics. For
example, ncs:hasPlural is used to specify that the classes can only have speci c
classes as their plural and vice versa using ncs:hasSingular and that these classes
are inverse of each other. Fig. 2 depicts some of the classes and an
annotation, and the ontology is accessible online at http://www.meteck.org/files/
ontologies/. Following this, the nouns of the names of the classes in an ontology
can now be annotated with their noun class.
Specifying Rules for Word Variation. The NCS determines how words of
grammatical categories such as nouns are in ected by adding pre xes to stems.
The lemon morphology module uses transformational rules to account for the
in ection of words in a particular language using PERL-like regular
expressions. The nominal morphology in Bantu languages is based on the pre xes of
noun classes while other syntactic categories like verbs and adjectives depend
on agreement markers prescribed by the noun classes. For example, a verb stem
in Chichewa can have over 20 di erent forms generated through the
concatenation of agreement markers and other syntactic components to roots, in a similar
way as the eats example for isiZulu in the previous section. Although rules can
be written once for each language using the lemon morphology module, the
transformation system for the arbitrary case is very complex and would both
overburden the lemon system and render such rules useless outside the setting
of multilingual ontologies. No computational version of such generative rules
exist yet (except initial theoretical results [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]), and in that light it will be more
e ective to develop a separate generic grammar engine that can be plugged in.
Nevertheless, it may be feasible to write a subset of the rules to generate lexica
for `limited' ontologies, like those that are mainly taxonomies and those where
object/data properties have their domain and range declared in detail. An
example is shown in Fig. 3 for partial noun morphology for nc:1 and nc:2 using
the lemon morphology module. The rule is partial, as the pre x can di er also
depending on the stem.
:NYNC1_2 a lemon:MorphPattern ;
lemon:transform [
lemon:rule "mu~" ;
lemon:generates [
lexinfo:number lexinfo:singular;
ncs:hasNounClass ncs:class1
]],[
lemon:rule "a~" ;
lemon:generates [
lexinfo:number lexinfo:plural;
ncs:hasNounClass ncs:class2]].
Agglutination. Bantu languages are highly agglutinative, to the extent that a
word can be composed of over ve constituents. For comparatively well-studied
languages within the Bantu language family, a few rules can be written based
on the Morphology module of lemon. However, some of the aspects of the Bantu
languages morphology cannot be handled using the proposed approach which
favours concatenation morphology. Due to space limitations, we do not discuss
those details here. The shortest and simplest example for an object property
(verb) is included in Fig. 5, ignoring such issues as past tense, variation due
to the domain being of a di erent noun class, and the habit in `English OWL
ontologies' to include a preposition in the name of the object property. It is
not doable in the general case as a form variation generator and novel methods
need to be developed to accommodate such phenomena for object properties.
More complications exist if one were to partially verbalise writing an axiom, as
in Protege's \class expression editor", as even conjunction `and' can be glued to
the second class, modifying it [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] (e.g., ushizi becomes noshizi for `and cheese').
4.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>FOAF and GoodRelations lexica in Chichewa</title>
        <p>The lexica for the selected ontologies were written manually due to
unavailability of language resources and tools for the two languages. The translation
of the term of each ontology element was collected and a further analysis was
done to determine how it can be lexicalised in lemon format. The lexicalisation
ended up to be a 1:1 correspondence between the lexical entry and the English
terms of ontology elements. Due to the absence of some terms in Chichewa,
more speci cally for technical terms, some of the elements were not lexicalised:
no equivalences were provided for FOAF phrases foaf:'sha1sum of a personal
mailbox URI name', foaf:DNAchecksum (a joke), and foaf:openid. An example of a fully
lexicalised entry is shown in Fig. 4 for foaf:Person (munthu), specifying its part
of speech, plural form and sense referenced through the FOAF foaf:Person.
:munthu a lemon:Word;
lexinfo:partOfSpeech lexinfo:noun;
lemon:canonicalForm [lemon:writtenRep "munthu"@ny;
lexinfo:number lexinfo:singular
ncs:hasNounClass ncs:class1] ;
lemon:otherForm [lemon:writtenRep "anthu"@ny;
lexinfo:number lexinfo:plural;
ncs:hasNounClass ncs:class2];
lemon:sense [lemon:reference foaf:Person].</p>
        <p>Most of the FOAF object properties consisted of multi-word phrases and
their lexicalisations turned out to be complete sentences in a natural language
and has the issue of variation depending on the noun class of the noun of the
name of the OWL class associated to it. Where possible, the lemon Morphology
module was used to write the rules for the verbs, as illustrated in Fig. 5 for
foaf:knows (dziwa). This was doable, as both the domain and range of foaf:knows
are declared to be foaf:Person, therewith greatly reducing the set of possible rules
for this object property. Data properties in the FOAF vocabulary are associated
with aspects of people such as birthday, rst name and surname, which were
easier to translate. Overall, the FOAF lemon lexicon in Chichewa covers 90%
of the classes and individuals of the original English FOAF, and 80% of its
properties.
:NYNC1_2_verb_stem a lemon:MorphPattern ;
lemon:transform
:present_transform ,
:agreement_transform.
:present_transform
lemon:rule "ma~" ;
lemon:rule "i~/ma~";
lemon:nextScope :agreement_transform.
lemon:generates [lexinfo:tense lexinfo:present].
:agreement_transform</p>
        <p>lemon:rule "a~" ;
lemon:generates [lexinfo:number lexinfo:plural],
[lexinfo:number lexinfo:singular]].
:dziwa a lemon:LexicalEntry.
:dziwa lemon:pattern :NYNC1_2_verb_stem.
lemon:abstractForm [lemon:writtenRep "dziwa"@ny;
lexinfo:partOfSpeech lexinfo:verb];
lemon:sense [lemon:reference foaf:knows].</p>
        <p>
          The GoodRelations ontology is an e-commerce vocabulary for product, price,
store and company data and is widely used on the web for annotating resources
and has the support of Yahoo and Google [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. The ontology is highly domain
speci c and lexicalisation was a big challenge, partly due to the absence of
suitable terms in Chichewa. Only 25% of the entities were lexicalised in Chichewa,
which mostly included nouns of classes that belong to noun classes that were
not already covered in the FOAF vocabulary.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>
        The primary challenge of building lemon lexica in Bantu languages is how to
handle the NCS and morphological structure of these languages. Modelling the NCS
as a gender classi cation is unsatisfactory and the gender grammatical category
in ISOcat data category registers and GOLD is irrelevant. Hence, the
requirement to build an ontology module to support this aspect. Bantu languages have
a complex morphological structure based on the NCS. However, lemon
morphology module provides a limited rule encoding feature. An option to separate the
rules from the lexica so as to foster their reusability would be bene cial. Other
challenges not limited to Bantu Languages include:
{ Limited multilingual support of OWL. OWL has limited support for
internationalization and ontology localization to build true multilingual ontologies
which can support deep semantic processing.
{ Limitations of the lexical-semantic relationship in lemon . The direct relation
of the ontology with the lexical resources cannot yet be fully exploited in
applications that require referencing non 1:1 mappings (e.g., there are 19
di erent translations for `part' in isiZulu). The variation module is limited
to model this aspect, and their associated e ects on inferences are unclear.
{ Heterogeneity of linguistic annotations. GOLD and ISOcat use di erent
formalisms for modelling linguistic annotations, and in most cases lexica can
use properties from di erent communities which cannot be aligned, creating
some confusion and incompatibilities. Although, some work exist on
linguistic annotations based on OWL DL [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], most of the resources are not based
on the more stricter pro les and no reasoning on properties can be done.
{ Poor tool support and documentation. Tools that can be incorporated into
existing Semantic Web platforms have to be developed. More documentation
would also be helpful.
      </p>
      <p>
        Bantu languages have their own language speci c challenges. There is still some
work to be done in the eld of linguistics on topics such as semantic classi cation
of nouns and morphology in general. Additionally, no satisfactory method for
word generation has yet been proposed. Previous studies in other Bantu
languages have shown that the regular expression methods are limited, and the
morphology module of lemon may serve less for Bantu languages lexica [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In
addition, there is relatively little syntactic knowledge and a few language
resources to enable other avenues of research in this area [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. lemon can currently
be used in building lexica for small general ontologies that are in general domains
or those relevant for the application context.
6
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion and future work</title>
      <p>
        This paper has presented some challenges of building lemon lexica in Bantu
languages. It required the building of a noun class system ontology module, ncs.owl,
to allow annotation for noun classes. While rules can be encoded, the
complexity of rules for Bantu languages makes it a challenge for lemon. Localisation of
FOAF and GoodRelations into Chichewa have been experimented with,
resulting in a near-full coverage of FOAF, mainly through hard-coding the classes and
using only comparatively short rules for the object properties. Some remaining
challenges were outlined. The ontology and FOAF and GoodRelations
lexicalizations in Chichewa are available at www.meteck.org/files/ontologies/. The
next step focuses on the use of multilingual ontologies for tasks such as ontology
verbalisation, of which the rst results have been obtained [
        <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
        ] and
multilingual access to data as well as evaluate other models such as ontolex-lemon.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alberts</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fogwill</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keet</surname>
            ,
            <given-names>C.M.:</given-names>
          </string-name>
          <article-title>Several required OWL features for indigenous knowledge management systems</article-title>
          .
          <source>In: Proc. of OWLED'12. CEUR-WS</source>
          , vol.
          <volume>849</volume>
          , p.
          <source>12p</source>
          (
          <year>2012</year>
          ),
          <fpage>27</fpage>
          -28 May, Heraklion, Crete, Greece
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bentley</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulemeka</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : Chichewa. Lincom Europa, Munchen (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bosch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pretorious</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Towards Machine-Readable Lexicons for South African Bantu languages</article-title>
          .
          <source>Nordic Journal of African Studies</source>
          <volume>16</volume>
          (
          <issue>2</issue>
          ),
          <volume>131</volume>
          {
          <fpage>145</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brickley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Libby</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <source>FOAF Vocabulary Speci cation 0</source>
          .99 . http://xmlns.com/ foaf/spec/ (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Broeder</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>A data category registry-and component-based metadata framework</article-title>
          .
          <source>In: Proc. of LREC'10</source>
          .
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          (
          <year>2010</year>
          ),
          <fpage>19</fpage>
          -21 May,Valletta, Malta
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sintek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>LingInfo:Design and Applications of a Model for the Integration of Linguistic Information in Ontologies</article-title>
          .
          <source>In: Proc. of LREC'06</source>
          .
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          (
          <year>2006</year>
          ),
          <fpage>22</fpage>
          -28 May, Genoa, Italy
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sintek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A multilingual/multimedia lexicon model for ontologies</article-title>
          .
          <source>In: Proc. of ESWC'06, LNCS</source>
          , vol.
          <volume>4011</volume>
          , pp.
          <volume>502</volume>
          {
          <fpage>513</fpage>
          . Springer (
          <year>2006</year>
          ), budva, Montenegro
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chiarcos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Ontologies of linguistic annotation: Survey and perspectives</article-title>
          . In: Calzolari,
          <string-name>
            <surname>N.</surname>
          </string-name>
          , et al.
          <source>(eds.) Proc. of LREC'12. ELRA</source>
          , Istanbul, Turkey (may
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sintek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Lexinfo: A declarative model for the lexicon-ontology interface</article-title>
          .
          <source>Web Semant</source>
          .
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <volume>29</volume>
          {
          <fpage>51</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haase</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herold</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mantel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Lexonto: A model for ontology lexicons for ontology-based NLP</article-title>
          .
          <source>In: Proc. of OntoLex07 Workshop held in conjunction with ISWC'07</source>
          .
          <string-name>
            <surname>ISW</surname>
          </string-name>
          (
          <year>2007</year>
          ), 11 Nov, Busan, South-Korea
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Exploiting ontology lexica for generating natural language texts from RDF data</article-title>
          .
          <source>In: Proc. of the 14th European Workshop on NLG</source>
          . pp.
          <volume>10</volume>
          {
          <fpage>19</fpage>
          .
          <string-name>
            <surname>ACL</surname>
          </string-name>
          (
          <year>2013</year>
          ),
          <article-title>so a</article-title>
          ,
          <source>Bulgaria</source>
          ,
          <fpage>8</fpage>
          -
          <lpage>9</lpage>
          Aug 2013
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Cobert</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mtenje</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gender Agreement in Chichewa</article-title>
          .
          <source>Studies in African Linguistics</source>
          <volume>18</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>38</fpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ehrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cecconi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vannella</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.:
          <article-title>Representing multilingual data as linked data: the case of BabelNet 2.0</article-title>
          .
          <source>In: Proc. of LREC'14</source>
          . pp.
          <volume>401</volume>
          {
          <fpage>408</fpage>
          .
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          (
          <year>2014</year>
          ), reykjavik, Iceland, May
          <volume>26</volume>
          -31
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Espinoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ontology localization</article-title>
          .
          <source>In: Proc. of K-CAP'09</source>
          . pp.
          <volume>33</volume>
          {
          <fpage>40</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -4 Sept,
          <year>2009</year>
          , Redondo Beach, California, USA
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Farrar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langendoen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Markup and the GOLD ontology</article-title>
          .
          <source>In: Proc. of workshop on digitizing and annotating text and eld recordings</source>
          . pp.
          <volume>845</volume>
          {
          <fpage>863</fpage>
          . LSA Institute (
          <year>2003</year>
          ),
          <source>july 11th -13th</source>
          , Michigan, USA
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.):
          <article-title>WordNet: An electronic lexical database</article-title>
          . MIT press, Cambridge (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Francopoulo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>George</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monachini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bel</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pet</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Lexical Markup Framework</article-title>
          .
          <source>In: Proc. of LREC'06</source>
          . vol.
          <volume>17</volume>
          , pp.
          <volume>233</volume>
          {
          <fpage>236</fpage>
          .
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          (
          <year>2010</year>
          ),
          <fpage>genoa</fpage>
          - Italy,
          <fpage>24</fpage>
          -26 may
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Getao</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miriti</surname>
          </string-name>
          , E.:
          <article-title>Computational Modelling in Bantu Language</article-title>
          .
          <source>Advances in Systems Modelling and ICT Applications</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ),
          <volume>128</volume>
          {
          <fpage>138</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Guthrie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Comparative Bantu: An Introduction to the Comparative Linguistics and Prehistory of the Bantu Languages</article-title>
          . No. v. 1-
          <issue>2</issue>
          ,
          <string-name>
            <surname>Gregg</surname>
          </string-name>
          (
          <year>1971</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Hepp</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>GoodRelations: An Ontology for Describing Products</article-title>
          and
          <article-title>Services Offers on the Web</article-title>
          .
          <source>In: Proc. of EKAW'08. LNCS</source>
          , vol.
          <volume>5268</volume>
          , pp.
          <volume>332</volume>
          {
          <fpage>347</fpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Keet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khumalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Basics for a grammar engine to verbalize logical theories in isiZulu</article-title>
          .
          <source>In: Proc. of RuleML'14. LNCS</source>
          , vol.
          <volume>8620</volume>
          , pp.
          <volume>216</volume>
          {
          <fpage>225</fpage>
          . Springer (
          <year>2014</year>
          ),
          <fpage>18</fpage>
          -
          <lpage>20</lpage>
          August, Prague, Czech Republic
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Keet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khumalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Toward verbalizing logical theories in isiZulu</article-title>
          .
          <source>In: Proc. CNL'14. LNAI</source>
          , vol.
          <volume>8625</volume>
          , pp.
          <volume>78</volume>
          {
          <fpage>89</fpage>
          . Springer (
          <year>2014</year>
          ),
          <fpage>20</fpage>
          -
          <lpage>22</lpage>
          August, Galway, Ireland
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spohr</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Linking lexical resources and ontologies on the semantic web with lemon</article-title>
          .
          <source>In: Proc. of ESWC'11. LNCS</source>
          , vol.
          <volume>6643</volume>
          , pp.
          <volume>245</volume>
          {
          <fpage>259</fpage>
          . Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Interchanging lexical resources on the Semantic Web</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>46</volume>
          (
          <issue>4</issue>
          ),
          <volume>701</volume>
          {719 (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>The lemon cookbook</article-title>
          .
          <source>Technical report, Monnet Project (June</source>
          <year>2012</year>
          ), www.lemon-model.net
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Mchombo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Chichewa. In: Spencer,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Zwicky</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.M.</surname>
          </string-name>
          <article-title>(eds.) The Handbook of Morphology</article-title>
          . Blackwell Publishers (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Modelling multilinguality in ontologies</article-title>
          .
          <source>In: Coling</source>
          <year>2008</year>
          : Companion volume -
          <source>Posters and Demonstrations</source>
          . pp.
          <volume>67</volume>
          {
          <issue>70</issue>
          (
          <year>2008</year>
          ), manchester,UK, August
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Ngcobo</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          :
          <article-title>Zulu noun classes revisited: A spoken corpus-based approach</article-title>
          .
          <source>South African Journal of African Languages</source>
          <volume>1</volume>
          ,
          <issue>11</issue>
          {
          <fpage>21</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A lemon lexicon for DBpedia</article-title>
          .
          <source>In: Proc. of the NLP &amp; DBpedia workshop. CEUR-WS</source>
          , vol.
          <volume>1064</volume>
          , pp.
          <volume>21</volume>
          {
          <issue>25</issue>
          (
          <year>2013</year>
          ), sydney, Australia
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>