<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DBlexipedia: A nucleus for a multilingual lexical Semantic Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sebastian Walter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christina Unger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Cimiano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Semantic Computing Group, CITEC, Bielefeld University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>A huge amount of datasets on the Semantic Web are linked to a few datahubs, the most prominent of which is DBpedia. What makes the exploitation of DBpedia challenging for natural language-based applications, however, is that such NLP applications require knowledge about how the ontology elements are verbalized in natural language. In order to provide such knowledge at the required scale and thereby leverage the use of DBpedia in di erent applications, we construct a lexicon for the DBpedia 2014 ontology by means of existing automatic methods for lexicon induction. It contains 11,998 lexical entries for 574 di erent properties in three languages: English, German, and Spanish. Just like DBpedia provides a hub for Semantic Web datasets, this lexicon can provide a hub for the lexical Semantic Web, an ecosystem in which ontology lexica are published, linked, and re-used across applications.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology lexicalization</kwd>
        <kwd>DBpedia</kwd>
        <kwd>lemon</kwd>
        <kwd>M-ATOLL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The amount of datasets on the Semantic Web is evermore increasing, and large
part of these datasets are linked to central hubs. The biggest and arguably most
important of those hubs is DBpedia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a general-purpose, multi-domain dataset
extracted from Wikipedia, which is, for example, heavily used by systems that
employ structured data for applications like web-based information retrieval or
search. But what makes the exploitation of DBpedia challenging for a variety of
natural language-based applications (such as question answering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and natural
language generation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) is that they usually require knowledge about how the
ontology elements are verbalized in natural language, including lexical variants
across di erent languages. In order to provide such knowledge at the required
scale and thereby leverage the use of DBpedia and all connected datasets in
applications, we construct a lexicon for the DBpedia 2014 ontology by means of
existing automatic methods for lexicon induction. Just like DBpedia provides a
hub for Semantic Web datasets, this lexicon can provide a hub for the lexical
Semantic Web, an ecosystem part of the linguistic linked data cloud1 in which
ontology lexica are published, linked, and re-used across applications.
      </p>
      <p>For an example of a lexical entry consider the property http://dbpedia.
org/ontology/spouse from the DBpedia 2014 ontology. This property expresses</p>
      <sec id="sec-1-1">
        <title>1 http://linguistic-lod.org/llod-cloud</title>
        <p>that two persons are married to each other. The property can be expressed in
natural language by the following expressions (lexical entries).</p>
        <p>{ X is the husband of Y e.g. Barack Obama is the husband of Michelle Obama.
{ X is the wife of Y e.g. Michelle Obama is the wife of Barack Obama.
{ X is married to Y e.g. Barack Obama is married to Michelle Obama.</p>
        <p>All these lexical entries can be interpreted as expressing the spouse property.
We say that the property spouse is the reference of the lexical entry.</p>
        <p>To our knowledge, DBlexipedia is the rst wide-coverage lexicon of
DBpedia. It contains 11,998 lexical entries for 574 di erent RDF-properties from the
DBpedia 2014 ontology in three languages: English, German, and Spanish.</p>
        <p>
          In contrast to an earlier manually crafted DBpedia lexicon [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], DBlexipedia
is automatically constructed and therefore can easily be updated as DBpedia
evolves.
        </p>
        <p>The remainder of the paper is structured as follows. In the next section, we
present our approach to generating ontology lexica, followed by a description
of the resulting lexicon in Section 3. We conclude with perspectives for future
work.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>
        The lexicon published on http://dblexipedia.org is the result of applying
M-ATOLL2 [
        <xref ref-type="bibr" rid="ref10 ref11">11, 10</xref>
        ] to the DBpedia ontology and a Wikipedia text corpus.
MATOLL creates ontology lexica in lemon [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] format, and it is designed to be
employed in a semi-automatic fashion, i.e. automatically constructing lexical
entries that are then manually checked and corrected by a human.
      </p>
      <p>In the following we brie y explain the main corpus-based approach as well as
an extension of it that deals with a special case of adjective entries. We call this
second approach label-based, as we use the label of a property in combination
with machine learning techniques to generate the entries.
2.1</p>
      <sec id="sec-2-1">
        <title>Corpus-based Approach</title>
        <p>
          M-ATOLL takes as input an ontology and a dependency parsed text corpus in
the target language. For DBlexipedia, the input was the DBpedia 2014 ontology
and a Wikipedia text corpus parsed for three target languages: English, German,
and Spanish. As dependency parser we use the MaltParser [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] for English, the
ParZu parser [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for German, and an online service with an instance of the
Spanish MaltParser [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] for Spanish.
        </p>
        <p>M-ATOLL performs three steps in order to nd lexicalizations of ontology
properties in the accompanying text corpus:
1. Retrieving all triples for a given property from the ontology. For example, the
results for the property spouse include the triple &lt;Barack Obama,spouse,
Michelle Obama&gt;.</p>
        <sec id="sec-2-1-1">
          <title>2 https://github.com/ag-sc/matoll</title>
          <p>2. Retrieving all sentences from the parsed text corpus which contain mentions
of the subject and object of the triples discovered in Step 1.
3. Searching for prede ned patterns in those sentences, in order to extract
candidate lexicalizations of the property.</p>
          <p>So far, M-ATOLL covers entries that describe transitive verbs (e.g. to cross),
intransitive verbs with a prepositional object (e.g. to live in), relational nouns
with prepositional object (e.g. capital of ), and relational adjectives (e.g. similar
to) in all three languages: English, German, and Spanish. Important to note is
that one entry can have multiple references, e.g. the relational noun entry village
in3 can refer to hometown, birthplace, and location, among others.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Label-based Approach</title>
        <p>Examining the lexical entries created by the above approach and the kind of
lexicalizations needed for question answering, for example, it is clear that
MATOLL does not yet cover all relevant patterns. Consider the request Give me all
female Danish politicians. In this case a lexical entry for female is needed, which
w.r.t. DBpedia refers to all individuals that are related to the resource Female
by means of the property gender. Similarly, Danish refers to all individuals
that are related to the resource Denmark by means of the property country or
birthPlace.</p>
        <p>
          As these adjective lexicalizations are not handled by the standard
corpusbased approach, we extended M-ATOLL with a dedicated module, extending
[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. This approach is based on the observation that a lot of those adjectives
actually occur in the labels of the objects, e.g. female in &lt; ,gender,Female&gt; or
Catholic in &lt; ,religion,Catholic Church&gt;. We therefore check for adjectives
occurring in object labels, training an SVM in order to decide whether those
adjectives are valid laxicalizations of the restriction class in question.
        </p>
        <p>
          Resulting entries are, for example, the above mentioned female4, and blue5.
For more information about this approach and the implemented features, see
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>This approach currently works only for English, but will be soon adapted to
German and Spanish.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Dataset</title>
      <p>The dataset we present in this paper is the result of the above approaches and
contains entries in three languages: English, German, and Spanish. Table 1 shows
how many properties were lexicalized for each language. Overall, 574 di erent
properties were lexicalized, but note that not every property is covered for each
language. This is due to the fact that for some properties no sentences are found
3 http://dblexipedia.org/LexicalEntry_village_as_Noun_withPrep_in
4 http://dblexipedia.org/LexicalEntry_female_as_AdjectiveRestriction
5 http://dblexipedia.org/LexicalEntry_blue_as_AdjectiveRestriction
in the text corpus that match the prede ned lexicalization patterns. For
English, 567 properties were lexicalized, 224 by the corpus-based approach and 445
by the label-based approach. For German and Spanish, 145 and 50 properties,
respectively, were lexicalized using the corpus-based approach (the label-based
approach is currently not implemented for these languages).</p>
      <p>Table 2 shows the number of entries generated by the di erent approaches
for the di erent languages. Overall our dataset contains 11,998 entries. It is
important to mention that one lexical entry can lexicalize multiple properties.</p>
      <p>English
German
Spanish</p>
      <p>Corpus-based Label-based Total
224 445 567
145 n.a. 145
50 n.a. 50</p>
      <p>Attached to the entries we also publish meta-data, in particular about
provenance, specifying from which pattern an entry was created (for the corpus-based
approach), with which frequency, and with which con dence (for the label-based
approach). In the future we also intend to include example sentences for each
entry (for the corpus-based approach) and the set of corresponding features which
led to the entry (for the label-based approach).</p>
      <p>Moreover, the lexical entries are linked to dbnary6 and lemon UBY 7,
considering the canonical form and the part of speech as relevant information for
comparison.</p>
      <p>The dataset is published at http://dblexipedia.org, which was generated
by a modi ed version of YUZU 8. The website enables a user to browse through
the lexical entries and supports search over them. The whole dataset can also be
downloaded at http://dblexipedia.org/public/all.nt.gz (as N-Triples). In</p>
      <sec id="sec-3-1">
        <title>6 http://kaiko.getalp.org/about-dbnary/development/ 7 http://www.lemon-model.net/lexica/uby/ 8 https://github.com/jmccrae/yuzu</title>
        <p>addition to the dataset, the most current version of M-ATOLL is available at
http://dblexipedia.org/public/MATOLL.jar.</p>
        <p>
          We compared the published entries of the English corpus-based approach
with those of the manually created DBpedia lexicon [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Out of the 224
properties lexicalised by M-ATOLL, 84 were also lexicalised in the manually created
lexicon. We therefore evaluated on those 84 properties and found that 48 % of
the automatically constructed entries were consistent with the corresponding
manually created ones. This shows that on a larger number of properties, on
average half of the generated entries are good, whereas the other half has to be
manually corrected. For the example above with the property spouse, M-ATOLL
creates the three lexical entries mentioned in the introduction. Additionally we
nd for this property the expressions X is the widow of Y and X is reunited with
Y.
        </p>
        <p>The main limitation of the presented dataset is the small coverage of the
DBpedia 2014 ontology, containing around 2796 properties. However, for 1424
properties of this ontology, no data is available at the o cial DBpedia SPARQL
endpoint. Considering only properties with at least one received data item, our
very rst release of the lexicon covers already 42% of the properties.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>
        In this paper we presented the rst multilingual, automatically generated lexicon
for DBpedia, covering 574 properties from the DBpedia 2014 ontology. The
evaluation showed that on average half of the generated entries are correct. In order
to further improve the quality of the resulting lexica across languages, we will
work on the adaptation of the label-based approach to German and Spanish, and
in the future also Japanese. Moreover, we plan to evaluate to which extent an
increase in coverage can improve tasks like question answering (see, e.g. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). We
will also continue to publish updates of the dataset at http://dblexipedia.org,
o ering it to the community as a resource that can support natural
languagebased applications over the Semantic Web and that can serve as a hub for other
lexical resources on the linguistic linked data cloud.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>This work was supported by the Cluster of Excellence Cognitive Interaction
Technology CITEC (EXC 277) at Bielefeld University, which is funded by the
German Research Foundation (DFG).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Nadjet</given-names>
            <surname>Bouayad-Agha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gerard</given-names>
            <surname>Casamayor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Leo</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <article-title>Natural language generation in the context of the semantic web</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <volume>493</volume>
          {
          <fpage>513</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Sherzod</given-names>
            <surname>Hakimov</surname>
          </string-name>
          , Christina Unger,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Walter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>Applying semantic parsing to question answering over linked data: Addressing the lexical gap</article-title>
          . In Chris Biemann, Siegfried Handschuh, Andre Freitas, Farid Meziane, and Elisabeth Metais, editors,
          <source>Natural Language Processing and Information Systems</source>
          , volume
          <volume>9103</volume>
          of Lecture Notes in Computer Science, pages
          <volume>103</volume>
          {
          <fpage>109</fpage>
          . Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Jens</given-names>
            <surname>Lehmann</surname>
          </string-name>
          , Chris Bizer, Georgi Kobilarov, Soren Auer, Christian Becker, Richard Cyganiak, and
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann</surname>
          </string-name>
          .
          <article-title>DBpedia - a crystallization point for the web of data</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>7</volume>
          (
          <issue>3</issue>
          ):
          <volume>154</volume>
          {
          <fpage>165</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Vanessa</given-names>
            <surname>Lopez</surname>
          </string-name>
          , Victoria Uren, Marta Sabou, and
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Motta</surname>
          </string-name>
          .
          <article-title>Is question answering t for the semantic web?: A survey</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <volume>125</volume>
          {
          <fpage>155</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Montserrat</given-names>
            <surname>Marimon</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nuria</given-names>
            <surname>Bel</surname>
          </string-name>
          .
          <article-title>Dependency structure annotation in the IULA Spanish LSP Treebank</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>49</volume>
          (
          <issue>2</issue>
          ):
          <volume>433</volume>
          {
          <fpage>454</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>John</surname>
            <given-names>McCrae</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Dennis</given-names>
            <surname>Spohr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>Linking lexical resources and ontologies on the semantic web with lemon</article-title>
          .
          <source>In The Semantic Web: Research and Applications</source>
          , pages
          <volume>245</volume>
          {
          <fpage>259</fpage>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <article-title>An e cient algorithm for projective dependency parsing</article-title>
          .
          <source>In Proceedings of the 8th International Workshop on Parsing Technologies (IWPT</source>
          , pages
          <volume>149</volume>
          {
          <fpage>160</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Rico</given-names>
            <surname>Sennrich</surname>
          </string-name>
          .
          <article-title>The UZH system combination system for WMT 2011</article-title>
          .
          <source>In Proceedings of the Sixth Workshop on Statistical Machine Translation, WMT '11</source>
          , pages
          <fpage>166</fpage>
          {
          <fpage>170</fpage>
          ,
          <string-name>
            <surname>Stroudsburg</surname>
          </string-name>
          , PA, USA,
          <year>2011</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Christina</surname>
            <given-names>Unger</given-names>
          </string-name>
          ,
          <string-name>
            <surname>John McCrae</surname>
            ,
            <given-names>Sebastian</given-names>
          </string-name>
          <string-name>
            <surname>Walter</surname>
            , Sara Winter, and
            <given-names>Philipp</given-names>
          </string-name>
          <string-name>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>A lemon lexicon for DBpedia</article-title>
          .
          <source>In Proceedings of 1st International Workshop on NLP and DBpedia</source>
          , co
          <article-title>-located with the 12th International Semantic Web Conference (ISWC</article-title>
          <year>2013</year>
          ),
          <source>October 21-25</source>
          , Sydney, Australia,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sebastian</surname>
            <given-names>Walter</given-names>
          </string-name>
          , Christina Unger, and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>ATOLL { a framework for the automatic induction of ontology lexica</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <volume>94</volume>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>B</given-names>
          </string-name>
          (
          <volume>0</volume>
          ):
          <volume>148</volume>
          {
          <fpage>162</fpage>
          ,
          <year>2014</year>
          . Special Issue following the 18th
          <source>International Conference on Applications of Natural Language Processing to Information Systems (NLDB'13).</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Sebastian</surname>
            <given-names>Walter</given-names>
          </string-name>
          , Christina Unger, and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>M-ATOLL: a framework for the lexicalization of ontologies in multiple languages</article-title>
          .
          <source>In The Semantic Web { ISWC</source>
          <year>2014</year>
          , volume
          <volume>8796</volume>
          of Lecture Notes in Computer Science, pages
          <volume>472</volume>
          {
          <fpage>486</fpage>
          . Springer International Publishing,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Sebastian</surname>
            <given-names>Walter</given-names>
          </string-name>
          , Christina Unger, and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>Automatic acquisition of adjective lexicalizations of restriction classes: a machine learning approach</article-title>
          .
          <source>In Journal on Data Semantics</source>
          (to appear),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sebastian</surname>
            <given-names>Walter</given-names>
          </string-name>
          , Christina Unger, Philipp Cimiano, and
          <string-name>
            <given-names>Bettina</given-names>
            <surname>Lanser</surname>
          </string-name>
          .
          <article-title>Automatic acquisition of adjective lexicalizations of restriction classes</article-title>
          .
          <source>In Proceedings of 2st International Workshop on NLP and DBpedia</source>
          , co
          <article-title>-located with the 13th International Semantic Web Conference (ISWC</article-title>
          <year>2014</year>
          ),
          <source>October 19-23, Riva del Garda</source>
          , Italy,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>