<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Latest enhancements in the Spanish DBpedia?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Sanz-Lucio</string-name>
          <email>fsara.sanz.lucio@alumnos</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oussama Tahiri-Alaoui</string-name>
          <email>oussama.talaoui@alumnos</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>no Ri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ontology Engineering Group, Universidad Politecnica de Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Spanish DBpedia is a data source used initially to support the Spanish community. However, our logs show that the Spanish language goes beyond Spanish speakers and many non-Spanish speakers use the Spanish DBpedia on a daily basis. In the last months we have made two important enhancements to the Spanish DBpedia: (1) we publish a nonstandard dataset containing the type of resources that in the standard distribution have no type, and (2) we update automatically our data every week by using the DBpedia databus. In this way, we satisfy a frequent request made by companies and we foster the usage of the Spanish language, the second mother language by the number of speakers (after Chinese), and the second in scienti c papers (after English).</p>
      </abstract>
      <kwd-group>
        <kwd>Spanish DBpedia</kwd>
        <kwd>Resource type</kwd>
        <kwd>DBpedia data bus</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <sec id="sec-2-1">
        <title>The rising of the Spanish language</title>
        <p>
          The data published by the Cervantes Institute in its 2020 report [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] are
overwhelming: Spanish speakers have increased by 30% in the last decade, and the
number of foreigners who study it has grown by 60%. More than 585 million
people speak Spanish. Of these, almost 489 million are native Spanish speakers.
Furthermore, Spanish is the second mother tongue by number of speakers after
Mandarin Chinese, and the third language in the global count of users after
English and Mandarin Chinese. On the Internet, it is the third most used and is
the second language, behind English, publishing scienti c texts.
        </p>
        <p>
          The DBpedia project has long generated semantic information from English
Wikipedia. Since June 2011, the information generation process has extracted
information from Wikipedia in 111 of its languages, but only 18 languages have a
DBpedia chapter with a website. One of them is Spanish. The DBpedia
Internationalization Committee has assigned a website and a SPARQL [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] endpoint for
each of these languages1. In the case of Spanish (with website es.dbpedia.org),
the extraction process produces more than 100 million RDF triples from the
Spanish Wikipedia. All these triples are available on the SPARQL endpoint
es.dbpedia.org/sparql using Semantic Web [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and Linked Data [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
technologies.
1.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>The DBpedia datasets</title>
        <p>As we have mentioned previously, DBpedia extracts data from 111 di erent
language editions of Wikipedia. Then, for each language we have a knowledge
base (a \Knowledge Graph" in modern terminology, abbreviated as KG). The
largest DBpedia KG is extracted from the English edition of Wikipedia, with
around 400 million facts (triples) that describe 3.7 million resources (Wikipedia
entries). The DBpedia knowledge graphs that are extracted from the other 110
Wikipedia editions together consist of 1.46 billion facts and describe 10 million
additional resources. Therefore, two-thirds of the information in DBpedia comes
from non-English Wikipedias.</p>
        <p>
          From a technical perspective, the DBpedia project maps Wikipedia infoboxes [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
from 27 di erent language editions into the DBpedia ontology, a single shared
ontology consisting of 320 classes and 1,650 properties. The mappings are created
via a worldwide crowd sourcing e ort and enable knowledge from the di
erent Wikipedia editions to be combined. The DBpedia project publishes regular
releases of all DBpedia knowledge bases for download and provides SPARQL
query access to 18 out of the 111 language editions via a global network of local
DBpedia chapters.
        </p>
        <p>In addition to the regular releases, the project maintains a live knowledge
base which is updated whenever a page in Wikipedia changes. DBpedia sets
27 million links pointing to many external data sources (e.g. Wikidata, Yago,
Freebase) and thus enables data from these sources to be used together with
DBpedia data. Several hundred data sets on the Web publish RDF links pointing
to DBpedia and thus make DBpedia the central interlinking hub in the Linked
Open Data (LOD) cloud2.
1.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>The Spanish DBpedia datasets</title>
        <p>
          Around 40% of the resources (entries) in the Spanish Wikipedia are not pointed
(do not have links) by the English Wikipedia [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which means they are
\noncanonical" datasets, that is, 40% of the information stored in the Spanish
DBpedia is exclusively stored in the Spanish DBpedia and is not available in the
English DBpedia. This fact places the Spanish DBpedia as a valuable and
exclusive source of local information. Additionally, we have to remark that the
English DBpedia does not contain all the information stored in local DBpedias,
but only a minimum part comprising labels and abstracts.
1 See https://www.dbpedia.org/members/chapter-overview/
2 See https://lod-cloud.net
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>New contributions to the Spanish DBpedia</title>
      <p>The contributions made to the Spanish DBpedia during the last year are
described in the next sections.
2.1</p>
      <sec id="sec-3-1">
        <title>On new DBpedia types</title>
        <p>Each DBpedia resource usually has more than one type. For example, the
resource Cervantes has types Agent, Person and Artist, this is because the
DBpedia ontology de nes a well-known hierarchy of classes.</p>
        <p>
          A previous study [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] showed that a large number of DBpedia resources,
around 16%, do not have any type. DBpedia, in its 3.9 (English) version, has
more than four million resources, but only around 50% of them have a type
beyond level 1 (we consider Thing as level 0). Having correct types is important
when working with semantic information, because it allows for better data
queriability and discoverability. Thus, the more types we have and the more precise
they are, the better data quality.
        </p>
        <p>In the study mentioned previously, made by members of our research group, it
was presented a new technique that improved the overall quality of the English
DBpedia dataset by (1) providing type(s) to those resources lacking a type,
and (2) adding more specialized types to already typed resources. Following the
previous example about Cervantes, this means to infer that it also has the type
Writer, a more speci c type in the DBpedia hierarchy, although this fact was
not in the original dataset. These inferred types are stored in a new dataset so
that now it is made publicly available in Spanish and English.</p>
        <p>
          This method surpasses the so-called SDTypes dataset as shown in gure 1.
The small circle represents the number of types predicted by the SDType
approach (1.38M). The big circle represents our approach, which produces 56.7%
more types (2.15M). Our approach predicts most of the types predicted by the
SDTypes approach[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], speci cally 96.5% of them. Therefore, we conclude that
our approach predicts most of the types predicted by the SDTypes approach.
        </p>
        <p>App2 C5.0 - 156.7%
2,151,900
only SDtype - 3.5% → 48,152
common - 96.5% → 1,325,252
only App2 C5.0 - 60.2% → 826,760
SDtype - 100%
1,373,404
Fig. 1. Overlapping between the types predicted by our approach (light gray) and the
ones predicted by the SDTypes approach (dark gray).</p>
        <p>
          The contribution of the present work is to notify that we have overcome a
technical issue that limited the number of resources that could be analyzed. This
limitation restricted our study to old versions of DBpedia, in which we had less
than 2.4 million resources. The master thesis [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] of the rst author of this paper,
Sara, was able to solve a limitation of the C5.0 library (a popular multiclass
classi er). Now we can analyze any DBpedia and provide a new dataset to the
DBpedia community, but we are focused mainly on the English and Spanish
datasets. The publication mechanism is the DBpedia databus, mentioned in the
next section.
2.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>On DBpedia automatic update process</title>
        <p>In the past, the update process was based on a manual process in which we
had to download the Spanish datasets provided periodically by DBpedia. This
periodicity was the refresh rate of the DBpedia extraction process, and was in
the range of months.</p>
        <p>
          However, now we have an automated process and we get fresh data
every week. This is possible by using the DBpedia Databus [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], which aims at
providing an end-to-end pipeline easing auto-extraction, metadata-generation
and publishing of Linked Data Knowledge Graphs at scale. The Databus
platform provides two tools to connect consumers and producers, one for consumers
(https://databus.dbpedia.org) and the SPARQL API (https://databus.db
pedia.org/repo/sparql) serves as a user interface to con gure data set retrieval
and combination in catalogs and the other for providers, the Databus Maven
plugin (http://dev.dbpedia.org/Databus/Derive/Maven/Integration) which
enables systematic upload and release of datasets on the bus.
        </p>
        <p>
          The idea behind using Databus to get our data is the fact that it is able
to generate metadata about datasets and then upload this metadata, so that
anybody can query, download, derive, and build applications with this data
via the Databus. Since the integration of data is easy with the Databus, many
additional datasets have been integrated and loaded alongside DBpedia for the
world to query. The process to load datasets on the bus comprises these phases:
{ Acquisition: data is downloaded from the source and logged in.
{ Conversion: data is converted to N-Triples and cleaned (Syntax parsing,
datatype validation and SHACL).
{ Mapping: the vocabulary is mapped on the DBpedia Ontology and converted
(this was being done for Wikipedia's Infoboxes and Wikidata [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], but now
it is done for other datasets as well).
{ Linking: Links are mainly collected from the sources, cleaned and enriched.
{ IDying: All entities found are given a new Databus ID for tracking.
{ Clustering: IDs are merged onto clusters using one of the Databus IDs as
cluster representative.
{ Data Comparison: Each dataset is compared with all other datasets. We
have an algorithm that decides on the best value, but the main goal here is
transparency, i.e. to see which data value was chosen and how it compares
to the other sources.
{ A main knowledge graph fused from all the sources, i.e. a transparent
aggregate.
{ For each source, a local fused version called the \Databus Complement" is
being produced. This is a major feedback mechanism for all data providers,
where they can see what data they are missing, what data di ers in other
sources and what links are available for their IDs.
        </p>
        <p>{ The possibility to compare all data via a web service.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and future work</title>
      <p>
        We share with the community the enhancements made to the Spanish DBpedia
in the last year. This has been the biggest change since its creation in 2011.
The changes include a new user interface, automatic weekly updates, and the
creation of a high-quality dataset on resource's types. All of this is publicly
available (code, datasets and instructions) on the Internet [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngonga</surname>
            <given-names>Ngomo</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.C.</surname>
          </string-name>
          :
          <article-title>Introduction to linked data and its lifecycle on the web</article-title>
          .
          <source>In: Reasoning on the Web in the Big Data Era</source>
          . vol.
          <volume>6848</volume>
          , pp.
          <volume>1</volume>
          {
          <issue>75</issue>
          (01
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lassila</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The semantic web</article-title>
          .
          <source>Scienti c American</source>
          <volume>284</volume>
          (
          <issue>5</issue>
          ),
          <volume>34</volume>
          {
          <fpage>43</fpage>
          (
          <year>2001</year>
          ), http://www.jstor.org/stable/26059207
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Frey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obraczka</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>DBpedia FlexiFusion the best of wikipedia &gt; wikidata &gt; your data</article-title>
          .
          <source>In: ISWC 2019</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Garc</surname>
            a-Montero,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Spanish: a living language</article-title>
          .
          <source>Report 2020. Tech. rep., Instituto Cervantes</source>
          (
          <year>2020</year>
          ), https://cvc.cervantes.es/lengua/espanol lengua viva/pd f/espanol lengua viva
          <year>2020</year>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mihindukulasooriya</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rico</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a Castro,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An analysis of the quality issues of the properties available in the spanish dbpedia</article-title>
          .
          <source>In: AEPIA conference</source>
          . pp.
          <volume>198</volume>
          {
          <issue>209</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Paulheim</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Improving the quality of linked data using statistical distributions</article-title>
          .
          <source>IJSWIS</source>
          <volume>10</volume>
          (
          <issue>2</issue>
          ),
          <volume>63</volume>
          {
          <fpage>86</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Prud'hommeaux</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>SPARQL query language for RDF, W3C recommendation (</article-title>
          <year>2008</year>
          ), http://www.w3.org/TR/rdf-sparql-query/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rico</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santana-Perez</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pozo-Jimenez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Inferring types on large datasets applying ontology class hierarchy classi ers: The dbpedia case</article-title>
          .
          <source>In: EKAW 2018</source>
          . pp.
          <volume>322</volume>
          {
          <issue>337</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Sanz-Lucio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Deteccion de tipos en DBpedia.
          <source>Master's thesis</source>
          , Universidad Politecnica de Madrid (
          <year>2021</year>
          ), http://oa.upm.es/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tahiri-Alaoui</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>An approach to automatically update the Spanish DBpedia using DBpedia Databus</article-title>
          .
          <source>Master's thesis</source>
          , Universidad Politecnica de Madrid (
          <year>2020</year>
          ), http://oa.upm.es/63646/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Vrandecic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krotzsch, M.:
          <article-title>Wikidata: A free collaborative knowledgebase</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <issue>10</issue>
          ),
          <volume>78</volume>
          {
          <fpage>85</fpage>
          (
          <year>2014</year>
          ). https://doi.org/10.1145/2629489
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatically re ning the wikipedia infobox ontology</article-title>
          .
          <source>In: Proceeding of the 17th International Conference on World Wide Web</source>
          <year>2008</year>
          , WWW'08. pp.
          <volume>635</volume>
          {
          <issue>644</issue>
          (04
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>