<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Description of Knowledge Graphs Construction Automation: Status &amp; Challenges</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Chaves-Fraga</string-name>
          <email>david.chaves@upm.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasia Dimou</string-name>
          <email>anastasia.dimou@kuleuven.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Flanders Make - DTAI-FET</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Knowledge Graphs, Automation, Explainable AI, Declarative Rules</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KGCW'22: International Workshop on Knolwedge Graph Construction</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>KU Leuven, Department of Computer Science</institution>
          ,
          <addr-line>Sint-Katelijne-Waver</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Leuven.AI - KU Leuven institute for AI</institution>
          ,
          <addr-line>B-3000 Leuven</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidad Politécnica de Madrid</institution>
          ,
          <addr-line>Campus de Montegancedo, Boadilla del Monte</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>Nowadays, Knowledge Graphs (KG) are among the most powerful mechanisms to represent knowledge and integrate data from multiple domains. However, most of the available data sources are still described in heterogeneous data structures, schemes, and formats. The conversion of these sources into the desirable KG requires manual and time-consuming tasks, such as programming translation scripts, defining declarative mapping rules, etc. In this vision paper, we analyze the trends regarding the automation of KG construction but also the use of mapping languages for the same process, and align the two by analyzing their tasks and a few exemplary tools. Our aim is not to have a complete study but to investigate if there is potential in this direction and, if so, to discuss what challenges we need to address to guarantee the maintainability, explainability, and reproducibility of the KG construction.</p>
      </abstract>
      <kwd-group>
        <kwd>Challenges</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        A lot of works on knowledge graph (KG) construction are focused on defining mapping languages
to declaratively describe the transformation process, and on optimizing the execution of such
declarative rules. The mapping languages rely on either dedicated syntaxes, such as the family
of languages around the W3C recommended R2RML1 (e.g., RML [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or R2RML-F [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]), or on
re-purposing existing specifications, such as query languages like the W3C recommended
SPARQL2 (e.g., SPARQL-Generate [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or SPARQL-Anything [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), or constraints languages like
ShEx3 (e.g., ShExML [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]).
      </p>
      <p>
        Despite the plethora of mapping languages and the increasing number of optimizations
for the execution of the declarative rules, these rules are still defined through a manual and
time-consuming process, afecting negatively their adoption. Diferent solutions were proposed
to automate the definition of mapping rules that describe how a KG should be constructed.
CEUR
Workshop
Proceedings
On the one hand, MIRROR [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], D2RQ [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Ontop [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] follow a similar approach, extracting
from the RDB schema a target ontology and the mapping correspondences. On the other hand,
AutoMap4OBDA [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and BootOX [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] consider an input ontology and generate actual R2RML
mappings from the RDB. However, these solutions are focused on declarative solutions only
for relational databases, while recent solutions investigate non-declarative automation of KG
construction.
      </p>
      <p>
        Beyond relational databases, the recent SemTab challenge4 presents a set of tabular datasets [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
with the aim of matching them automatically to external KGs, such as DBpedia and Wikidata.
The proposed solutions [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13, 14, 15</xref>
        ] address the problem using diferent techniques, such as
heuristic rules, fuzzy searching over the KGs or knowledge graph embeddings. Although their
ifnal objective is the same (to obtain high precision and recall results) and they perform similar
procedures, each solution implements its own workflow and addresses each proposed task
by SemTab in diferent ways. Hence, making a fair and fine-grained comparison among the
diferent solutions to understand how they obtain the actual results is not an easy task.
      </p>
      <p>
        In this vision paper, we align tasks followed by solutions for the automation of the semantic
table annotation with concepts of existing declarative solutions. We indicatively select and
analyze a few tools for the automation of KG construction and identify common steps. We
discuss whether they can be declaratively described relying on existing mapping languages, and
what the challenges are to proceed in this direction. We consider the RDF Mapping Language
(RML) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as a high-level and general representation to describe schema transformations and its
extension, the Function Ontology (FnO) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to describe data transformations.
      </p>
      <p>Our objective is not to present a complete study but to investigate if there is potential in this
direction. By describing the steps followed by diferent solutions in a more fine-grained and
standard manner, we make the steps comparable, and we can better discuss what challenges we
need to address to guarantee the maintainability, explainability, and reproducibility of the KG
construction, as well as to ensure the provenance of each performed task.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task alignment with mapping languages</title>
      <p>We analyze the diferent steps of the SemTab challenge, inspect the relationship between the
SemTab challenge tasks, and align them with concepts from the declarative construction of
RDF graphs (Figure 1). To achieve this, we include the relationship between each of the tasks
and their potential declarations within a mapping language. We consider the RML mapping
language because it is commonly used and the authors are more familiar with it, but we are
confident that the other mapping languages could express the same concepts. Before we proceed
with the alignment, we give a small introduction on the SemTab challenge and RML:
SemTab challenge The SemTab challenge consists of three tasks: (i) cell to KG entity
matching (CEA), which matches cells to individuals; (ii) column to KG class matching (CTA),
which matches cells to classes; and (iii) column pair to KG property matching (CPA), which
captures the relationships between pairs of columns.</p>
      <sec id="sec-2-1">
        <title>4https://www.cs.ox.ac.uk/isg/challenges/sem-tab/</title>
        <p>CEA</p>
        <p>Col0</p>
        <p>Col1</p>
        <p>Col2
Union Depot</p>
        <p>
          Spier &amp; Rohns
RML The RDF mapping language (RML), a superset of the W3C recommended R2RML,
expresses schema transformations from heterogeneous data to RDF. An RML mapping contains
one or more Triple Maps which on their own turn contain a Subject Map to generate the subjects
of the RDF triples, and zero or more Predicate Object Maps with pairs of Predicate and Object
Maps to generate the predicates and the objects respectively for each incoming data record. RML
was aligned with the Function Ontology (FnO) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] to describe the data transformations which
are required to construct the desired RDF graph, ensuring that the functions are independent
from any implementation.
        </p>
        <p>We analyze how the diferent tasks of the challenge contribute in constructing a part of an
RDF triple, and we align these tasks with the corresponding concepts of the RML mapping
language that construct the same part of an RDF triple.</p>
        <p>Cell-Entity Annotation (CEA): This task identifies the URI of an entity from a cell. In the
target RDF graph, this is the subject or the object of the RDF triple. In Fig. 1, the Col0 values
are used to obtain the subjects of the triples while the Col3 values generate the objects (both
green colored in the RDF extract of Fig. 1). If a declarative approach is considered to generate
these triples, for example in RML, the rr:subjectMap property is used (line 5 of RML doc in
Fig. 1), which declares how the subjects of the triples are generated and the rr:objectMap (line
8 of RML doc in Fig. 1), when the expected objects are in the form of URIs.</p>
        <p>Column-Type Annotation (CTA): This task predicts the common class of a set of items
given a column from the table. SemTab assumes that a table only generates one kind of entity
(i.e. the first column is used for CTA). In Figure 1, we can observe that the URIs retrieved
using Col0 are considered for obtaining the corresponding shared concept (i.e., restaurant)
(red colored in the RDF extract of Fig. 1). Declaring the class in RML can be done through the
shortcut rr:class property within the rr:SubjectMap or using a rr:predicateObjectMap
with a rdf:type fixed predicate (line 7 of RML doc in Fig. 1).</p>
        <p>Columns-Property Annotation (CPA): This task aims to predict the property that relates
the CTA column (subjects) to the rest of the columns. Fig. 1 shows a CPA task that relates Col0
with Col3 through the property architectural style (wdt:P149, yellow colored in the RDF
extract). In RML, the predicates of the triples are declared using the rr:predicateMap property
(line 8 of RML doc in Fig. 1), and unlike typical mapping rules, where it is usually assumed that
predicates are constants (as they are declared in the input ontology), the predicates depend on
the data, hence they are dynamically defined.</p>
        <p>Based on the aforementioned analysis, we conclude that the tasks performed to automate the
KG construction can be aligned with concepts from declarative mapping languages. The CEA
task is aligned with the RDF term construction for the subject or the object of the RDF triple,
the CTA task assigns the class and the CPA task aligns with the Predicate and Object Map.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Comparing semantic tabular matching systems</title>
      <p>In this section, we analyze in detail the steps performed by some of the tools proposed for
solving the SemTab challenge. The comparative analysis among the three selected engines
(summarized in Table 1), is not meant to be exhaustive. We aim to identify if there are common
steps and functions that the engines perform to accomplish the challenge’s tasks and ultimately
if it is possible and desired to declaratively describe them with mapping languages.</p>
      <sec id="sec-3-1">
        <title>3.1. Selected Systems</title>
        <p>
          We indicatively selected the systems that: (i) obtained good results in the SemTab 2021
challenge5; and (ii) have the source code openly available. Therefore, we included in this comparison
JenTab [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], MTab [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and MantisTable V [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The use of diferent terminologies for describing
similar tasks (e.g., majority vote in Mantis V is referred as frequency) and the complexity of the
proposed workflows, where the results from one of the task influence the others in a iterative
way, create dificulties to compare the approaches and reproduce their results.
        </p>
        <p>JenTab6 participated in SemTab 2020 and 2021, and it was always positioned among the top
ifve solutions for most rounds. It follows a heuristic-based approach proposing the CFS (Create,
Filter, Select) approach for all tasks and with diferent configurations and workflows.</p>
        <p>MTab7 participated in all SemTab editions, winning the first prize in 2019 and 2020. Apart
from the support of multilingual datasets, MTab implements several approaches for performing
the entity search (i.e. CEA): keyword search, fuzzy search, and aggregation search8.</p>
        <sec id="sec-3-1-1">
          <title>5https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2021 6https://github.com/fusion-jena/JenTab 7https://github.com/phucty/mtab_tool 8https://mtab.app/mtabes/docs</title>
          <p>
            MantisTable V9 is an extended and improved version of MantisTable [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]. Similarly to
JenTab, MantisTable has also participated in SemTab 2020 and 2021 editions. It implements a
set of heuristic rules (similar as JenTab) and complex string similarity functions for the entity
recognition task (like MTab). Additionally, it provides a general and eficient tool (LamAPI) to
fetch the necessary data for all SemTab tasks, independently of the target KG.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Observations</title>
        <p>The systems we inspected follow the same steps: they perform a preprocessing step, and setup
lookup and datatype prediction services. Then the CEA task is performed followed by the CTA
and CPA tasks which depend on the CEA task. Given that the systems follow the same steps,
we could map the three main tasks (CEA, CPA, CTA) to the Create-Filter-Select (CFS) procedure
proposed by JenTab (see Table 1).</p>
        <p>We observe similarities in most tasks among the engines. The subtasks performed in the
preprocessing step, are very similar in the three engines. Preprocessing tasks include several
functions, such as fixing encoding issues, removing HTML tags or special characters, and
detecting missing white spaces (see Table 1), and they usually delegate them to third-party
libraries (e.g., ftfy 10). We observe similar tasks are performed when declarative solutions are
used for cleaning and preparing the data. These preprocessing tasks are described with FnO
in the case of RML and executed either together with the schema transformations or as a
preprocessing task too.</p>
        <p>The same occurs for the datatype prediction, where regular expressions are often used to
detect if cell values are entities or literals, and what type of literals (string, date, or numbers). In
the case of declarative solutions, this datatype inspection task is performed manually. However,
adjusting the datatype is possible by relying on functions for data transformations.</p>
        <p>
          Most of them also incorporate a lookup step to retrieve the necessary data from the KGs (e.g.,
using SPARQL queries), including similarity functions or fuzzy search. The search engine for
the KG lookups in JenTab and Mantis V is ElasticSearch, although the former implements the
Jaro Winkler distance [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] while the latter embeds it in a more eficient engine and exploits
its query capabilities. Lookups were also incorporated in the case of declarative solutions [20],
where lookup services retrieve a URI to identify an entity instead of assigning a new one.
        </p>
        <p>
          As far as the actual tasks are concerned, each engine performs its own approach for the CEA,
CTA, and CPA tasks, although we also find some similarities. The most important ones that
are implemented in the three engines are: (i) the Levenshtein distance [21] to filter candidates
and (ii) the majority vote (called frequency in Mantis V) to select the final annotations. We
believe that the use of declarative approaches, such as the Function Ontology [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] for describing
common functions (e.g., Levenshtein), could make the solutions more comparable. It would
also be clearer if they perform the same function and more explainable, as current solutions
for the automation of KG construction act like blackboxes: neither their implementations are
open sourced nor the declarative descriptions of what they execute are available. Providing at
least declarative descriptions of the tasks performed would enhance the transparency of these
solutions.
        </p>
        <sec id="sec-3-2-1">
          <title>9https://bitbucket.org/disco_unimib/mantistable-v/ 10https://pypi.org/project/ftfy/</title>
          <p>Fix encoding
Special characters
Preprocessing Restore
missing spaces
Remove
HTML tags
Remove
non-cell-values
Datatype</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Challenges for a declarative automation of KG Construction</title>
      <p>We identify a set of challenges to be addressed to declaratively describe solutions for automatic
KG construction. These challenges can be divided into two categories: technical and conceptual.</p>
      <p>
        On the technical side, there is a major diference between the solutions for the automation of
KG construction and the execution of declarative KG construction solutions: The solutions for
automatic KG construction rely on iterative processes that continuously refine and improves a
task, while the diferent tasks influence each other. To the contrary, the declarative KG
construction is a linear process that is executed only once. Not all declarative rules are executed linearly,
solutions that restructure [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] or parallelize them [22, 23] are increasingly encountered, but no
iterative solutions were proposed so far. Thus, if the solutions for automatic KG construction
are declaratively described, their iterative execution needs to be described as well. How do we
do that with the mapping languages? When is it meaningful?
      </p>
      <p>Besides the overall execution process, the iteration patterns are diferent. The solutions for
automatic KG construction are applied to all directions, both per column and per row, and even
combined. To the contrary, the declarative solutions are applied only per row, and the mapping
languages are designed under this assumption. Should the mapping languages be extended to
support more iteration patterns? If so, would the rml:iteration for RML and the relevant
constructs in the other mapping languages be suficient or more adjustments are required?</p>
      <p>The solutions for automatic KG construction rely on interrelated tasks which may produce
intermediate representations, e.g., probabilistic methods, and their results impact the rest tasks.
The declarative KG construction solutions then need to deal with dynamic and recursive
steps (e.g., intermediate representation of the input data sources and mapping rules, multiple
function execution, etc.) that can negatively impact the generation process. Hence, declaratively
describing is a challenge. Should the mapping languages be further extended then?</p>
      <p>On the conceptual side, there are two major diferences with respect to the training data and
target KGs. In most real projects, the input data and sometimes the target ontology are only
provided, but there is neither similar data to train the solutions nor existing KGs to target that
can be used to find entities or to predict the relationships. In the past, alternative approaches for
KG construction were discussed depending on what is available where the process starts (e.g.,
data, ontologies, target KGs), but it is not investigates neither how these editing approaches
afect the KG generation nor its automation. How are the automated solutions proposed for the
SemTab but not only afected by the lack of training data and target ontologies and KGs. How
does this afect their declarative representation?</p>
      <p>While relying on ontology matching techniques between existing KGs (e.g., DBPedia,
Wikidata) and the target ontology or exploiting NLP approaches between ontology and input sources
documentation could be a solution for the latter, would it be realistic given that most ontologies
are not aligned and not all of them provide documentation?</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Work</title>
      <p>In this paper, we analyze the KG construction solutions and compare the automatic with the
declarative. While the tasks can be aligned with respect to what they achieve, their execution is
fundamentally diferent and a direct alignment is not feasible.</p>
      <p>Automatic solutions for KG construction are required to facilitate the adoption of KGs, but
there are also merits when the automation tasks are declaratively described, with respect to
maintenability, sustainability, and reproducibility. However, directly aligning the automatic
solutions with the declarative solutions might be technically and conceptually challenging
considering their diferent execution and iteration patterns. Extending the existing mapping
languages would be a solution, but it would also require to address the identified challenges
and not only. Would such extensions be feasible and desired or would they lead them beyond
their purpose? Although, mapping languages are not the only approach to have declarative
descriptions. Declarative descriptions of workflows emerge as well. Would that be a more viable
solution? If so, would the automatic and declarative solutions keep on growing in diferent
directions? These are questions that would be nice to reflect and discuss during the workshop.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>David Chaves-Fraga is supported by the Spanish Minister of Universities (Ministerio de
Universidades) and by the NextGenerationEU funds through the Margarita Salas postdoctoral fellowship.
Both authors are supported by Flanders Make.
[20] S. Jozashoori, A. Sakor, E. Iglesias, M.-E. Vidal, EABlock: A Declarative Entity Alignment
Block for Knowledge Graph Creation Pipelines, in: Proceedings of the 37th ACM/SIGAPP
Symposium On Applied Computing, 2022.
[21] V. I. Levenshtein, et al., Binary codes capable of correcting deletions, insertions, and
reversals, in: Soviet physics doklady, volume 10, Soviet Union, 1966, pp. 707–710.
[22] G. Haesendonck, W. Maroy, P. Heyvaert, R. Verborgh, A. Dimou, Parallel RDF generation
from heterogeneous big data, in: Proceedings of the International Workshop on Semantic
Big Data, 2019, pp. 1–6.
[23] J. Arenas-Guerrero, D. Chaves-Fraga, J. Toledo, M. S. Pérez, O. Corcho, Morph-kgc: Scalable
knowledge graph materialization with mapping partitions, Semantic Web (2022).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dimou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Vander</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Colpaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Verborgh</surname>
          </string-name>
          , E. Mannens, R. Van de Walle,
          <article-title>RML: a generic language for integrated RDF mappings of heterogeneous data</article-title>
          ,
          <source>in: Ldow</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Debruyne</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>O'Sullivan, R2RML-F: towards sharing and executing domain logic in R2RML mappings</article-title>
          , in: LDOW@ WWW,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lefrançois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zimmermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bakerally</surname>
          </string-name>
          ,
          <article-title>A SPARQL extension for generating RDF from heterogeneous formats</article-title>
          ,
          <source>in: European Semantic Web Conference</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Daga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Asprino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulholland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <article-title>Facade-X: an opinionated approach to SPARQL anything</article-title>
          ,
          <source>Studies on the Semantic Web</source>
          <volume>53</volume>
          (
          <year>2021</year>
          )
          <fpage>58</fpage>
          -
          <lpage>73</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Iglesias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jozashoori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chaves-Fraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Collarana</surname>
          </string-name>
          , M.-E. Vidal,
          <article-title>SDM-RDFizer: An RML Interpreter for the Eficient Creation of RDF Knowledge Graphs</article-title>
          ,
          <source>in: Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>3039</fpage>
          -
          <lpage>3046</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jozashoori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chaves-Fraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Iglesias</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-E. Vidal</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Corcho</surname>
          </string-name>
          , Funmap:
          <article-title>Eficient execution of functional mappings for knowledge graph creation</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2020</year>
          , pp.
          <fpage>276</fpage>
          -
          <lpage>293</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L. F.</given-names>
            <surname>d. Medeiros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Priyatna</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Corcho,</surname>
          </string-name>
          <article-title>MIRROR: Automatic R2RML mapping generation from relational databases</article-title>
          , in: International Conference on Web Engineering, Springer,
          <year>2015</year>
          , pp.
          <fpage>326</fpage>
          -
          <lpage>343</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Seaborne</surname>
          </string-name>
          ,
          <article-title>D2RQ-treating non-RDF databases as virtual RDF graphs</article-title>
          ,
          <source>in: Proceedings of the 3rd international semantic web conference (ISWC2004)</source>
          , volume
          <volume>2004</volume>
          , Springer Hiroshima,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Calvanese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Cogrel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Komla-Ebri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kontchakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lanti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rezk</surname>
          </string-name>
          , M. RodriguezMuro, G. Xiao,
          <article-title>Ontop: Answering SPARQL queries over relational databases</article-title>
          ,
          <source>Semantic Web</source>
          <volume>8</volume>
          (
          <year>2017</year>
          )
          <fpage>471</fpage>
          -
          <lpage>487</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Á. Sicilia</surname>
          </string-name>
          , G. Nemirovski,
          <article-title>AutoMap4OBDA: Automated generation of R2RML mappings for OBDA</article-title>
          ,
          <source>in: European Knowledge Acquisition Workshop</source>
          , Springer,
          <year>2016</year>
          , pp.
          <fpage>577</fpage>
          -
          <lpage>592</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kharlamov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zheleznyakov</surname>
          </string-name>
          , I. Horrocks,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pinkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Skjaeveland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Thorstensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mora</surname>
          </string-name>
          ,
          <article-title>Bootox: Bootstrapping OWL 2 ontologies and R2RML mappings from relational databases</article-title>
          , in: International Semantic Web
          <string-name>
            <surname>Conference (P&amp;D)</surname>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hassanzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Srinivas</surname>
          </string-name>
          ,
          <year>Semtab 2019</year>
          :
          <article-title>Resources to benchmark tabular data to knowledge graph matching systems</article-title>
          ,
          <source>in: European Semantic Web Conference</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>514</fpage>
          -
          <lpage>530</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , I. Yamada,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kertkeidkachorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ichise</surname>
          </string-name>
          , H. Takeda,
          <string-name>
            <surname>SemTab</surname>
          </string-name>
          <year>2021</year>
          :
          <article-title>Tabular Data Annotation with MTab Tool</article-title>
          , SemTab@ ISWC (
          <year>2021</year>
          )
          <fpage>92</fpage>
          -
          <lpage>101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Abdelmageed</surname>
          </string-name>
          , S. Schindler, JenTab Meets SemTab 2021'
          <article-title>s New Challenges</article-title>
          , in: SemTab@ ISWC,
          <year>2021</year>
          , pp.
          <fpage>42</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Huynh</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Deuzé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <article-title>DAGOBAH: Table and Graph Contexts For Eficient Semantic Annotation Of Tabular Data</article-title>
          , in: SemTab@ ISWC,
          <year>2021</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>B. De Meester</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Seymoens</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Dimou</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Verborgh</surname>
          </string-name>
          ,
          <article-title>Implementation-independent function reuse</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>110</volume>
          (
          <year>2020</year>
          )
          <fpage>946</fpage>
          -
          <lpage>959</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          , M. Cremaschi, MantisTable V:
          <article-title>a novel and eficient approach to Semantic Table Interpretation, SemTab@ ISWC (</article-title>
          <year>2021</year>
          )
          <fpage>79</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Spahiu</surname>
          </string-name>
          ,
          <article-title>A fully automated approach to a complete semantic table interpretation</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>112</volume>
          (
          <year>2020</year>
          )
          <fpage>478</fpage>
          -
          <lpage>500</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>W. E.</given-names>
            <surname>Winkler</surname>
          </string-name>
          ,
          <article-title>String comparator metrics and enhanced decision rules in the Fellegi-Sunter model of record linkage (</article-title>
          <year>1990</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>