<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MTab4Wikidata at SemTab 2020: Tabular Data Annotation with Wikidata</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Phuc Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ikuya Yamada</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natthawut Kertkeidkachorn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryutaro Ichise</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hideaki Takeda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Advanced Industrial Science and Technology</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute of Informatics</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Studio Ousia</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper introduces an automatic semantic annotation system, namely MTab4Wikidata, for the three semantic annotation tasks, i.e., Cell-Entity Annotation (CEA), Column-Type Annotation (CTA), Column Relation-Property Annotation (CPA), of Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab 2020). In particular, we introduce (1) a novel fuzzy entity search to address misspelling table cells, (2) a fuzzy statement search to deal with ambiguous cells, (3) a statement enrichment module to address the Wikidata shifting issue, (4) an e cient and e ective post-processing for the matching tasks. Our system achieves impressive empirical performance for the three annotation tasks and wins the rst prize at SemTab 2020. MTab4Wikidata is ranked 1st in the two tasks of CEA and CPA, and 2nd rank in the CTA task on the round 1, 2, 3 datasets and 1st rank on the round 4 dataset and the Tough Tables (2T) dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Thanks to the Open Data movements, many tabular resources have been made
available on the Web and data portals. However, it is di cult to use such tabular
data because of missing or incomplete metadata, heterogeneous table schema,
table cell ambiguity, or misspelling. A promising solution to improve these
tabular data's usability is to generate semantic annotations for table elements using
knowledge graph concepts. As a result, such annotations could be useful for other
downstream tasks such as data management and knowledge discovery.</p>
      <p>CEA</p>
      <p>CTA</p>
      <p>CPA</p>
      <p>Cell
Wikidata Item
(Entity)</p>
      <p>Column
Wikidata Item
(Type)</p>
      <p>Relation
Wikidata
Property
Wikidata Knowledge Graph: Given Wikidata as a knowledge graph of
G = (Q; P; S), where Q and P are the set of Wikidata items, Wikidata
properties, respectively. A Wikidata item e (e 2 Q) is represented with a unique
identi er as Q###; for example, Q2 is the \Earth" item. The P### unique
identi er represents a Wikidata property p (p 2 P ); for example, P1082 is for the
\population" property. S is a set of statements represented as a triple format of
Subject - Predicate - Object. These statements describe Wikidata item
information such as literal information or connect Wikidata items1. For example, the Q5
(Earth) item could have a literal statement as Q5 (Earth) - P1082 (population)
7,655,957,369; and a item statement as Q5 (Earth) - P276 (location) - Q7879772
(inner Solar System).</p>
      <p>Tabular Data: Let a table T be a two-dimensional tabular structure consisting
of an ordered set of N rows and M columns. ni is a row of a table (i = 1:::N ),
mj is a column of the table (j = 1:::M ). The intersection between a row ni and
a column mj is ci;j is a value of the cell Ti;j .</p>
      <p>Annotation Tasks: We formalize the tabular data to Wikidata annotations as
the three matching tasks as follows.</p>
      <p>{ Cell-Entity Annotation (CEA): Matching a table cell ci;j into a Wikidata
item (entity) e.
{ Column-Type Annotation (CTA): Matching a column mj into a Wikidata
item (type) e.
{ Column Relation-Property Annotation (CPA): Matching the relation
between two columns mj1 and mj2 (j1; j2 2 [1; M ]; j1 6= j2) to a relation p.
ci;j</p>
      <p>CEA</p>
      <p>! Q
mj</p>
      <p>CTA</p>
      <p>! Q:
rmj1 ;mj2</p>
      <p>CPA
! P
(1)
(2)
(3)
1 Wikidata has about 89.75 million items, 7,972 properties, and 1.14 billion statements
in October 2020.
SemTab 2020 Challenges SemTab 2020 tasks are challenging for many
reasons as follows.</p>
      <p>Tabular data di culties:
{ Tabular data does not have metadata to describe the semantic meaning of
table elements.
{ Table headers are ambiguous; for example, the most popular table
headers are \col0", \col1", and \col2"; therefore, it is hard to understand table
schema if we only rely on the headers.
{ Table cells are ambiguous, contain many spelling errors and abbreviations.</p>
      <p>As a result, directly performing entity search with cell values (e.g., \Tokyo")
using standard APIs (e.g., Wikidata API search2 or Wikidata query3) could
return whereas too many relevant entities or no relevant entities in the
responding list. For example, searching with Wikidata API for an ambiguous
cell value of \Tokyo" could get 15,780 relevant entities, while there are no
relevant entities for a misspelling cell of \6C 124133+40580".</p>
      <p>Knowledge graph di culties:
{ Noisy Schema: Wikidata is a data-oriented knowledge graph. It has many
entity types annotated by humans; therefore, these types are noisy and not
consistent for many Wikidata items. As a result, Wikidata schema
standardization is still an open challenge.
{ Wikidata Shifting: Wikidata is a fast-evolving knowledge graph, with around
13 million edited Wikidata items per month4. According to our analysis on
the Round 1 data (Wikidata from March to August 2020), 56.09% of the
matching items had changed their information statements (possibly a ect
the CEA and CPA tasks), and 4.49% changes are related to Wikidata item
types (possibly e ect on the CTA task).</p>
      <p>
        Related work SemTab 2019 is the last year's challenge on tabular data to
DBpedia matching [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This year's challenge is tabular data to Wikidata matching; it
has a new set of di culties such as larger-scale knowledge graph setting,
knowledge graph data shifting, and noisy schema structure of Wikidata. Additionally,
this year's challenge also has a more challenging dataset (the tough tables [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]),
which is manually curated, o ering realistic issues than the last challenge.
      </p>
      <p>
        The last winner of SemTab 2019 is MTab system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] based on an
aggregation of multiple cross-lingual lookup services and probabilistic graphical model.
2 Wikidata search: https://www.wikidata.org/w/api.php
3 Wikidata query: https://query.wikidata.org/
4
https://stats.wikimedia.org/#/wikidata.org/content/edited
      </p>
      <p>pages/normal|line|2-year| total|monthly</p>
      <p>
        CSV2KG (IDLab) also uses multiple lookup services to improve matching
performance [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Tabular ISI implements the lookup part with Wikidata API and
Elastic Search on DBpedia labels, and aliases [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. ADOG [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] system also uses
Elastic Search to index knowledge graph. LOD4ALL rst checks whereas there is
available entity which has a similar label with table cell using ASK SPARQL, else
perform DBpedia entity search [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. DAGOBAH system performs entity linking
with a lookup on Wikidata and DBpedia; the authors also used Wikidata entity
embedding to estimate the entity type candidates [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Mantis Table provides a
Web interface and API for tabular data matching [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>Assumption</title>
        <p>This section describes MTab4Wikidata's assumptions (Section 2.1) and the
overall framework in Section 2.2.</p>
        <p>Assumption 1 MTab4Wikidata is built on a closed-world assumption.
Assumption 2 Input tables are in the vertical relation type.</p>
        <p>Assumption 3 The table cells in a column have the same entity types and data
types.</p>
        <p>Assumption 4 The rst row of the table (n1) is the table header. The rst cell
of column is header of this column, c1;j 2 mj .</p>
        <p>Assumption 5 The rst column of the table (m1) is the core attribute where
a cell value in this column could be matched into a Wikidata item, and other
remaining cells in the same row could be matched into the item statements.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Framework</title>
        <p>
          MTab4Wikidata is an automatic annotation system that could address the three
annotation tasks of CEA, CTA, and CPA. Unlike our previous system MTab
(DBpedia matching) built on the joint probability distribution of table element
signals[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], this work focus on improving the entity search capability as follows:
{ Wikidata shifting: Since all history revisions (entity labels, statements) are
available at Wikidata item history revision, and we use those revisions to
enrich the Wikidata dump; to this end, enhance matching performance.
{ Table cell noisiness and misspelling: Due to a high level of spelling mistakes
in table cells, we build a novel fuzzy entity lookup using edit distances that
could handle fuzzy queries.
{ Table cell ambiguity: We also build a novel fuzzy statement search that could
nd relevant statements from two cells taken from a table. This searcher
is built based on the assumption that there is an existing logical relation
occurring in the knowledge graph between the table's two cells. The searcher
helps reduce the number of relevant candidates for ambiguous queries.
        </p>
        <p>MTab4Wikidata contains three steps pipeline as Figure 2. Step 0 is to prepare
Wikidata resources with the entity dump and Wikidata history revisions. We
use these resources to create the two indexes of fuzzy entity search and fuzzy
statement search. In Step 1, we perform table cells lookup to nd relevant entity
candidates for each table cell (using the fuzzy entity search) and the two cells
in the same row (using the fuzzy statement search). In Step 3, we perform a
post processing with a value matching module, then majority voting to select
the CEA, CPA, and CTA annotations. The details of each step are described in
Section 2.3 (Step 0), Section 2.4 (Step 1), and Section 2.5 (Step 2).</p>
        <sec id="sec-2-2-1">
          <title>Input: Table</title>
          <p>Step 0
Step 1</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Resources</title>
          <p>Entity dump
History revisions</p>
          <p>Indexing
 Lookup
Fuzzy Entity Search
Fuzzy Statement Search
Step 2</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Post Processing</title>
          <p>Value Matching
Property
Voting
Entity
Voting
Type
Voting
CPA
CEA
CTA</p>
        </sec>
        <sec id="sec-2-2-4">
          <title>Annotation Results</title>
          <p>In this module, we perform the following operations:
{ Extracting Wikidata items and these statements: We extract this
information from the Wikidata 29 June 2020 entity dump. Note that this entity
dump is not available on the Wikidata server now, since they only keep the
two months' latest dumps.
{ Getting history revision of Wikidata items: We crawl Wikidata item
history revisions as a statement form (Subject - Predicate - Object) from 1
March to 29 June 20205. We use those revisions to enrich the Wikidata
5 Revision history: https://www.mediawiki.org/wiki/API:Revisions.
dump statements. Since we do not know when the challenge's tabular data
is generated, we only use the adding operation for all history statements of
Wikidata items. In other scenarios, where we know the exact time point of
Wikidata, we can reconstruct Wikidata items at this time point by using
adding, removing, modifying operations.
{ Building an entity index: We build this index using a hash table with the
entity dump data using multilingual labels, aliases, and identi ers (the
identi ers to other sources). In total, the fuzzy entity search index about 150
million entity names for 87 million items of Wikidata.
{ Building a statement index: we use a sparse matrix to build an index for
Wikidata statements. To be simple, we only encode the item statements
(item - property - item) and ignore the statements' property information. In
the general case, we can also index the literal statements and properties into
the sparse matrix. Overall, we index around 500 million statements.
2.4</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Step 1: Lookup</title>
        <p>In this step, we perform entity candidate generation for each target cells
using the fuzzy entity search, and two target cells in the same rows using the
fuzzy statement search. Figure 3 depicts the semantic annotations of the table
of 00KKIZPQ in the round 3 data. We use this table as an example to describe
the two following searching module.</p>
        <p>col0
V*!AY Psc
col1
-4.593
col2
Pisces
col3
Fig. 3. Semantic Annotations of the table of 00KKIZPQ in the round 3 data
A cell search In this searcher, we perform a fuzzy entity search on the entity
index. Given a table, this searcher returns a ranking list of relevant entities ranked
with an edit distance, speci cally Levenshtein distance. We rst start with two
edits; if the searcher does not have answers, the number of edits is gradually
increased to a maximum of six edits. For example: in Table 3, we have four queries
as two misspelling queries: \V*!AY Psc", and \SDSS J153509.57+360054.5".
This fuzzy search returns the correct answers for the two misspelling queries
as Q85702771 (V* AY Psc) and Q81115852 (SDSS J152509.57+360054.5). This
searcher achieves coverage of 99.89% on average when performing entity lookup
for table cells on SemTab 2020 datasets.</p>
        <p>
          Two cells search We design this searcher to handle the ambiguity of table
cells. We assume a logical relation equivalent to a statement between two cells
of a table row. Therefore, we perform the two cells' search on the target cells in
the same rows. The rst cell is in the core attribute column, and the second cell
is in the remaining columns. We ignore the rows do not have enough two target
cells. This idea is inspired by our novel idea of entity-entity column matching in
MTab [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Unlike our original idea in MTab that performs the two cells matching
in the post-processing step, this work adapts this idea into the lookup step and
makes the post-processing more e cient. Using statement search in the lookup
step is also e cient because of the sparse matrix searching. This idea also works
e ectively, and robust on DBpedia [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>For example: in Table 3, we have two queries: the two cells of row 1 \V*!AY
Psc", \Pisces", and row 2 \SDSS J153509.57+360054.5", \Bootes". We rst
search one cell then check all combinations between the two respond lists,
whereas the statements are available. In the end, we only keep entity
candidates at two cells that have a statement available. Regarding the cells in the
core attribute column, we concatenate all the rst cell response candidates.</p>
        <p>Finally, for each cell that need to be annotated, we priority select the entity
candidates from two cells search. If we do not have candidates in the \two cells"
searches, we select the candidates from a cell search.
2.5</p>
      </sec>
      <sec id="sec-2-4">
        <title>Step 2: Post processing</title>
        <p>We perform value matching of the cells in the core attribute with the
remaining cells in their corresponding rows in the post processing step. This module
performs context similarity calculation between candidate statements and table
row values then ranks the entity candidates based on these similarities.
Value matching This module applies to each data row of a table. We calculate
the similarities between the candidate statements in the core attribute column
with other corresponding values in the same row for each table's data row. If
many statements have the same similarity score, we select all these statements.
We calculate the context similarity of the candidate in the core attribute by
taking the average of all statement similarities in the same row.
CEA: If the target cells are in the core attribute, we select a candidate with the
highest context similarity as a CEA annotation. If the target cells are not in the
core attribute, we infer these annotations from CEA results in the core attribute
using value matching similarities.</p>
        <p>CPA: To get the CPA annotations, we aggregate all properties of statement
candidates in the same rows, then using majority voting to select the CPA
annotations.</p>
        <p>CTA: To get the CTA annotation, we get the direct types from the CEA
annotations in a column and vote for the majority types to get CTA annotations.</p>
        <p>In the CPA task and CTA tasks, if the system returns many answers for one
target annotation, we randomly select one answer as the nal annotation.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        SemTab 2020 has four rounds, where round 1, 2, 3 were implemented as a form
of AICrowd leaderboard, and round 4 was implemented as a blind setting. There
are ve datasets in total, including the four-round datasets and the Tough Tables
(2T) dataset 6. The four rounds datasets were generated by an automatic dataset
generator [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], while the 2T dataset is manually curated for table annotations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
6 The 2T dataset is in the round 4 dataset
      </p>
      <p>Table 1 reports the overall results of MTab4Wikidata for three matching
tasks (CEA, CTA, and CPA) in SemTab 2020 datasets (the four datasets of four
rounds and the 2T dataset). Overall, these results show that MTab4Wikidata
achieves impressive performances for the three matching tasks: the 1st rank in
the two tasks of CEA and CPA, and the 2nd rank in the CTA task in round 1,
2, 3 and the 1st rank in round 4 and the Tough Tables (2T) dataset.</p>
      <p>The details of MTab4Wikidata's performance on the 2T dataset is shown
in Figure 5. The overall performance decrease in both tasks of CEA and CTA
compared to the four-round datasets. Although the tough tables dataset is more
di cult with a high level of table noisiness, our system still performs e ectively.
The F1, Precision, and Recall are quite similar in our results because our system
tries to make as many annotations as possible. Even though we do not have a
corresponding entity in the knowledge graph, we also return the most relevant
entities.</p>
      <p>TOUGH_MISC TOUGH_HOMO
TOUGH_MISC TOUGH_HOMO</p>
      <p>Fig. 5. MTab4Wikidata results on the Tough Tables dataset</p>
      <p>TOUGH_N0O.9IS6E1
TOUGH_SOR0T.E8D8
0.76
MTab4Wikidata</p>
      <p>ALL
0.91
0.75
0.50
0.25
0.62</p>
      <p>0.74
0.87
(a) CEA
0.84CTRL_DBP</p>
      <p>TOUGH_NOISE1
0.8C3TRL_NOISE2 TOUGH_SORTED 0.66
TOUGH_T2D</p>
      <p>TOUGH_MISSP</p>
      <p>0.75
TOUGH_NOISE2
0.6
0.72
0.6
f1
precision
recall</p>
      <p>MTab4Wikidata</p>
      <p>ALL</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>This paper presents an automatic annotation system called MTab4Wikidata
to match tabular elements: i.e., cells, columns, relations between columns to
Wikidata concepts. In this work, we focus on improving the standard entity
search with the fuzzy search and fuzzy statement search to deal with misspelling
and ambiguity of table content. The fuzzy search works e ectively and achieves
an average of 99.89% recall on the three rounds of SemTab 2020. The statement
search also gives a tremendous e cient improvement where it could eliminate
non-statements candidates. To this end, MTab4Wikidata wins the rst prize at
SemTab 2020.</p>
      <p>Acknowledgements This work was supported by the Cross-ministerial
Strategic Innovation Promotion Program (SIP) Second Phase, \Big-data and AI-
enabled Cyberspace Technologies" by the New Energy and Industrial Technology
Development Organization (NEDO).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Ernesto</given-names>
            <surname>Jimenez-Ruiz</surname>
          </string-name>
          , Oktie Hassanzadeh, Vasilis Efthymiou, Jiaoyan Chen, and
          <string-name>
            <given-names>Kavitha</given-names>
            <surname>Srinivas</surname>
          </string-name>
          .
          <source>Semtab</source>
          <year>2019</year>
          :
          <article-title>Resources to benchmark tabular data to knowledge graph matching systems</article-title>
          .
          <source>In ESWC</source>
          , pages
          <volume>514</volume>
          {
          <fpage>530</fpage>
          . Springer,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Vincenzo</given-names>
            <surname>Cutrona</surname>
          </string-name>
          , Federico Bianchi, Ernesto Jimenez-Ruiz, and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Palmonari</surname>
          </string-name>
          .
          <article-title>Tough tables: Carefully evaluating entity linking for tabular data</article-title>
          .
          <source>In ISWC</source>
          , pages
          <volume>328</volume>
          {
          <fpage>343</fpage>
          . Springer,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Phuc</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , Natthawut Kertkeidkachorn, Ryutaro Ichise, and
          <string-name>
            <given-names>Hideaki</given-names>
            <surname>Takeda</surname>
          </string-name>
          . Mtab:
          <article-title>Matching tabular data to knowledge graph using probability models</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>7</volume>
          {
          <fpage>14</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Gilles</given-names>
            <surname>Vandewiele</surname>
          </string-name>
          , Bram Steenwinckel, Filip De Turck, and Femke Ongenae. Cvs2kg:
          <article-title>Transforming tabular data into semantic knowledge</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>33</volume>
          {
          <fpage>40</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Avijit</given-names>
            <surname>Thawani</surname>
          </string-name>
          , Minda Hu, Erdong Hu, Husain Zafar, Naren Teja Divvala, Amandeep Singh,
          <string-name>
            <surname>Ehsan Qasemi</surname>
          </string-name>
          ,
          <article-title>Pedro A Szekely,</article-title>
          and
          <string-name>
            <given-names>Jay</given-names>
            <surname>Pujara</surname>
          </string-name>
          .
          <article-title>Entity linking to knowledge graphs to infer column types and properties</article-title>
          . In SemTab@ ISWC, pages
          <volume>25</volume>
          {
          <fpage>32</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Daniela</given-names>
            <surname>Oliveira</surname>
          </string-name>
          and
          <article-title>Mathieu d'Aquin. Adog-annotating data with ontologies and graphs</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>1</volume>
          {
          <issue>6</issue>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Hiroaki</given-names>
            <surname>Morikawa</surname>
          </string-name>
          .
          <article-title>Semantic table interpretation using lod4all</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>49</volume>
          {
          <fpage>56</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Yoan</given-names>
            <surname>Chabot</surname>
          </string-name>
          , Thomas Labbe, Jixiong Liu, and Raphael Troncy.
          <article-title>Dagobah: an endto-end context-free tabular data semantic annotation system</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>41</volume>
          {
          <fpage>48</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Marco</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          , Roberto Avogadro, and David Chieregato.
          <article-title>Mantistable: an automatic approach for the semantic table interpretation</article-title>
          .
          <source>In SemTab@ ISWC</source>
          , pages
          <volume>15</volume>
          {
          <fpage>24</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Phuc</surname>
            <given-names>Nguyen</given-names>
          </string-name>
          , Natthawut Kertkeidkachorn, Ryutaro Ichise, and
          <string-name>
            <given-names>Hideaki</given-names>
            <surname>Takeda</surname>
          </string-name>
          . Tabeano:
          <article-title>Table to knowledge graph entity annotation</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .
          <year>01829</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>