<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MTab: Matching Tabular Data to Knowledge Graph with Probability Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Phuc Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natthawut Kertkeidkachorn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryutaro Ichise</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hideaki Takeda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Advanced Industrial Science and Technology</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute of Informatics</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>SOKENDAI (The Graduate University for Advanced Studies)</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the design of our system, namely MTab, for Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab 2019). MTab combines the voting algorithm and the probability model to solve critical bottlenecks of the matching task. Results on SemTab 2019 show MTab obtains the promising performance. Tabular Data to Knowledge Graph Matching (SemTab 2019) 4 is a challenge on matching semantic tags from table elements to knowledge bases (KBs), especially DBpedia. Fig. 1 depicts the three sub-tasks for SemTab 2019. Given a table data, CTA (Fig. 1a) is the task of assigning a semantic type (e.g., a DBpedia class) to a column. In CEA (Fig. 1b), a cell is linked to an entity in KB. The relation between two columns is assigned to a property in KB in CPA (Fig. 1c).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>To address the three tasks of the challenge, we designed our system (MTab) by
the 4-steps pipeline as shown in Fig. 2.</p>
      <p>
        Step 1 is to pre-process a table data by predicting languages of the table with
fasttext [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], correcting spelling, predicting data types (e.g., number or text), and
searching relevant entities in DBpedia. Due to the heterogeneous problem, we
utilize entity searching on many services including DBpedia Lookup, DBpedia
4 http://www.cs.ox.ac.uk/isg/challenges/sem-tab/
      </p>
      <p>Copyright 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>Phuc Nguyen et al.
endpoint. Also, we search relevant entities on Wikipedia and Wikidata by
redirected links to DBpedia to increase the possibility of nding the relevant entities.
We assume that cells in a column have the same type. We then use information
from Step 1 to estimate the probability of the types for the column in Step 2.
The type candidate which has the highest probability is the result for CTA task.
In Step 3, the result of Step 2 and information of Step 1 are used to estimate
the probability for entities. Similarly, the entity candidate which has the highest
probability is the result for CEA task. In Step 4, we use the result from Step
3 to estimate the property between two entities, and then, adopt the voting
technique to estimate the probability for all rows of two columns. The result for
CPA is the highest probability of property candidate in Step 4. We repeatedly
execute Step 2, 3 and 4 to nd the best candidates for columns, cells, and the
relation between two columns.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Results and Conclusion</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Bag of tricks for e cient text classi cation</article-title>
          .
          <source>In: EACL 2017</source>
          . pp.
          <volume>427</volume>
          {
          <fpage>431</fpage>
          .
          <string-name>
            <surname>ACL</surname>
          </string-name>
          (April
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>