<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Table Interpretation using LOD4ALL</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Fujitsu Laboratories Limited</institution>
          ,
          <addr-line>4-1-1 Kamikodanaka, Nakahara-ku, Kawasaki, Kanagawa</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we describe Semantic Table Interpretation using LOD4ALL. LOD4ALL is an LOD search engine developed by Fujitsu laboratories. This engine crawls Linked Open Data from the Web and provides a high-speed search service. There are many tabular data on the Web, and these data are important sources of Knowledge Graphs. Therefore, we have enhanced a crawler that is a component of LOD4ALL for taking in these tabular data. This crawler is able to construct Knowledge Graphs to a tabular data. To evaluate of the function of this crawler, we have participated in the challenge \Semantic Web Challenge on Tabular Data to Knowledge Graph Matching".</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
      <p>Presentation of the system</p>
    </sec>
    <sec id="sec-2">
      <title>General statement</title>
      <p>1.2</p>
    </sec>
    <sec id="sec-3">
      <title>Speci c techniques used</title>
      <p>
        In this section, we describe the overview of our proposed approach in the
following ve steps:
Step 1 Extract candidate entities
Step 2 Resolve an entity type
Step 3 Determine a column type
Step 4 Determine an entity for each cell
Step 5 Extract a relationship of entities for each row
Our proposed approach is similar to that of Efthymiou et al.[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and, we have
imitated the approach of Zwicklbauer et al.[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] from Step 1 to Step 3, and we
have enhanced Step 3 in our approach. Step 4 and Step 5 are original processes.
Step 0 builds some databases as a preparatory step. Fig. 1 is an overview of our
proposed system.
      </p>
      <p>Step 1
Extract candidate
entities</p>
      <p>Step 2
Resolve an entity
type</p>
      <p>Step 3
Determine a
column type</p>
      <p>Step 4</p>
      <p>Determine an
entity for each cell</p>
      <p>Step 5
Extract a relationship
of entities for each row</p>
      <p>RDF store
Score DB
Keyword
store</p>
      <sec id="sec-3-1">
        <title>CTA result</title>
      </sec>
      <sec id="sec-3-2">
        <title>CEA result</title>
      </sec>
      <sec id="sec-3-3">
        <title>CPA result</title>
        <p>Step 0
Build Score DB</p>
        <p>Step 0
Build Keyword
store</p>
        <p>Input/output
Output Result</p>
        <p>Reference
Step 0: Build Databases Step 0 builds some databases for a predict on.
Our proposed approach can only build these from a target Knowledge Graph.
Therefore, our approach does not build predicted databases using annotated
tables like T2Dv21. Our approach has two databases. One is the Score DB to
resolve CTA task in the challenge. Another is the Keyword Store for Step 1.</p>
        <p>
          In DBpedia, an entity has some classes. For example, dbr:Barack_Obama
has dbo:Person, dbo:Politician, and dbo:President. For recognizing a more
detailed class, we have adopted Okapi BM25[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which is a traditional retrieval
score as the Score DB. To calculate this score, we have considered entities as
terms and classes as documents. Adopting this score to DBpedia, the Okapi
BM25 score of dbr:Barack_Obama is Table 1. By using the Okapi BM25, we
could be assigned a higher score to a more detailed class.
1 http://webdatacommons.org/webtables/goldstandardV2.html
Step 1: Extract candidate entities This step obtains candidate entities for
each cells in a target tabular data. We have used the literal search function
of LOD4ALL. The original literal search function of LOD4ALL uses
Elasticsearch that returns a subject (a candidate entity) corresponding to an object
that matches a cell value. We have enhanced this literal search function for this
challenge. This function can obtain candidate entities in the following steps.
Step 1-1 Direct search: First, by combining http://dbpedia.org/resource/
with the cell value, we create a candidate entity. Next, we execute ASK
query to RDF store. For example, if the cell value is \Japan",
executing SPARQL like ASK f &lt;http://dbpedia.org/resource/Japan&gt;
?p ?o . g.
        </p>
        <p>
          Step 1-2 Keyword search: We have enhanced the literal search function of
LOD4ALL that can retrieve subjects that have rdfs:label, foaf:
name, foaf:surname and foaf:givenName in the triple. Further more,
we have added keywords for a person that combines a rst
character, the object value corresponding to foaf:givenName, \." and the
object value corresponding to foaf:surname. Through this
enhancement, we can retrieve dbr:Barack_Obama from the cell value of \B.
Obama". In addition, a similar string search step has been added. We
have adopted SimString[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for the similar string search.
        </p>
        <p>Step 1-3 Output candidate entities: To Output in order of highest TOP-K
score. We have set 100.0 for the score of an entity found by Direct
search, and set a value multiplying the Elasticsearch's score and the
SimString's score for the score of an entity that have found by
Keyword search.</p>
        <p>The ow of this step is Fig. 2
Step 2: Resolve an entity type In this step, classes for each entity are
obtained from the Score DB. This process is similar to those in to previous
studies.</p>
        <p>Step 3: Determine a column type In this step, estimating a column type
(CTA result) using candidate entities extracted in Step 2 and the Score DB.
This step consists of the three following steps:
Step 3-1 Calculate a cover ratio</p>
        <p>RDF store
Keyword
store</p>
        <p>Step 1-1</p>
        <p>Step 1-2
SimString</p>
        <p>DB</p>
        <p>Input a cell value</p>
        <p>Direct search</p>
        <p>Not found</p>
        <p>Not found
partial match
partial match
extended
keywords</p>
        <p>using
SimString DB</p>
        <p>Hit
Hit
Hit</p>
        <p>Candidate
entities</p>
        <p>Step 1-3</p>
        <p>Filter
TOP-k</p>
        <p>Candidate
entities
result
Step 3-2 Search a class score from the score DB
Step 3-3 Calculate a predict score using both a score ratio and a class
score</p>
        <p>We illustrate the running example of Step 2 and Step 3 using a part of the
\List of residences of Presidents of the United States#Summer White House#0.csv"
in Round2 in Fig. 3. We have set Step 1's K as 1 in this gure to explain easily.
In this gure, our approach outputs dbo:President as the result of this step.
Step 4: Determine an entity for each cell In this step, we determine an
entity for each cell that has the class (the column type) determined in Step 3
from candidate entities that extract by Step 1. The output of this step is the
result of the CEA task.</p>
        <p>Step 5: Extract a relationship of entities for each row In this step, we
rst collect candidate predicates using entities determined in Step 4 by executing
SPARQL (Listing 1.1) to RDF store. %URI1% and %URI2% in SPARQL are
replaced entities determined in Step 4. Next, by calculating a frequency, we
obtain the predicate of the greatest frequency. The output of this step is the
result of the CPA task.</p>
        <p>We illustrate the running example of the Step 5 in Fig. 4.</p>
        <p>Listing 1.1. SPARQL query used the Step 5
PREFIX dbr :</p>
        <p>&lt;http :// dbpedia . org / resource /&gt;
select distinct ? predicate where {</p>
        <p>% URI1 % ? predicate % URI2 %
}
Class</p>
        <p>Cover ratio
dbo:Person
dbo:Politician
dbo:President
dbo:OfficeHolder
1.0
0.8
0.8
0.2
average score per
a class (STEP3-2)
0.23
1.31
5.17
1.95</p>
        <p>Predict score
0.23
1.05
4.14
0.39</p>
        <p>Get TOP-1 class</p>
        <p>dbo:President
%URI2%
We designed predict score functions for Step 3 through trial and error. The best
result in our trial is Equation 1. Here, the normalizedClassScore is the class
score that has scaled between 0 and 1. and are hyperparameters.</p>
        <p>P redictScore =
normalizedClassScore +
ratioScore
(1)
In our experiment, by setting =1.2 and =2.5, the system has output the best
of results. In Step 1, the top k number is K=10, and the similarity parameter
for SimString is 0.7.
1.4</p>
        <p>Link to the system and parameters le
We plan to publish our modules in the GitHub2 .
2</p>
        <sec id="sec-3-3-1">
          <title>Results</title>
          <p>2 https://github.com/lod4all/semanticTableInterpretation
Round1
Round2
Round3
Round4</p>
          <p>CTA
CEA
CPA
CTA
CEA
CPA
CTA
CEA
CPA
CTA
CEA</p>
          <p>CPA</p>
          <p>Round4
(After closing) CTA</p>
          <p>CEA</p>
          <p>CPA
3</p>
          <p>General comments
In these challenge datasets, there were several cases where it was difficult to
determine the column type in the CTA task. One of them is shown in Fig. 5.
This table is the ranking of Forbes Korea Power Celebrity in 20123. We found
this in the Round2 dataset. The Name column in this table may be assigned
dbo:Person and dbo:Group. It is difficult to determine a unique column type for
CTA task. To deal with this case, we believe it is necessary to ues an evaluation
method that allows multiple perfect annotations for a column.
4</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Conclusions</title>
          <p>
            In this paper, we described Semantic Table Interpretation using LOD4ALL. Our
proposed approach is still in the early stages, so there are a lot of problems to
be resolved. In the future, we will improve the accuracy of collecting candidate
entities in Step 1 and the prediction process in Step 3 in these problems in
particular. Further more, we will improve our system by developing an ensemble
approach between our approach and others. For CTA task, we believe it is
necessary to improve DBpedia. For example, we will assign a class to an entity that
has not been assigned a class by adopting Fang's method[
            <xref ref-type="bibr" rid="ref9">9</xref>
            ].
3 https://en.wikipedia.org/wiki/Forbes Korea Power Celebrity
          </p>
          <p>Beast*
10</p>
          <p>Park Tae-Hwan</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Lehmberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Oliver</surname>
          </string-name>
          , et al. \
          <article-title>A large public corpus of web tables containing time and context metadata</article-title>
          .
          <source>" Proceedings of the 25th International Conference Companion on World Wide Web. International World Wide Web Conferences Steering Committee</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kruit</surname>
            , Benno,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Boncz</surname>
          </string-name>
          , and Jacopo Urbani. \
          <article-title>Extracting Novel Facts from Tables for Knowledge Graph Completion."</article-title>
          <source>International Semantic Web Conference</source>
          . Springer, Cham,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Naseer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Aisha</surname>
          </string-name>
          , et al. \
          <article-title>LOD for all: Unlocking in nite opportunities." Semantic Web Challenge (</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Semantic</given-names>
            <surname>Web</surname>
          </string-name>
          <article-title>Challenge on Tabular Data to Knowledge Graph Matching http</article-title>
          : //www.cs.ox.ac.uk/isg/challenges/sem-tab/.
          <source>(accessed Sep</source>
          .
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <surname>Vasilis</surname>
          </string-name>
          , et al. \
          <article-title>Matching web tables with knowledge base entities: from entity lookups to entity embeddings."</article-title>
          <source>International Semantic Web Conference</source>
          . Springer, Cham,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Zwicklbauer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Stefan</surname>
          </string-name>
          , et al. \
          <source>Towards Disambiguating Web Tables." International Semantic Web Conference (Posters &amp; Demos)</source>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Robertson</surname>
          </string-name>
          , Stephen, and Hugo Zaragoza. \
          <article-title>The probabilistic relevance framework: BM25 and beyond</article-title>
          .
          <source>" Foundations and Trends in Information Retrieval 3</source>
          .4 (
          <year>2009</year>
          ):
          <fpage>333</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Okazaki</surname>
          </string-name>
          , Naoaki, and
          <article-title>Jun'ichi Tsujii. \Simple and efficient algorithm for approximate dictionary matching</article-title>
          .
          <source>" Proceedings of the 23rd International Conference on Computational Linguistics. Association for Computational Linguistics</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Fang</surname>
            , Lu,
            <given-names>Qingliang</given-names>
          </string-name>
          <string-name>
            <surname>Miao</surname>
          </string-name>
          , and Yao Meng. \
          <article-title>DBpedia Entity Type Inference Using Categories."</article-title>
          <source>International Semantic Web Conference (Posters &amp; Demos)</source>
          .
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>