<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>WInte.r - A Web Data Integration Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oliver Lehmberg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Brinkmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bizer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science Group, Universitat Mannheim B6 26</institution>
          ,
          <addr-line>68159 Mannheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Web provides a plethora of structured data, such as semantic annotations in web pages, data from HTML tables, datasets from open data portals, or linked data from the Linked Open Data Cloud. For many use cases, it is necessary to integrate such web data with existing local datasets. This integration entails schema matching, identity resolution, as well as data fusion. As an alternative to using a combination of partial or ad hoc solutions, this poster presents the Web Data Integration Framework (WInte.r ), which supports end-to-end data integration by providing algorithms and building blocks for data pre-processing, schema matching, and identity resolution, as well as data fusion. While being fully usable out-of-the box, the framework is highly customisable and allows for the composition of sophisticated integration architectures such as T2K Match, which is used to match millions of web tables against DBpedia. A second use case for which WInte.r was employed is the task of stitching (combining) web tables from the same web site into larger tables as a preprocessing step before matching. The WInte.r framework is written in Java and is available as open source under the Apache 2.0 license.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Many web-based systems need to combine data from various sources, such as
semantic annotations in web pages, data from HTML tables, datasets from open
data portals, or linked data from the Linked Open Data Cloud. These data often
do not use the same schema and miss explicit links between their entities. Such
heterogeneous data sources must hence be integrated before they can be used
for any further use case.</p>
      <p>
        While there are tools available to solve speci c data integration problems,
such as schema matching, identity resolution or data fusion, using these tools
in order to realise a complete data integration work ow can be cumbersome
and requires researchers to spend quite some time on writing glue code. Also,
scalability is often a problem as many tools are designed for use cases with
only a few datasets and not for web-scale scenarios where thousands of data
sources need to be integrated. As an alternative, this poster presents the Web
Data Integration (WInte.r ) Framework. The WInte.r framework supports
endto-end data integration processes by providing algorithms and building blocks
for web data pre-processing, schema matching, identity resolution, as well as
data fusion. As an example, the integration of millions of tables collected from
web pages with the DBpedia knowledge base has been implemented using the
WInte.r framework in the T2K Match [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] project.1
      </p>
      <p>The WInte.r framework is fully usable out-of-the box with implementations
for all mentioned data integration tasks and provides highly customisable
functionality that allows for the composition of sophisticated integration
architectures. This allows researchers to focus on the actual task instead of starting from
scratch and repeatedly implementing the same functionality in di erent projects.
The WInte.r framework is written in Java and is available as open source under
the Apache 2.0 license.2
2</p>
      <p>Functionality
The WInte.r framework covers all steps of the data integration process, including
data loading, pre-processing, schema matching, identity resolution, and data
fusion. This section gives an overview of the methods that are provided for each
of these steps. Figure 1 shows the components that are involved in the integration
process.
The WInte.r data sets act as an interchangeable data management and
processing component for the rest of the framework. They implement essential
processing operations such as iteration, ltering, transformations, aggregations and
joins. By exchanging the data set implementation, a user can, for example, switch
between single-threaded or parallel execution.</p>
      <p>Data Loading. WInte.r provides readers for standard data formats such as CSV,
XML and RDF. In addition, WInte.r implements a specialized JSON format for
representing tabular data from the Web together with meta-information about
the origin and context of the data, as used by the Web Data Commons (WDC)
web tables corpora.3 Speci cally, the context represents data such as the original
URL, page and section headings and the surrounding text from the web page
from which the tables were extracted.</p>
    </sec>
    <sec id="sec-2">
      <title>1 https://github.com/olehmberg/T2KMatch 2 https://github.com/olehmberg/winter 3 http://webdatacommons.org/webtables/</title>
      <p>Pre-processing. During pre-processing, data is prepared for the methods that
are applied later on in the integration process. WInte.r provides specialized
preprocessing methods for data with missing schema information, such as: data
type detection, unit of measurement normalization, header detection, and entity
name detection (In cases where no explicit \rdfs:label" property is available, the
entity name detection nds the property that most likely contains the entity
names).
2.2</p>
      <sec id="sec-2-1">
        <title>Matching</title>
        <p>The matching components are used for schema matching and identity resolution.
First, an optional blocking step generates candidate pairs of records to reduce the
total number of comparisons. Then, a matching rule evaluates each candidate
pair and decides whether or not to create a correspondence. Matching rules
are speci ed by the user and can, for example, calculate the weighted sum of
di erent similarity values and apply a threshold. Alternatively, matching rules
can be learned using supervised machine learning. Given a labelled training set,
the framework will apply a user-speci ed machine learning algorithm from the
WEKA machine learning library to learn the matching rule.</p>
        <p>Schema Matching. Schema matching methods nd attributes in two schemata
that have the same meaning. WInte.r provides three pre-implemented schema
matching algorithms which either rely on attribute labels (label-based schema
matching) or data values (instance-based schema matching), or exploit an
existing mapping of records (duplicate-based schema matching) in order to nd
attribute correspondences.</p>
        <p>Identity Resolution. Identity resolution methods (also known as data matching
or record linkage methods) identify records that describe the same real-world
entity. The pre-implemented identity resolution methods can be applied to a single
dataset for duplicate detection or to multiple datasets in order to nd
recordlevel correspondences. WInte.r provides the following pre-implemented methods:
standard blocking by single/multiple blocking key(s), Sorted-Neighbourhood
Method, token-based identity resolution and rule-based identity resolution.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Data Fusion</title>
        <p>
          Data fusion methods combine the data from multiple sources into a single,
consolidated dataset. For this task, they rely on the schema- and record-level
correspondences that were discovered in the previous steps of the integration process.
However, di erent sources may provide con icting data values. WInte.r provides
a mechanism to resolve such data con icts (i.e., deciding which value to
include in the nal dataset) by applying di erent con ict resolution functions for
strings, numbers, lists of values, data type independent functions, and functions
that consider metadata such as creation time or dataset speci c trust scores.
A user can specify a fusion strategy that speci es how data con icts should be
handled for each property.
This section gives an overview of existing work based on the WInte.r framework.
Integration of Large Numbers of Data Sources: Augmenting the DBpedia
Knowledge base with Web Table Data. Many web sites provide data in the form of
HTML tables. Data from these tables can be used to ll missing values in large
cross-domain knowledge bases such as DBpedia. An example of how pre-de ned
building blocks from the WInte.r framework are combined into an advanced,
use-case speci c integration method is the T2K Match algorithm [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].4 The
algorithm matches millions of web tables against a central knowledge base describing
millions of instances belonging to hundreds of di erent classes [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
Preprocessing for Large-Scale Matching: Stitching Web Tables for Improving
Matching Quality. Tables on web pages (\web tables") cover a diversity of topics
and can be a source of information for di erent tasks such as knowledge base
augmentation or the ad-hoc extension of datasets. The challenges that matching
methods for this purpose have to overcome are the high heterogeneity and the
small size of the tables. To counter these problems, web tables from the same
web site can be stitched (combined) before running any of the existing matching
systems, which improves the matching quality especially for small tables [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].5
Data Search 4 Data Mining (DS4DM). Analysts increasingly have the problem
that they know that some data which they need for a project is available
somewhere on the Web or in the corporate intranet but they are unable to nd the
data. The goal of the Data Search 4 Data Mining (DS4DM) project is to extend
the data mining platform Rapidminer with data search and data integration
functionalities which enable analysts to nd relevant data in potentially very
large data corpora and to semi-automatically integrate the discovered data with
existing local data.6
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 https://github.com/olehmberg/T2KMatch</title>
      <p>5 https://github.com/olehmberg/WebTableStitching
6 http://ds4dm.de/en/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>O.</given-names>
            <surname>Lehmberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Stitching Web Tables for Improving Matching Quality</article-title>
          . PVLDB,
          <volume>10</volume>
          (
          <issue>11</issue>
          ):
          <volume>1502</volume>
          {
          <fpage>1513</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ritze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lehmberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Matching HTML Tables to DBpedia</article-title>
          .
          <source>In Proc. of the 5th Int. Conference on Web Intelligence</source>
          , Mining and Semantics, page
          <volume>10</volume>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ritze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lehmberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Oulabi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Pro ling the potential of web tables for augmenting cross-domain knowledge bases</article-title>
          .
          <source>In Proceedings of the 25th WWW</source>
          , pages
          <volume>251</volume>
          {
          <fpage>261</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>