<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cross-Fertilizing Deep Web Analysis and Ontology Enrichment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CNRS LTCI Te´ le´ com ParisTech</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CNRS LTCI T e´le´ com ParisTech</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CNRS LTCI Paris</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France Paris</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France Paris</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Deep Web databases, whose content is presented as dynamicallygenerated Web pages hidden behind forms, have mostly been left unindexed by search engine crawlers. In order to automatically explore this mass of information, many current techniques assume the existence of domain knowledge, which is costly to create and maintain. In this article, we present a new perspective on form understanding and deep Web data acquisition that does not require any domain-specific knowledge. Unlike previous approaches, we do not perform the various steps in the process (e.g., form understanding, record identification, attribute labeling) independently but integrate them to achieve a more complete understanding of deep Web sources. Through information extraction techniques and using the form itself for validation, we reconcile input and output schemas in a labeled graph which is further aligned with a generic ontology. The impact of this alignment is threefold: first, the resulting semantic infrastructure associated with the form can assist Web crawlers when probing the form for content indexing; second, attributes of response pages are labeled by matching known ontology instances, and relations between attributes are uncovered; and third, we enrich the generic ontology with facts from the deep Web.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>ONTOLOGIES AND THE DEEP WEB</title>
      <p>The deep Web consists of dynamically-generated Web pages that
are reachable by issuing queries through HTML forms. A form is
a section of a document with special control elements (e.g.,
checkboxes, text inputs) and associated labels. Users generally interact
with a form by modifying its controls (entering text, selecting menu
items) before submitting it to a Web server for processing.</p>
      <p>
        Forms are primarily designed for human beings, but they must
also be understood by automated agents for various applications
such as general-purpose indexing of response pages, focused
indexing [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], extensional crawling strategies (e.g., Web archiving),
automatic construction of ontologies [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], etc. However, most
existing approaches to automatically explore and classify the deep
Web crucially rely on domain knowledge [
        <xref ref-type="bibr" rid="ref10 ref12 ref30">10, 12, 30</xref>
        ] to guide form
understanding. Moreover, they tend to separate the steps of form
interface understanding and information extraction from result pages,
although both contribute [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] to a more authentic vision on the
backend database schema. The form interface exposes in the input
schema some attributes describing the query object, while response
pages present this object instantiated in Web records that outline
the form output schema. In this paper, we determine a mapping
beVLDS’12 August 31, 2012. Istanbul, Turkey.
      </p>
      <p>Copyright c 2012 for the individual papers by the papers’ authors. Copying
permitted for private and academic purposes. This volume is published and
copyrighted by its editors.
tween the input and output schemas which associates the data types
corresponding to form elements in the input schema to instances
aligned in the output schema.</p>
      <p>
        A harder challenge is to understand the semantics of these data
types and how they relate to the object of the form. The input–output
schema mapping may give us hints, such as the input schema labels,
but this information cannot suffice by itself. This has been addressed
in related work using heuristics [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] or an assumed domain
knowledge [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] which is either manually crafted or obtained by merging
different form interface schemas belonging to the same domain.
Domain knowledge is, however, not only hard to build and maintain,
but also often restricted to a choice of popular domain topics, which
may lead to biased exploration of the deep Web.
      </p>
      <p>
        We present a new way to deal with this challenge: we initially
probe the form in a domain-agnostic manner and transform the
information extracted from response pages into a labeled graph. This
graph is then aligned with a general-domain ontology, YAGO [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
using the PARIS ontology alignment system [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. This allows us to
infer the semantics of the deep Web source, to obtain new,
representative query terms from YAGO for the probing of form fields, and to
possibly enrich YAGO with new facts.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Merging input schemas of deep Web interfaces has been used
to acquire domain ontologies automatically [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] and perform Web
database classification and query routing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The main drawback
of these approaches is that data integration dramatically relies on
the interface schema, whose shallow features (the form structure
and labels) are neither complete, nor representative enough for the
actual response records [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        To obtain response pages, the form has to be filled in and
submitted first. Most approaches described in the literature are
domainspecific and use dictionary instances [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Domain-agnostic probing
approaches are more powerful because they do not make such
assumptions and incrementally build knowledge that tends to improve
the probing and the quality of response pages. However, existing
domain-agnostic techniques do not aim at understanding the
intensional purpose of the form, but at extensional crawling [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Deep Web response pages are an extremely rich source of
semistructured information. Works dealing with response pages assume
the form probing mechanism understood and focus on information
extraction (IE) from Web records [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Extracting the schema from
response pages [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] is possible due to the structural similarity of
records. Because this schema has been obtained by probing the form
and analyzing the response pages, it is called the output schema of
the form.
      </p>
      <p>
        The data extracted from deep Web sources through IE
processing is typically used to build and/or enrich ontologies [
        <xref ref-type="bibr" rid="ref2 ref21 ref24">2, 21, 24</xref>
        ],
gazzetters [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or to expand sets of entities [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. ODE [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] in
particular gets closer to our work by its holistic approach, but still needs a
domain ontology built by matching different deep Web interfaces. A
more important difference appears in the annotation of the extracted
data from response pages using heuristic rules for label assignment,
similar to [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Comparatively, we use PARIS alignment algorithm.
      </p>
      <p>
        The next step is the discovery of the semantic relationships
between the entity of the form and the record attributes; for this, several
techniques are proposed in the literature. Traditionally, statistical
and rule-based methods use the instances in a textual context in
order to infer the relation between them [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Another option [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
is to match the terminology of a given term with a known concept
using semantic resources such as DBpedia or WordNet [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Yet
another trend is to use classifiers that can predict specific relations
(e.g., subClassOf ) given enough training and test data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The
closest work to ours may be [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], an approach relying on supervised
learning that uses a generic ontology to infer types and relations
among the data in a Web table. We deal with the more general
setting of deep Web interfaces here, and we propose a fully automatic
approach that does not require human supervision.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>ENVISIONED APPROACH</title>
      <p>
        We now present our vision of a holistic deep Web semantic
understanding and ontology enrichment process, which is summarized
in Figure 1: a Web form is analyzed and probed, record attribute
values are extracted from result pages, and their types are mapped to
input fields. While these steps are rather standard and we follow the
well-established best practices, they have never been analyzed in a
holistic manner without the assumption of domain knowledge that
describes the form interface. The novelty of studying these steps
in connection comes from their contribution to the formation of a
labeled graph which encompasses data values of unknown types
and implicit semantic relations. This graph is further aligned with a
generic ontology for knowledge discovery using PARIS.
Form Analysis and Probing. The form interface is
presented as an input schema which gives a prescriptive description
of the object that the user can query through the form. The input
schema is the ordered list of labels corresponding to form elements,
possibly together with constraints and possible values (for
dropdown lists and other non-textual input fields). Important data
constraints or properties of the backend Web database can be discovered
through well-designed probing and response page analysis. Some
may be precious for a crawler that interacts with the form: Are stop
words indexed? Which Boolean connectors are used (conjunctive or
disjunctive)? Is the search affected by spelling errors? We perform
form probing in an agnostic manner (i.e., without domain
knowledge) following [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We try to set non-textual input elements or
to fill in a textual input field with stop words or with contextual
terms extracted from non-textual input controls (e.g., drop-down list
entries) or surrounding text (e.g., indications to the user). We rely
on the fact that many sites provide a generous index (i.e., a response
page can be obtained by inputting a single letter). A more elaborate
idea is to use AJAX auto-completion facilities.
      </p>
      <p>
        Record Identification. If the form has been filled in correctly,
we obtain a result page. Otherwise, to identify possible error pages,
our method infers a characteristic XPath expression by submitting
the form with a nonsense word and tracing its location in the DOM
of the response page. This approach uses the fact that the nonsense
word will usually be repeated in the error page to present the
erroneous input to the user. If not, techniques such as those of [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
can be applied. If the probing yields a response page which does
not contain the error pattern, then we determine the generic XPath
location of Web records using [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Output Schema Construction. A way to build the output
schema is to use the reflection of a given domain knowledge in
response pages [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Another method is to perform attribute
alignment [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for records obtained from different pages. Since Web
records represent subtrees which are structurally similar at DOM
level, we extract the values of their textual leaf nodes and cluster
these values based on their DOM path. The rationale is that the
values found under the same record internal path are attributes of
the same type. For instance, “Great Expectations” and “David
Copperfield” in Figure 1 both represent literals of the title attribute of
a book and have a common location pattern. We define a record
feature as the association between a relevant record internal path and
its cumulated bag of instances. The output schema for a response
page is then defined by the ordered sequence of record features. In
practice, we remove uninformative record features from the output
schema by restricting ourselves to paths which contain different
instances across various response pages.
      </p>
      <p>Input and Output Schema Mapping. We align input fields
of the form with record features of the result pages in the following
fashion. For non-textual form elements such as drop-down lists,
we check if their values do not trivially match one of the record
features of the output schema. For textual form elements, we use a
more elaborate idea. Due to binding patterns, query instances which
appear at a certain record internal path should appear again at the
same location when they are submitted in the “right” input field
for this path. If we submit them in an unrelated field, however, we
should obtain an error page or unsuitable results. Formally, given
a record feature f of the output schema, we can see if it maps to
a textual input t by filling in t with one of the initial instances of
f (say i) and submitting the form. Either we obtain an error page,
which means f and t should not be mapped, or we obtain a result
page in which we can use f ’s record internal path to extract a new
bag I of instances for f . In this case, we say that t and f are mapped
if all instances in I are equal to i or contain it as a substring (i.e., i
appears again at f ’s location pattern). We obtain the mapping by
performing these steps for all couples ( f ; t).</p>
      <p>
        Most of the time, the input–output schemas do not match exactly.
The attributes that cannot be matched are usually explicit in the input
schema (e.g., given by non-textual inputs, like drop-down lists), or
only present in the output schema (e.g., the price of a book).
Graph Generation. We represent the data extracted from the
Web records as RDF triples [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], in the following manner:
1. each record is represented as an entity;
2. all records are of the same class, stated using rdf:type;
3. the attribute values of records are viewed as literals;
4. each record links to its attribute values through the relation
(i.e., predicate) that corresponds to the record internal path of
the attribute type in the response page;
Since the triples form a labeled directed graph, it is possible to
add much more information to the representation, provided that we
have the means to extract it. An idea would be to include a more
detailed representation of a record by following the hyperlinks that
we identify in its attribute values and replacing them in the original
response page with the DOM tree of the linked page. In this way,
the extraction can be done on a more complete representation of
the backend database. We can also add complementary data from
various sources, e.g., Web services or other Web forms belonging to
the same domain.
      </p>
      <p>Book
rdfs:type
rdfs:type
rdfs:type</p>
      <p>Othello</p>
      <p>Great
Expectations</p>
      <p>David
Copperfield
(novel)
y:hasName</p>
      <p>y:created
y:hasName
y:created
y:created
y:hasName
List of records</p>
      <p>Great Expectations
Charles Dickens
Dover Thrift Editions
David Copperfield
by Charles Dickens
Penguin Classics</p>
      <p>RDF
triples
generation</p>
      <p>Yago
"Othello"
"Great Expectations"
"David Copperfield"
ontology
alignment
?class
rdfs:type</p>
      <p>?e1
rdfs:type ?e2
Shakespeare y:hasName</p>
      <p>"Shakespeare"
Charles
Dickens
y:hasName</p>
      <p>"Charles Dickens"
Labeled graph</p>
      <p>
        ontology
enrichment
"Great Expectations"
"Charles Dickens"
"Dover Thrift Editions"
"David Copperfield"
"by Charles Dickens"
"Penguin Books"
Ontology Alignment. The ontology that we compile from the
result pages is aligned with a large reference ontology. We use
YAGO [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], though our approach can be applied to any reference
ontology. We use PARIS [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] to perform the ontology alignment.
Unlike most other systems, PARIS is able to align both entities
and relations. It does so by bootstrapping an alignment from the
matching literals and propagating evidence based on relation
functionalities. Through the alignment, we discover the class of entities,
the meaning of record attributes and the actual relation that exists
between them. Two main adaptations are needed to use PARIS in
the deep Web data alignment process. First, extracted literals
usually differ from those of YAGO because of alternate spellings or
surrounding stop words. A typical case on Amazon is the addition
of related terms, e.g., “Hamlet (French Edition)” instead of just
“Hamlet”. To mitigate this problem we normalize the literals,
eliminate punctuation and stop words. Pattern identification in the data
values of the same type could increase the probability of extracting
cleaner values. We are working on a way to index YAGO literals in
a manner that is resilient to the small differences we wish to ignore.
A promising approach to do this is shingling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Second, an entity-to-literal relation in the labeled graph may not
necessarily correspond to a single edge in the reference ontology,
but to a sequence of edges. This amounts to a join of the involved
relations; a typical case in our prototype is the “author” attribute
which is linked to a record entity through a two-step YAGO path
“y:created y:hasPreferredName”. To ensure that the alignment with
joins, typically costly, can be performed in practice, we limit the
maximal length of joins. A consequence is that PARIS will explore
a smaller fraction of YAGO in the search for relations relevant to the
data of our labeled graph. In addition to the use of record attribute
values as literals, PARIS could use the form labels (through the
input–output mappings) to guide the alignment and favor YAGO
relations with a similar name. Some record instances do not align
with any literal of the ontology. The cause is that they represent
information which is unknown to YAGO.</p>
      <p>Form Understanding and Ontology Enrichment.
Ontology alignment gives us knowledge about the data types, the
domains and ranges of record attributes, and their relation to the object
of the form (in our case, a book). The propagation of this
knowledge to the input schema through the input–output mapping (for
the form elements that have been successfully mapped) results in a
better understanding of the form interface. On the one hand, we can
infer that a given field of the Amazon advanced search form expects
author names, and leverage YAGO to obtain representative author
names to fill in the form. This is useful in intensional or extensional
automatic crawl strategies of deep Web sources. On the other hand,
we can generate new result pages for which data location patterns
are already known and enrich YAGO through the alignment that we
once determined.</p>
      <p>There are three main possibilities to enrich the ontology. First,
we can add to the ontology the instances that did not align. For
instance, we can use the Amazon book search results to add to YAGO
the books for which it has no coverage. Second, we can add facts
(triples) that were missing in YAGO. Third, we can add the relation
types that did not align. For instance, we can add information
about the publisher of a book to YAGO. This latter direction is
more challenging, because we need to determine if the relation
types contain valuable information. One safe way to deal with this
relevance problem is to require attribute values to be mapped to a
form element in the input schema. We can then use the label of the
element to annotate them.
4.</p>
    </sec>
    <sec id="sec-4">
      <title>PRELIMINARY EXPERIMENTS</title>
      <p>We have prototyped this approach for the Amazon book advanced
search form1. Obviously, we cannot claim any statistical
significance of the results we report here, but we believe that the approach,
because it is generic, can be successfully applied to other sources of
the deep Web.</p>
      <p>Our preliminary implementation performed agnostic probing of
the form, wrapper induction, and mapping of input–output schemas.
It generated a labeled graph with 93 entities and 10 relation types
out of which 2 (title and author) are recognized by YAGO. Literals
underwent a semi-heuristic normalization process (lowercasing,
removal of parenthesized substrings). We then replaced each extracted
1http://www.amazon.com/gp/browse.html?node=241582011
literal with a similar literal in YAGO, if the similarity (in terms of the
number of common 2-grams) was higher than an arbitrary threshold.</p>
      <p>We aligned this graph with YAGO by running PARIS for 15
iterations, i.e., a run time of 7 minutes (most of it was spent loading
YAGO, the proper computation took 20 seconds). Though the vast
majority of the books from the dataset were not present in YAGO,
the 6 entity alignments with best confidence were books that had
been correctly aligned through their title and author. To limit the
effect of noise on relation alignment, we recomputed relation
alignments on the entity alignments with highest confidence; the system
was thus able to properly align the title and author relations with
“y:hasPreferredName” and “y:created y:hasPreferredName”,
respectively. These relations were associated to the record internal paths
of the output schema attributes and propagated to form input fields.</p>
    </sec>
    <sec id="sec-5">
      <title>DISCUSSION</title>
      <p>Our vision is that of a holistic system for deep Web understanding
and ontology enrichment, where each stage of the process (form
analysis, information extraction, schema matching, ontology
alignment, etc.) would benefit of every other part. This is an ambitious
project, but our current prototype already exhibits promising results.</p>
      <p>Many challenges remain to be tackled: resilience to outliers and
noise resulting from imperfect literal matching and information
extraction; proper management of the confidence in the results of
each automatic task, especially when they are used as the input of
another task; identification of new relation types of interest among
those extracted from a Web source; integration of the information
contained in several different deep Web sources of the same domain.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We acknowledge Fabian Suchanek for initial discussions on this
topic. The research has been funded by the European Union’s
seventh framework programme, in the setting of the European Research
Council grant Webdam, agreement 226513, and the FP7 grant
ARCOMEM, agreement 270239.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Raposo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bellas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Cacheda</surname>
          </string-name>
          .
          <article-title>Extracting lists of data records from semi-structured Web pages</article-title>
          .
          <source>Data and Knowledge Engineering</source>
          ,
          <volume>64</volume>
          (
          <issue>2</issue>
          ),
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Chun</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-C. Huang</surname>
            , and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Geller</surname>
          </string-name>
          .
          <article-title>Enriching ontology for deep Web search</article-title>
          .
          <source>In Proc. DEXA</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geller</surname>
          </string-name>
          , Y.-T. Wu, and
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Chun</surname>
          </string-name>
          .
          <article-title>Semantic deep Web: automatic attribute extraction from the deep Web data sources</article-title>
          .
          <source>In Proc. SAC</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          and
          <string-name>
            <surname>S. Kambhampati.</surname>
          </string-name>
          <article-title>SourceRank: Relevance and trust assessment for deep Web sources based on inter-source agreement</article-title>
          .
          <source>In Proc. WWW</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Barbosa</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Freire</surname>
          </string-name>
          .
          <article-title>Siphoning hidden-Web data through keyword-based interfaces</article-title>
          .
          <source>J. Information and Data Management</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Beisswanger</surname>
          </string-name>
          .
          <article-title>Exploiting relation extraction for ontology alignment</article-title>
          .
          <source>In Proc. ISWC</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. Z.</given-names>
            <surname>Broder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Glassman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Manasse</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Zweig</surname>
          </string-name>
          .
          <article-title>Syntactic clustering of the Web</article-title>
          .
          <source>Computer Networks</source>
          ,
          <volume>29</volume>
          (
          <fpage>8</fpage>
          -
          <lpage>13</lpage>
          ),
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Caverlee</surname>
          </string-name>
          , L. Liu, and
          <string-name>
            <given-names>D.</given-names>
            <surname>Buttler</surname>
          </string-name>
          . Probe, cluster, and
          <article-title>discover: Focused extraction of QA-pagelets from the deep Web</article-title>
          .
          <source>In Proc. ICDE</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          , G. Ladwig, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          .
          <article-title>Gimme' the context: Context-driven automatic semantic annotation with C-PANKOW</article-title>
          .
          <source>In Proc. WWW</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Furche</surname>
          </string-name>
          , G. Gottlob,
          <string-name>
            <given-names>G.</given-names>
            <surname>Grasso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Guo</surname>
          </string-name>
          , G. Orsi, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schallhart</surname>
          </string-name>
          .
          <article-title>Real understanding of real estate forms</article-title>
          .
          <source>In Proc. WIMS</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Furche</surname>
          </string-name>
          , G. Grasso, G. Orsi,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schallhart</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Automatically learning gazetteers from the deep Web</article-title>
          .
          <source>In Proc. WWW</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. C.-C. Chang</surname>
          </string-name>
          , and J. Han.
          <article-title>Discovering complex matchings across Web query interfaces: A correlation mining approach</article-title>
          .
          <source>In Proc. KDD</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Yadav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bharti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Choudhary</surname>
          </string-name>
          .
          <article-title>Accurate and efficient crawling the deep Web: Surfacing hidden value</article-title>
          .
          <source>International J. Computer Science and Information Security</source>
          ,
          <volume>9</volume>
          (
          <issue>5</issue>
          ),
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Limaye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarawagi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Chakrabarti</surname>
          </string-name>
          .
          <article-title>Annotating and searching Web tables using entities, types and relationships</article-title>
          .
          <source>Proc. VLDB</source>
          ,
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nestorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Motwani</surname>
          </string-name>
          .
          <article-title>Extracting schema from semistructured data</article-title>
          .
          <source>In ACM International Conference on Management of Data (SIGMOD</source>
          <year>1998</year>
          ),
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Oita</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Senellart</surname>
          </string-name>
          .
          <article-title>Own work undergoing double-blind reviewing</article-title>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Resource</given-names>
            <surname>Description</surname>
          </string-name>
          <article-title>Framework (RDF): Concepts and abstract syntax</article-title>
          .
          <source>W3C Recommendation</source>
          . http://www.w3. org/TR/2004/REC-rdf-concepts-
          <volume>20040210</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Reynaud</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Safar</surname>
          </string-name>
          .
          <article-title>Exploiting WordNet as background knowledge</article-title>
          .
          <source>In Proc. ISWC Ontology Matching (OM-07) Workshop</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Senellart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Muschick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gilleron</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tommasi</surname>
          </string-name>
          .
          <article-title>Automatic wrapper induction from hidden-Web sources with domain knowledge</article-title>
          .
          <source>In Proc. WIDM</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Stoilos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. B.</given-names>
            <surname>Stamou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Kollias</surname>
          </string-name>
          .
          <article-title>A string metric for ontology alignment</article-title>
          .
          <source>In Proc. ISWC</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>W.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. H.</given-names>
            <surname>Lochovsky</surname>
          </string-name>
          . ODE:
          <article-title>Ontology-assisted data extraction</article-title>
          .
          <source>ACM Trans. Database Syst</source>
          .,
          <volume>34</volume>
          (
          <issue>2</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Senellart</surname>
          </string-name>
          . PARIS:
          <article-title>Probabilistic alignment of relations, instances, and schema</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .,
          <volume>5</volume>
          (
          <issue>3</issue>
          ),
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          , G. Kasneci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum. YAGO</surname>
          </string-name>
          :
          <article-title>A core of semantic knowledge unifying WordNet and Wikipedia</article-title>
          .
          <source>In Proc. WWW</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Thiam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pernelle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Bennacer</surname>
          </string-name>
          .
          <article-title>Contextual and metadata-based approach for the semantic annotation of heterogeneous documents</article-title>
          .
          <source>In Proc. SeMMA</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tiezheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Derong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yue</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <article-title>Extracting result schema based on query instances in the deep Web</article-title>
          . Wuhan University J.
          <source>Natural Sciences</source>
          ,
          <volume>12</volume>
          (
          <issue>5</issue>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>F. H.</given-names>
            <surname>Lochovsky</surname>
          </string-name>
          .
          <article-title>Data extraction and label assignment for Web databases</article-title>
          .
          <source>In Proc. WWW</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lochovsky</surname>
          </string-name>
          , and W.-Y. Ma.
          <article-title>Instance-based schema matching for Web databases by domain-specific query probing</article-title>
          .
          <source>In Proc. VLDB</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <article-title>Language-independent set expansion of named entities using the Web</article-title>
          .
          <source>In Proc. ICDM</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Meng</surname>
          </string-name>
          .
          <article-title>Bootstrapping domain ontology for semantic Web services from source Web sites</article-title>
          .
          <source>In Proc. VLDB Workshop on Technologies for E-Services</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wen</surname>
          </string-name>
          .
          <article-title>Understanding the search interfaces of the deep Web based on domain model</article-title>
          .
          <source>In Proc. ICIS</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>