<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semanti cation of Query Interfaces to Improve Access to Deep Web Content</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>German National Library of Economics Leibniz Information Centre for Economics</institution>
          ,
          <addr-line>Dusternbrooker Weg 120, 24105 Kiel</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>104</fpage>
      <lpage>111</lpage>
      <abstract>
        <p>This position paper as part of a PhD thesis is a contribution to an automatic retrieval of information from the Deep Web. Addressing current limitations of the Deep Web Information Retrieval leads to the prevailing lack of semantics regarding the retrieval process. Focusing this problem from the information providing services perspective, indicates the signi cant potential of additional semantic annotations provided by websites. Web query interfaces, the interfaces to the majority of available information on the Deep Web, are interpreted as Semantic Deep Web Services (SDWS). The introduction of a SDWS annotation leads to great potential for Information Retrieval services based on the large variety of information available on the Deep Web.</p>
      </abstract>
      <kwd-group>
        <kwd>Deep Web</kwd>
        <kwd>Semantic Deep Web Service</kwd>
        <kwd>web query interface</kwd>
        <kwd>semantic annotation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A continuously increasing amount of content on the web is not directly accessible
and indexable by search engines. The content might, for example, be hidden in
non-public, inaccessible areas or might be stored in background databases and
therefore only accessible through web query interfaces. This part of the web is
known as the Deep Web (or Hidden Web) in contrast to the Surface Web which
can be easily accessed and indexed by common search engines [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The Surface Web consists of mostly static content, which is directly
interlinked with static hyperlinks. "Search engines rely on hyperlinks to discover new
webpages [...]" [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], but static websites are outnumbered by dynamic websites on
an extremely large scale and the web has been rapidly deepened [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The content
as part of dynamic websites is mostly not accessible through static hyperlinks,
as this content is dynamically enwrapped into web pages as the response to a
query submitted through a web query interface. These are intended to be used
by human users to retrieve content from a background database often
containing highly relevant content of a speci c domain. Common search engines do not
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Arne Martin Klemenz, Klaus Tochtermann</title>
      <p>
        reach this part of the web. This is caused by the fact, that search engines "[...]
typically lack the ability to perform form submissions" [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Considering current arising services on the web like the Google Knowledge
Graph, "we can use [...] [these services] to answer questions you never thought to
ask and help you discover more"1. These services are related to Knowledge
Discovery, but in general the bene t from the automatic discovery of new knowledge
from existing information on the web is depending on an excellent Information
Retrieval. As the retrieval of information from the Deep Web is still limited,
Knowledge Discovery services are also still limited in their potential. Therefore,
more e cient and targeted retrieval mechanisms for the Deep Web are needed
to achieve full potential of Knowledge Discovery services.</p>
      <p>
        The usage of semantic annotations for information on the web play a
crucial role "to assimilate information from multiple knowledge sources" [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This
challenge has been addressed, resulting in standards like Resource Description
Framework in attributes (RDFa) and Microdata markups like schema.org
initiated by the search engine big players Bing, Google, Yahoo! and Yandex.
Therefore, this paper addresses the improvement of accessing this semantically
annotated content on the Deep Web.
2
      </p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>The retrieval and indexing of Deep Web content have been addressed from di
erent perspectives in the past. The e ort has mostly focused speci c applications
to discover, retrieve and index structured data from the Deep Web. This includes
special emphasis on the automatic web query interface interpretation.</p>
        <p>
          Common approaches focusing on exposing Deep Web content can be
classi ed to surfacing and virtual integration approaches. The surfacing approach
focuses a search engine initiated process to index the search result pages for
precomputed (randomized) queries to discover Deep Web content on large scale [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
The virtual integration approach follows the data integration paradigm, using
a mediator system to map queries to relevant web query interfaces [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The
content, that is retrieved, is brought to the user by the virtual integration to the
search result page. Both of these approaches have been approved as useful in
some cases. But in general the virtual integration approach is related to a lot of
manual e ort setting up query mapping rules for each Deep Web query interface
in the mediator system. Furthermore, the surfacing approach is too imprecise or
ine ective and therefore not scalable regarding the pre-computation of queries
for domain independent sets of Deep Web websites.
        </p>
        <p>
          Regarding the discovery and cataloging of Deep Web sources Hicks et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
highlight the challenges and demonstrate via prototype implementation, that
their Deep Web discovery framework can achieve high precision using domain
dependent knowledge for probing web query interfaces. Wenye et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] focus
"Manufacturing Deep Web Service Management [...] [by] Exploring Semantic
1 http://www.google.com/insidesearch/features/search/knowledge.html?hl=en
Web Technologies" by semantically annotating the Deep Web Services to
reect their hidden, dynamic, and heterogeneous contents while the relevance of
semantic annotations for the Deep Web has already been identi ed in 2003 by
Handschuh et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Whereas these publications as well as Chun et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] discuss
theses challenges from the information retrieving services perspective this paper
will set the focus to the information providing services perspective.
        </p>
        <p>
          Furche et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] introduced a promising automated form understanding
ontology based approach, which is far beyond heuristics to ll out search forms [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
combining "[...] signals from the text, structure, and visual rendering of a web
page". But according to Li, Xian et al. in "Truth Finding on the Deep Web: Is
the Problem Solved?" [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] the challenges arising from the Deep Web are regarded
as not yet solved. In general, current approaches are still limited either in being
domain speci c or limited in their e ciency.
        </p>
        <p>
          Until today, there still exists no general domain independent solution for
the Deep Web Information Retrieval problem. Just a fraction of total available
data in background databases may be covered by common state of the art
approaches. This is particularly due to the fact, that for large data sets there
exist nearly endless possible permutations of search results. This especially
applies to the retrieval of dynamic content. Therefore, it seems to be improbable
to improve retrieval and indexing mechanisms towards reaching a 100%
coverage of all available Deep Web content. Consequently, this is not the focus of
our current research. Currently still limited mechanisms have already "[...]
succeeded largely by targeting narrow domains where a search application can be
ne-tuned to query a relatively small number of databases and return highly
targeted results" [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. For that reason, we focus e.g. on the reduction of manual
e ort regarding the query mapping on the one hand and more precise query
generation or pre-computation for the targeted retrieval from broader domains on
the other hand. Therefore, this paper is intended to improve access to Deep Web
content by providing great potential for new Information Retrieval mechanisms
and for the signi cant improvement of previously existing mechanisms.
3
3.1
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Approach</title>
        <sec id="sec-2-2-1">
          <title>Research Focus</title>
          <p>To step forward towards a Semantic Deep Web, which is the superordinated
long-term objective, it is necessary to focus on additional research questions
resulting from previously identi ed limitations. For the targeted retrieval
especially of dynamic Deep Web content, the need of an e cient and in an ideal
case fully automatic approach is essential. Therefore, the focus needs to be set
to these challenges: content providing service Discovery, Invocation &amp; Execution
and Composition. Addressing these challenges will ensure the discovery of
appropriate web query interfaces providing access to relevant content (! Discovery ),
the appropriate query mapping and query submission (! Invocation &amp;
Execution) and the service interoperability (! Composition). Common approaches
for Deep Web Information Retrieval focus these challenges from the information
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Arne Martin Klemenz, Klaus Tochtermann</title>
      <p>retrieving services perspective. The conceptual idea being introduced in this
section focuses these challenges from the information providing services perspective.</p>
      <p>Common semantic annotation standards like RDFa and schema.org
microdata address particularly the annotation of web content and do not have means
for the prevailing lack of semantics at the crucial point of the Deep Web
Information Retrieval process. This crucial point is regarding the web query interfaces.
To improve common crawling, indexing and content retrieval mechanisms and to
ensure new mechanisms, a semantic annotation for query interfaces is suggested.
This will reuse the query interfaces originally intended for human users in a
combined computer and human readable format. The abstract concept, to describe
query interfaces in a computer readable format, is derived from the semantic
annotation of Web Services. Standards like Semantic Annotations for WSDL
and XML Schema (SAWSDL) provide a machine readable Web Service
annotation describing the functionality and retrievable data. A semantic annotation for
query interfaces will provide machine readable information for henceforth called
Semantic Deep Web Services (SDWS).
3.2</p>
      <sec id="sec-3-1">
        <title>Semantic Deep Web Service annotation</title>
        <p>An implementation of the SDWS annotation should meet the following
fundamental criteria: SDWS interface semantics, providing a generalization of SDWS
interfaces (! abstract ) with the ability to include own vocabularies as for
example thesauri (! extendable). Going more into details, a SDWS annotation
prototype should provide information about general properties regarding the
content that is provided by the SDWS (! content properties) and concrete
interface eld properties to describe the semantic structure and internal structural
dependencies of the SDWS interfaces (! eld properties).</p>
        <p>The prototype SDWS content properties describe the content domain of the
retrievable information, the content language, as well as the content type. The
content type attribute may be described based on schema.org microdata and
the supplementary usage of other vocabularies. An additional content property
might provide information about the amount of available data (property: count ).
These content properties are just the extendable basis for this prototype
providing general information about the retrievable content. Further ideas for the
extension of SDWS content properties will be discussed in the following section.</p>
        <p>A simple example for the SDWS content properties is provided in Fig. 1,
describing the basic SDWS eld properties for the interface of the subject portal
EconBiz 2. EconBiz provides highly relevant content for the domain of economics
and business studies and access to speci c content types (various types of
CreativeWork and information about Event s).</p>
        <p>The prototype SDWS eld properties describe the eld type (e.g. selectField,
inputField ), as well as the input domain and output range of each particular
SDWS interface eld. The input domain attribute describes valid input values of
a speci ed SDWS eld. Furthermore, it is a trigger for the output range attribute,
2 http://www.econbiz.de/en/
Semanti cation of Query Interfaces to Improve Access to Deep Web Content
5
highly relevant content for a specific domain
contentDomain: #economics [AND]</p>
        <p>#businessStudies
contentType: http://schema.org/CreativeWork [AND]</p>
        <p>http://schema.org/Event
contentLanuage: #multilingual
contentCount: 8.913.444
as its input value de nes the restriction set for the retrieval process at time of
form submission (examples, Fig. 2-4). Additionally, a vocabulary attribute may
reference for instance a thesaurus that can be used as suggest-value vocabulary
for the particular domain to ensure a targeted retrieval.</p>
        <p>The basic SDWS eld properties example in Fig. 2 refers to a standardized
vocabulary, the STW Thesaurus for Economics. The STW provides "vocabulary
on any economic subject" containing "[...] more than 6,000 standardized subject
headings and about 19,000 entry terms to support individual keywords"3. This
thesaurus is the basis for the annotation on metadata level in EconBiz and will
therefore ensure a targeted retrieval. The bene t of the vocabulary property
will especially apply to digital libraries but also to other domains. Therefore,
regarding simple SDWS interfaces, this might be one of the most appropriate
use cases for SDWS eld properties as there is no complex interface structure.
fieldType: inputField
fieldDomain: xsd:String
fieldRange: http://schema.org/CreativeWork#all [AND]</p>
        <p>http://schema.org/Event#all
fieldVocabulary: #STWthesaurus</p>
        <p>Focusing on more complex SDWS interfaces, the example in Fig. 3 contains
chunks of related elds that a ect each other. The rst selectField as part of the
marked chunk de nes the relation to the other chunks. The second selectField
as part of this chunk de nes the input eld domain and restricts the input eld
range of the inputField that is part of the focused chunk. Furthermore, Fig. 4
considers some exemplary e ects triggered by the selection of di erent select
values of the second selectField within the focused chunk in Fig. 3.</p>
        <p>More complex examples as illustrated in Fig. 3 and 4 demonstrate that the
semantic meaning behind a SDWS interface might be quite complex and
automated form understanding approaches will quickly reach their limits. Especially
3 http://zbw.eu/stw/versions/latest/about.en.html
6</p>
        <p>Arne Martin Klemenz, Klaus Tochtermann
chunk of related fields that affect each other
fieldType: selectField
fieldDomain: #predefined (xsd:boolean)
fieldRange: define #chunkRelation
fieldType: selectField
fieldDomain: #predefined (xsd:string)
fieldRange: define #inputFieldDomain</p>
        <p>restrict #inputFieldRange
the automated detection of related elds and the detection of complex relations
within chunks might be the most di cult part, where common approaches fail.</p>
        <p>In addition to the SDWS interface annotation, it is advisable to link every
SDWS interface from the websites root index. This will ensure a targeted SDWS
discovery and can be realized by using XML Sitemaps to de ne a SDWS retrieval
index. The SDWS annotation itself is suggested to be embedded directly to each
particular SDWS interfaces.
4</p>
        <sec id="sec-3-1-1">
          <title>Bene t and further Use Cases</title>
          <p>
            The introduced SDWS annotation will lead to great potential for new
information retrieval mechanisms and plays a signi cant role for the improvement of
current mechanisms. Queries might for example be automatically mapped to
various SDWS at the same time based on the SDWS annotation (! abstract
semantic querying ). In accordance with Heath et al. the vision of "users [...]
interacting with the Web as a data space" [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] will therefore also bene t from the
introduced SDWS annotation. In general, ensuring new user oriented services
especially new Knowledge Discovery services based on the large variety of
available information on the Deep Web is one of the major purposes. Furthermore,
the SDWS annotation might also be used for purposes, not directly focusing
on the retrieval itself, as for example reasoning processes for client side query
interface input validations. Another use case focuses content licensing issues as
these are a problematic topic for digital libraries. These may be addressed by
adding licensing information directly to the SDWS interfaces by extending the
introduced annotation prototype. This will be an appropriate possibility to
pro
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Semanti cation of Query Interfaces to Improve Access to Deep Web Content 7</title>
      <p>vide licensing information exactly at that point where the information itself is
being provided e.g. based on the Creative Commons licensing model.</p>
      <p>Overall, this approach will make webmasters aware of their responsibility
to add SDWS annotations to SDWS interfaces in addition to current semantic
content annotations. This process requires additional e ort on the one hand,
but on the other hand it also enables the webmasters to control the information
content that may be retrieved by various retrieving services like search engines.
For now webmasters may only use common HTML attributes like nofollow or
noindex and the Robots Exclusion Standard to control the crawling behavior on
their websites. The SDWS annotation ensures the targeted in uence of the
webmaster. Furthermore, only the webmaster knows the exact semantic statement
intended by the implemented SDWS interface. Regarding web content, search
engines rely on semantic content markups as it is more reliable than current
automatic content interpretation approaches. Therefore, it is obvious, that this
will also apply to the annotation of SDWS interfaces.
5</p>
      <sec id="sec-4-1">
        <title>Conclusion</title>
        <p>This paper addressed the lack of semantic information regarding web query
interfaces in the process of Information Retrieval from the Deep Web.
Transferring the concepts of semantic web content annotations on the one hand and
Semantic Web Service Descriptions on the other hand, leads to the great
potential of semantic annotations for SDWS interfaces. Equivalent to semantic web
content annotations, the SDWS annotation provides an unambiguous semantic
interpretation of the SDWS interface. A variety of current information retrieval
mechanisms and form understanding systems try to analyze SDWS interfaces
automatically by focusing the Deep Web Information Retrieval challenge from
the retrieving services perspective. Instead of relying on these, the introduced
SDWS interface annotation is focusing this challenge from the information
providing services perspective.</p>
        <p>
          In general, this approach follows the open knowledge sharing paradigm as
part of the Semantic Web vision from Berners-Lee et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. This is based on the
assumption, that the information provided on websites is intended to be retrieved
by various services. Any additional licensing issues restricting the retrieval and
further usage of the retrievable information have also been addressed.
        </p>
        <p>This approach will contribute to domain independent and automatic
Information Retrieval mechanisms based on the introduced SDWS annotation.
Manual e ort for currently still limited Deep Web Information Retrieval mechanisms
will be reduced or even eliminated. Furthermore, these retrieval mechanisms will
bene t regarding their e ciency and can be adapted targeting broader domains.
6</p>
      </sec>
      <sec id="sec-4-2">
        <title>Future Work</title>
        <p>Future work will especially focus on the critical evaluation based on further
research studies. The de nition of a concrete SDWS annotation syntax based</p>
        <p>Arne Martin Klemenz, Klaus Tochtermann
on the usage of existing annotation standards will concern the challenge how
retrieving services will learn to understand the SDWS annotation. Reduction
of manual e ort for the annotation process also requires further e ort. A
semiautomatic generation process may provide support for the de nition of SDWS
annotations. This process may be based on sampling and probing the background
database utilizing promising automated form understanding approaches. This
may lead to semi-automatic generation approaches for the SDWS annotation.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bergman</surname>
            ,
            <given-names>M. K.</given-names>
          </string-name>
          <article-title>White paper: The deep web: Surfacing hidden value</article-title>
          .
          <source>the journal of electronic publishing 7</source>
          ,
          <issue>1</issue>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lassila</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , et al.
          <article-title>The semantic web</article-title>
          .
          <source>Scienti c American</source>
          <volume>284</volume>
          ,
          <issue>5</issue>
          (
          <year>2001</year>
          ),
          <volume>28</volume>
          {
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chun</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Warner</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Semantic annotation and search for deep web services</article-title>
          .
          <source>In E-Commerce Technology and the Fifth IEEE Conference on Enterprise Computing</source>
          ,
          <article-title>E-Commerce and</article-title>
          <string-name>
            <surname>E-Services</surname>
          </string-name>
          ,
          <year>2008</year>
          10th IEEE Conference on (
          <year>2008</year>
          ), IEEE, pp.
          <volume>389</volume>
          {
          <fpage>395</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Furche</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gottlob</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grasso</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orsi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Schallhart</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Opal: Automated form understanding for the deep web</article-title>
          .
          <source>In Proceedings of the 21st international conference on World Wide Web</source>
          (
          <year>2012</year>
          ), ACM, pp.
          <volume>829</volume>
          {
          <fpage>838</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Annotation for the semantic web</article-title>
          , vol.
          <volume>96</volume>
          . IOS Press,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. C.-C.</surname>
          </string-name>
          <article-title>Accessing the deep web</article-title>
          .
          <source>Communications of the ACM 50</source>
          ,
          <issue>5</issue>
          (
          <year>2007</year>
          ),
          <volume>94</volume>
          {
          <fpage>101</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Semantic annotation and retrieval: Web of data</article-title>
          .
          <source>Handbook of Semantic Web Technologies</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hicks</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scheffer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngu</surname>
            ,
            <given-names>A. H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sheng</surname>
            ,
            <given-names>Q. Z.</given-names>
          </string-name>
          <article-title>Discovery and cataloging of deep web sources</article-title>
          .
          <source>In Information Reuse and Integration (IRI)</source>
          ,
          <year>2012</year>
          IEEE 13th International Conference on (
          <year>2012</year>
          ), IEEE, pp.
          <volume>224</volume>
          {
          <fpage>230</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lyons</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Truth nding on the deep web: Is the problem solved</article-title>
          ?
          <source>In Proceedings of the 39th international conference on Very Large Data Bases</source>
          (
          <year>2012</year>
          ), VLDB Endowment, pp.
          <volume>97</volume>
          {
          <fpage>108</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Madhavan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Afanasiev</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A. Harnessing</given-names>
          </string-name>
          <article-title>the deep web: Present and future</article-title>
          .
          <source>4th Biennial Conference on Innovative Data Systems Research (CIDR)</source>
          (
          <year>Jan</year>
          .
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Madhavan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganapathy</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rasmussen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Google's deep web crawl</article-title>
          .
          <source>Proceedings of the VLDB Endowment 1</source>
          ,
          <issue>2</issue>
          (
          <year>2008</year>
          ),
          <volume>1241</volume>
          {
          <fpage>1252</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Masanes</surname>
            ,
            <given-names>J. Archiving</given-names>
          </string-name>
          <article-title>the hidden web</article-title>
          .
          <source>In Web Archiving</source>
          . Springer,
          <year>2006</year>
          , pp.
          <volume>115</volume>
          {
          <fpage>129</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mukherjea</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Information retrieval and knowledge discovery utilising a biomedical semantic web</article-title>
          .
          <source>Brie ngs in Bioinformatics 6</source>
          ,
          <issue>3</issue>
          (
          <year>2005</year>
          ),
          <volume>252</volume>
          {
          <fpage>262</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ograph</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amanca</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Maahs</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Searching the deep web</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>51</volume>
          ,
          <issue>10</issue>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Wenyu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jianwei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ming</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jian</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lanfen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Manufacturing deep web service management: Exploring semantic web technologies</article-title>
          .
          <source>Industrial Electronics Magazine, IEEE 6</source>
          ,
          <issue>2</issue>
          (
          <year>2012</year>
          ),
          <volume>38</volume>
          {
          <fpage>51</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>