<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Integrated Retrieval over Structured and Unstructured Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Qiuyue Wang</string-name>
          <email>qiuyuew@ruc.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jinglin Kang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Key Lab of Data Engineering and Knowledge Engineering</institution>
          ,
          <addr-line>MOE, Beijing 100872</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information, Renmin University of</institution>
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We report our experiment results on the INEX 2012 Linked Data Track. We participated in the ad hoc and jeopardy tasks. As the new data collection on INEX 2012 Linked Data Track features a combination of unstructured and structured data, our first attempt is to investigate different strategies of combining the retrievals over structured and unstructured data, and compare the combined approaches with the traditional unstructured ones. In this paper, we discussed three types of combination strategies and we experimented two of them on the track. The experiment results show that ……</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Though Web is best known as an enormous collection of unstructured documents, it
also contains a huge amount of structured data, like HTML tables, data stored in Deep
Web databases, increasingly published RDF data due to the efforts of Linked Data
community, and so on. With more and more structured data became accessible to end
users, more intelligent search on the Web is expected. There are increasing interests
on semantic search on the Web, i.e. leveraging the semantics in structured data to
improve the Web search.</p>
      <p>
        The new data collection of INEX 2012 Linked Data Track is a fusion of Wikipedia
articles and their corresponding RDF data from DBpedia and YAGO2. Each
Wikipedia article corresponds to an entity/resource in DBpedia and YAGO2, while
DBpedia and YAGO2 contain structured data extracted from each article, e.g.
properties of entities and relationships with other entities. It can be viewed as an
integrated collection of unstructured and structured data covering a wide range of
topics. With such a data collection, we intended to investigate different strategies of
combining retrievals over unstructured and structured data so that the performance
would be better than that of unstructured retrieval or structured retrieval only.
Basically, there are three types of combination strategies. (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Parallel combination.
Retrieve the structured and unstructured data separately, and then combine the two
result lists. (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) Unstructured-structured serial combination. Retrieve the
unstructured data first. The top-k results, which correspond to entity nodes in the RDF
graph, then spread their activations over the RDF graph. Thus, some relevant results
which do not contain query terms may be retrieved. (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Structured-unstructured
serial combination. Retrieve the structured data first. The top-k returned entities or
subgraphs are then analyzed so that the original query could be better understood. For
example, the query is expanded with more effective terms, or is reformulated in terms
of related entities and so on. The newly transformed query is then used to retrieve the
unstructured data more accurately.
      </p>
      <p>Due to the limit of time, we only experimented on the first two strategies. Firstly,
we indexed the unstructured and structured data separately, and used language
modeling approaches to retrieve them respectively. We treat the unstructured run as
our baseline. Then we combined the unstructured and structured runs using weighted
sum approach. Secondly, we used the unstructured run as the input to the algorithm of
spreading actions on the RDF graph, and submitted the new ranked list of results after
spreading activation.</p>
      <p>The results show that ……
2</p>
    </sec>
    <sec id="sec-2">
      <title>Combined Retrieval Strategies</title>
      <p>In this section, we first present the retrieval models that we used to retrieve
unstructured data and structured data respectively, and then discuss different
strategies of combining the retrievals over unstructured and structured data.</p>
      <p>Given a keyword query and unstructured document collection, we can employ any
traditional IR models to retrieve relevant documents. In this paper, we use the
language modeling approach since it has the state of art performance among other
retrieval models.</p>
      <p>For a structured data collection, there are various approaches proposed to look for
relevant answers for a keyword query [1]. One of the key problems in structured
retrieval is that return units are not predefined as in document retrieval. There are
various ways to generate all possible results. Then the results are ranked using either
traditional content-based models, e.g. TF-IDF, vector space model, or
contentstructure-based ranking models. However, there is still lack of a general evaluation
campaign for comparing all these retrieval models for keyword search on structured
data. So it is very hard to draw any conclusions on these various approaches. In this
paper, we simply define the retrieval units of structured retrieval to be entities, which
actually correspond to Wikipedia articles in the collection of Linked Data Track. To
retrieve entities on RDF graphs, we first aggregate all information about an entity
together, i.e. the entity’s properties, subjects, objects, etc., and index it as a pseudo
document identified by the entity’s ID. Then we employ the language modeling
approach to rank each pseudo document with respect to the given keyword query.
2.1
2.2</p>
      <sec id="sec-2-1">
        <title>Parallel Combination</title>
      </sec>
      <sec id="sec-2-2">
        <title>Unstructured-Structured Serial Combination</title>
        <p>3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Results</title>
      <p>Due to the limit of time, we only experimented on the first two strategies. In this
section, we discuss the experiment results on the INEX 2012 Linked Data Track.
3.1</p>
      <sec id="sec-3-1">
        <title>Implementation</title>
        <p>We indexed the unstructured and structured data separately both using Indri with
Krovetz stemmer and a short stop word list {a, about, an, and, as, at, by, in, of, on, or,
that, the, to}. Remember that we generate a pseudo document for each entity in the
structured data set, and index this pseudo document for the entity.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results</title>
        <p>The evaluation results have not been released by the time when the author wrote this
abstract.
4
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The research work is supported by the “HGJ” National Science and Technology
Major Project of China under Grant No. 2010ZX01042-001-002-002.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Keyword search on structured and semi-structured data</article-title>
          .
          <source>In Proceedings of the 2009 ACM SIGMOD International Conference on Management of data (SIGMOD '09)</source>
          ,
          <source>Carsten Binnig and Benoit Dageville (Eds.)</source>
          . ACM, New York, NY, USA,
          <fpage>1005</fpage>
          -
          <lpage>1010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schwabe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Aragao</surname>
          </string-name>
          ,
          <article-title>A hybrid approach for searching in the semantic web</article-title>
          .
          <source>In Proceedings of the 13th international conference on World Wide Web (WWW '04)</source>
          . ACM, New York, NY, USA,
          <fpage>374</fpage>
          -
          <lpage>383</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lopez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Uren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vallet</surname>
          </string-name>
          , E. Motta, and
          <string-name>
            <given-names>P.</given-names>
            <surname>Castells</surname>
          </string-name>
          ,
          <article-title>Semantic Search Meets the Web</article-title>
          .
          <source>In Proceedings of the 2008 IEEE International Conference on Semantic Computing (ICSC '08)</source>
          . IEEE Computer Society, Washington, DC, USA,
          <fpage>253</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>