<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Descriptive Schema: Semantics-based Query Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>S. D. Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Yee</string-name>
          <email>kcyee@cs.hku.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David W. Cheung</string-name>
          <email>dcheung@cs.hku.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wenjun Yuan</string-name>
          <email>wjyuan@cs.hku.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, The University of Hong Kong</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We propose the novel concept of “descriptive schema” (DS). Unlike ordinary database schemas, a DS does not restrict the structure of the underlying database. Rather, it is just a probabilistic description of the structure. When answering keyword queries, DS can be used to improve semantics-based query answering and result ranking. Schema: To have or not to have? Wikipedia is a rich repository of information. However, facilities to exploit the information are still limited. Although typical search WWW search engines such as Google[1] allow users to look for information using keywords, they lack a schema for formulating the queries precisely. Besides hyperlinks among the Wikipedia pages, many pages have Category tags as well as Infoboxes, which can be exploited to perform more sophisticated searches. For example, the DBpedia community makes use of these tags to build a database of RDF triplets, allowing more expressive and precise queries in the form of SPARQL to be used to retrieve useful information [2]. The above are two extremes of search and query. In the former case, the user can perform a search easily using relevant keywords, without having to learn the schema's lexicon beforehand. In the latter case, a schema can be used to help specify the query more precisely, but it has a non-trivial learning curve. In this paper, we propose the approach of “descriptive schema” to address these shortcomings. We attempt to strike a balance between the ease of use of a schema-less approach and the high accuracy that a schema-based system can bring us.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In this paper, we propose a new concept called “Descriptive Schema” (DS).
Unlike XSD (XML Schema Definition), DS is not meant to prescriptively mandate
a structure on the underlying data. We want to retain the flexibility of free
format for the pages. Rather, DS, as its name implies, is descriptive. It is only
a summary of the structure exhibited by the underlying database. It does not
define the structure. The data may occasionally violate the DS.</p>
      <p>This tolerance to violations marks our biggest innovation, contrasting with
existing approaches. Existing approaches to data modelling use “Prescriptive
Schema”, which mandates a rigid structure on the underlying data, with little
(if any) tolerance to violations.</p>
      <p>We model a DS by a set of rules on the underlying data. There are many
possible ways to formulate the rules. One example rule is: “90% of the time, a
page of class ‘Countries’ has value for the field ‘capital’ in the infobox (infobox
for countries)”. Note that the rules defined in this way are probabilistic, because
they are not satisfied all the time. A DS may thus be considered a summary of
the patterns occurring in a database, instead of policies imposed on the data.</p>
      <p>
        The task of discovering a DS from a database is a mining task, which is
the problem of finding all rules satisfying a the specified syntax and support
thresholds, thus following the data mining model in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Applications</title>
      <p>Since a DS captures semantical information about the underlying data, it enables
a semantics-based approach to answering search queries. We can, for instance,
use the DS to help us disambiguate the query, enrich the query with
semantical information, as well as using the semantical information to rank the search
results. Applications of DS include, but are not limited to, the following:
– Keyword Disambiguation
– Query Augmentation
– Result Ranking
– Data Cleansing
– Guidelines for Authors
– Guided Query Building
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>We have proposed the concept of “descriptive schemas”, which is a set of rules
obeyed by most of the underlying data, with tolerance for violations. Although
the primary goal of devising this novel concept was to help answering keyword
queries with an accuracy comparable to databases with prescriptive schemas,
we have realized that DS can also be useful for other applications. Future works
include exploring further potentials of DS, developing a formalism for it,
devising efficient algorithms for mining DS, as well as more in-depth studies of the
applications mentioned in this paper.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The anatomy of a large-scale hypertextual web search engine</article-title>
          .
          <source>Computer Networks</source>
          <volume>30</volume>
          (
          <issue>1-7</issue>
          ) (
          <year>1998</year>
          )
          <fpage>107</fpage>
          -
          <lpage>117</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
          </string-name>
          , J.:
          <article-title>What have Innsbruck and Leipzig in common? extracting semantics from Wiki content</article-title>
          . In Franconi, E.,
          <string-name>
            <surname>Kifer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>May</surname>
          </string-name>
          , W., eds.
          <source>: ESWC</source>
          . Volume
          <volume>4519</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2007</year>
          )
          <fpage>503</fpage>
          -
          <lpage>517</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mannila</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toivonen</surname>
          </string-name>
          , H.:
          <article-title>Levelwise search and borders of theories in knowledge discovery</article-title>
          .
          <source>Data Min. Knowl. Discov</source>
          .
          <volume>1</volume>
          (
          <issue>3</issue>
          ) (
          <year>1997</year>
          )
          <fpage>241</fpage>
          -
          <lpage>258</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>