<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Wikipedia Bitaxonomy Explorer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tiziano Flati</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Navigli</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2004</year>
      </pub-date>
      <abstract>
        <p>We present WiBi Explorer, a new Web application developed in our laboratory for visualizing and exploring the bitaxonomy of Wikipedia, that is, a taxonomy over Wikipedia articles aligned to a taxonomy over Wikipedia categories. The application also enables users to explore and convert the taxonomic information into RDF format. The system is publicly accessible at wibitaxonomy.org and all the data is freely downloadable and released under a CC BY-NC-SA 3.0 license.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Knowledge modeling is a long-standing problem which has been addressed in a
variety of ways (see [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] for a survey). If we leave aside knowledge-lean taxonomy
learning approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], a typical and widespread model consists of knowledge
resources and multilingual dictionaries which provide concepts and relationships
between concepts. The scenario is characterized by two types of resources: those,
such as BabelNet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which provide general untyped relationships, and those,
such as DBpedia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in which edges model arbitrarily labelled predicates over
concepts (e.g., dbpedia-owl:birthPlace).
      </p>
      <p>
        In neither of these resource types, however, is any explicit attention paid to
hypernymy as a distinct relation type. Instead, hypernymy has been proven to
be a relevant relation type capable of ameliorating systems in several hard tasks
in Natural Language Processing [
        <xref ref-type="bibr" rid="ref2 ref7">2, 7</xref>
        ]. Indeed, even restricting to Wikipedia, no
high-quality, large-scale taxonomy is yet available, which exhibits high coverage
for both Wikipedia pages and categories.
      </p>
      <p>
        WiBi [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a project set up with the speci c aim of providing hypernymy
relations over Wikipedia and our tests con rm it as the best current resource
for taxonomizing both Wikipedia pages and categories in a joint fashion with
state-of-the-art results. Here we present a Web application for visualizing and
exploring our bitaxonomy of Wikipedia. The interface also o ers a customization
of the \view" and allows the export of data into RDF, in line with today's
Semantic Web trend.
WiBi [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is an approach which aims at building a bitaxonomy of Wikipedia, that
is, automatically extracting two taxonomies, one for Wikipedia pages and one
for Wikipedia categories, aligned to one another.
      </p>
      <p>
        The bitaxonomy is built thanks to a three-phase approach that i) rst builds
a taxonomy for the Wikipedia pages, then ii) leverages this partial information
to iteratively infer new hypernymy relations over Wikipedia categories while at
the same time increasing the page taxonomy, and nally iii) re nes the obtained
category taxonomy by means of three ad-hoc heuristics that cope with structural
problems a ecting some categories. As a result, a bitaxonomy is obtained where
each element - either page or category - is associated with one or more hypernyms
and where elements of one taxonomy are aligned (i.e, linked) to elements of the
other taxonomy. In order to transfer hypernymy knowledge from either one of
the two Wikipedia sides to the other side, the whole process remarkably, and as
a key feature, exploits categorization edges (here called cross-edges ) manually
provided by Wikipedians, which connect any page on one side to its categories
on the other side and vice versa. Extensive comparison has been carried out
on two datasets of 1,000 pages and categories each, against all the available
knowledge resources, including MENTA, DBpedia, YAGO, WikiTaxonomy and
WikiNet (for an extensive survey, see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). Results show that WiBi surpasses all
competitors not only in terms of quality, with the highest precision and recall,
but also in terms of coverage and speci city.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>The demo interface</title>
      <p>Here we present a Web-based visual explorer for displaying the two aligned
taxonomies of WiBi, centered on any given Wikipedia item of interest chosen
by the user. The interface easily integrates search facilities with customization
tools which personalize the experience from a user's point of view.
The home page. An excerpt of the interface's home page is shown in Fig.
1(a). As can be seen, this page has been kept very clean with as few elements
as possible. On the top of the page a navigation bar contains links to i) the
about page, which contains release information about the website content, ii) a
download area, where it is possible to obtain the data underlying the interface
and iii) the search page, which represents the core contribution of this work.</p>
      <p>The search page mainly contains a text area in which the user is requested
to input her query of interest, additionally opting for searching through either
the page inventory, the category inventory or both, thanks to dedicated radio
buttons. After the query is sent, the search engine tries to match the input text
against the whole database of Wikipedia pages (or categories) and, if a match
is found, the engine displays the nal result to the user. Otherwise, the query
is interpreted as a lemma and the user is returned with the (possible) list of all
Wikipedia pages/categories whose lemma matches against the query.
The result page. Starting from the Wikipedia element provided by the user,
the objective of the result page is to show a relevant excerpt of the bitaxonomy,
that is, the nearest (or more relevant) nodes connected to it, drawn from both
of the two taxonomies. To do this, WiBi Explorer performs a series of steps:
1. Start a DFS of maximum length 1 from the given element p of a taxonomy. As a
result, a subgraph ST1 = (SV1; SE1) is obtained;
(a) WiBi Explorer's home page.</p>
      <p>(b) Result for the ISWC Wikipedia page.
2. Collect all the nodes (p) belonging to the other taxonomy (i.e, those whose
crossedges are incident to p). Start a DFS of maximum length 2 from each element in
(p). As a result, a subgraph ST2 = (SV2; SE2) is obtained;
3. Display ST1 and ST2, as well as all the possible cross-edges linking nodes of the
two subgraphs. Prune out low-connected nodes from the displayed bitaxonomy.</p>
      <p>As a result, the interface displays a meaningful excerpt of the two taxonomies,
centered on the issued query. The result for the Wikipedia page International
Semantic Web Conference is shown in Fig. 1(b).</p>
      <p>Customization of the view Since a user might be interested in a more general
view of the bitaxonomy, two additional sliders are provided to the user in order
to manually adjust the two maximum depths 1 and 2 (see Fig. 1(b) on top).
Moreover, the interface provides the user with the capability to click on nodes
and interactively explore di erent parts of the taxonomy. The application thus
acts as a dynamic explorer that enables users to navigate through the structure
of the bitaxonomy and discover new relations as the visit proceeds.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Converting data to RDF</title>
      <p>
        Interestingly, data can also be exported in RDF format, in line with recent
work on (linguistic) linked open data and the Semantic Web [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. To this end, the
explorer is backed by the Apache Jena framework (https://jena.apache.org/)
and thus also integrates a single-click functionality that seamlessly converts the
displayed data into RDF format. The user can opt for Turtle, RDF/XML or
N-Triple format (see blue box in Fig. 1(b), bottom left). An excerpt of a view of
the bitaxonomy converted into RDF for the query ISWC is shown in Fig. 2. As
can be seen, several namespaces have been used: WiBi speci c entities encode
Wikipedia items, while standard SKOS's subsumption relations (skos:narrower
and skos:broader ) encode is-a relations.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>
        We have proposed the Wikipedia Bitaxonomy Explorer, a new, exible and
extensible Web interface that allows the navigation of the recently created Wikipedia
Bitaxonomy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In addition to default settings, several parameters concerning
the general appearance of the results can also be customized according to the
user's preferences. The demo is available at wibitaxonomy.org, it is seamlessly
integrated into the BabelNet interface (http://babelnet.org/) and the data
is freely downloadable under a CC BY-NC-SA 3.0 license.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The authors gratefully acknowledge the support of the</p>
      <p>ERC Starting Grant MultiJEDI No. 259234.</p>
      <p>The authors also acknowledge support from the LIDER project (No. 610782), a
Coordination and Support Action funded by the EC under FP7.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>DBpedia - a crystallization point for the Web of Data</article-title>
          .
          <source>Web Semantics</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <volume>154</volume>
          {
          <fpage>165</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cui</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chua</surname>
          </string-name>
          , T.S.:
          <article-title>Soft Pattern Matching Models for De nitional Question Answering</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          ) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ehrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cecconi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vannella</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mccrae</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.:
          <article-title>Representing Multilingual Data as Linked Data: the Case of BabelNet 2.0</article-title>
          .
          <source>In: Proc. of LREC 2014</source>
          . pp.
          <volume>401</volume>
          {
          <fpage>408</fpage>
          . Reykjavik, Iceland
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Flati</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vannella</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasini</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.:
          <article-title>Two Is Bigger (and Better) Than One: the Wikipedia Bitaxonomy Project</article-title>
          .
          <source>In: Proc. of ACL 2014</source>
          . pp.
          <volume>945</volume>
          {
          <fpage>955</fpage>
          .
          <string-name>
            <surname>Baltimore</surname>
          </string-name>
          , Maryland
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          :
          <article-title>Collaboratively built semi-structured content and Arti cial Intelligence: The story so far</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>194</volume>
          ,
          <issue>2</issue>
          {
          <fpage>27</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.:</given-names>
          </string-name>
          <article-title>BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>193</volume>
          ,
          <fpage>217</fpage>
          {
          <fpage>250</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Snow</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Semantic taxonomy induction from heterogeneous evidence</article-title>
          .
          <source>In: Proc. of the COLING-ACL 2006</source>
          . pp.
          <volume>801</volume>
          {
          <fpage>808</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Van</given-names>
            <surname>Harmelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Lifschitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Porter</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Handbook of knowledge representation</article-title>
          , vol.
          <volume>1</volume>
          .
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Velardi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faralli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.: OntoLearn Reloaded:
          <article-title>A graph-based algorithm for taxonomy induction</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>39</volume>
          (
          <issue>3</issue>
          ),
          <volume>665</volume>
          {
          <fpage>707</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>