<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Utilizing, creating and publishing Linked Open Data with the Thesaurus Management Tool PoolParty</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Schandl</string-name>
          <email>schandl@punkt.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Blumauer</string-name>
          <email>blumauer@punkt.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thesaurus</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Personal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>punkt. NetServices GmbH</institution>
          ,
          <addr-line>Lerchenfelder Gürtel 43, 1160 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We introduce the Thesaurus Management Tool (TMT) PoolParty based on Semantic Web standards that reduces the effort to create and maintain thesauri by utilizing Linked Open Data (LOD), text-analysis and easy-to-use GUIs. PoolParty's aim is to lower the access barriers to managing thesauri, so domain experts can contribute to thesaurus creation without needing knowledge about the Semantic Web. A central feature of PoolParty is the enriching of a thesaurus with relevant information from LOD sources. It is also possible to import and update thesauri from remote LOD sources. Going a step further we present a Personal Information Management tool built on top of PoolParty which relies on Open Data to assist the user in the creation of thesauri by suggesting categories and individuals retrieved from LOD sources. Additionally PoolParty has natural language processing capabilities enabling it to analyse documents in order to glean new concepts for a thesaurus and several GUIs for managing thesauri varying in their complexity. Thesauri created with PoolParty can be published as Open Knowledge according to LOD best practices.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Web</kwd>
        <kwd>Linking Information Management</kwd>
        <kwd>SKOS</kwd>
        <kwd>RDF</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Thesauri have been an important tool in Information Retrieval for decades and still
are [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. While they have the potential to greatly improve the information management
of organisations, professionally managed thesauri are rarely used in content
management systems, search engines or tagging systems.
      </p>
      <p>Important reasons frequently given for this are: (1) the difficulty of learning and
using TMT, (2) the lacking possibilities to integrate TMTs into existing information
systems, (3) it is laborious to create and maintain a thesaurus, and while TMTs often
support either automatic or manual methods to maintain a thesaurus they rarely
combine those two approaches, and (4) users don't have enough knowledge about
thesaurus building methodologies and/or worthwhile use cases utilizing semantic
knowledge models like SKOS thesauri.</p>
      <p>The TMT PoolParty1 addresses the first three issues. A central goal is to ease the
process of creating and maintaining thesauri by domain experts, that don’t have a
strong technical background, don’t know about semantic technologies and maybe
know little about thesauri. We see an important role for Linked Open Data in this area
and equipped PoolParty with the capability to enrich one’s own knowledge model
with relevant information from the LOD cloud. In combination with several GUIs
suited for varying levels of complexity PoolParty allows for low access barriers for
creating and utilizing thesauri and Open Data.</p>
      <p>PoolParty is a commercial application, but will have a version that can be used free
of charge. We are working on such a version that makes use of the Talis platform2. In
any version the user will have the option to publish thesauri as LOD and license them
under various Creative Commons licenses.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Use Cases</title>
      <p>
        PoolParty is based on Semantic Web technologies like RDF3 and SKOS4 (Simple
Knowledge Organisation System) allowing for multilingual thesauri to be represented
in a standardised manner [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While OWL5 would offer greater possibilities in
creating knowledge models, it is deemed too complex for the average information
worker.
      </p>
      <p>
        PoolParty was conceived to facilitate various commercial and non-commercial
applications for thesauri. In order to achieve this, it needs to publish them and offer
methods of integrating them with various applications [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In PoolParty this can be
realized on top of its RESTful web service interface providing thesaurus management,
indexing, search, tagging and linguistic analysis services.
      </p>
      <p>Some of these (semantic) web applications are:
- Semantic search engines
- Recommender systems (similarity search)
- Corporate bookmarking
- Annotation- &amp; tag recommender systems
- Autocomplete services and facetted browsing.
- Personal Information Management</p>
      <p>These use cases can be either achieved by using PoolParty stand-alone or by
integrating it with existing (Enterprise) Search Engines and Document Management
Systems.</p>
      <sec id="sec-2-1">
        <title>1 http://poolparty.punkt.at/</title>
        <p>2 http://www.talis.com/platform/
3 http://www.w3.org/RDF/
4 http://www.w3.org/2004/02/skos
5 http://www.w3.org/TR/owl-ref/
PoolParty is written in Java and uses the SAIL API6, whereby it can be utilized with
various triple stores, which allows for flexibility in terms of performance and
scalability.</p>
        <p>Thesaurus management itself (viewing, creating and editing SKOS concepts and
their relationships) can be done in an AJAX Frontend based on Yahoo User Interface
(YUI). Editing of labels can alternatively be done in a Wiki style HTML frontend.</p>
        <p>For key-phrase extraction from documents PoolParty uses a modified version of
the KEA7 5 API, which is extended for the use of controlled vocabularies stored in a
SAIL Repository (this module is available under GNU GPL). The analysed
documents are locally stored and indexed in Lucene8 along with extracted concepts
and related concepts.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 Thesaurus Management with PoolParty</title>
      <p>The main thesaurus management GUI of PoolParty (see Fig. 1) is entirely web-based
and utilizes AJAX to e.g. enable the quick merging of two concepts either via drag &amp;
drop or autocompletion of concept labels by the user. An overview over the thesaurus
can be gained with a tree or a graph view of the concepts.
Consistent with PoolParty's goal of relieving the user of burdensome tasks while
managing thesauri doesn't end with a comfortable user interface: PoolParty helps to
semi-automatically expand a thesaurus as the user can use it to analyse documents
(e.g. web pages or PDF files) relevant to her domain in order to glean candidate terms
for her thesaurus. This is done by the key-phrase extractor of KEA. The extracted</p>
      <sec id="sec-3-1">
        <title>6 http://www.openrdf.org/doc/sesame2/system/ch05.html 7 http://www.nzdl.org/Kea/index.html 8 http://lucene.apache.org/</title>
        <p>terms can be approved by the user, thereby becoming "free concepts" which later can
be integrated into the thesaurus, turning them into "approved concepts".</p>
        <p>Documents can be searched in various ways – either by keyword search in the full
text, by searching for their tags or by semantic search. The latter takes not only a
concept's preferred label into account, but also its synonyms and the labels of its
related concepts are considered in the search. The user might manually remove query
terms used in semantic search. Boost values for the various relations considered in
semantic search may also be adjusted. In the same way the recommendation
mechanism for document similarity calculation works.</p>
        <p>PoolParty by default also publishes an HTML Wiki version of its thesauri, which
provides an alternative way to browse and edit concepts. Through this feature anyone
can get read access to a thesaurus, and optionally also edit, add or delete labels of
concepts. Search and autocomplete functions are available here as well.</p>
        <p>The Wiki’s HTML source is also enriched with RDFa, thereby exposing all RDF
metadata associated with a concept to be picked up the RDF search engines and
crawlers.</p>
        <p>PoolParty supports the import of thesauri in SKOS (in serializations including
RDF/XML, N-Triples or Turtle) or Zthes format.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6 Linked Open Data Capabilities</title>
      <p>
        PoolParty not only publishes its thesauri as Linked Open Data (additionally to a
SPARQL endpoint), but it also consumes LOD in order to expand thesauri with
information from LOD sources. Concepts in the thesaurus can be linked to e.g.
DBpedia9 via the DBpedia lookup service [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which takes the label of a concept and
returns possible matching candidates. The user can select the DBpedia resource that
matches the concept from his thesaurus, thereby creating an owl:sameAs relation
between the concept URI in PoolParty and the DBpedia URI. The same approach can
be used to link to other SKOS thesauri available as Linked Data.
      </p>
      <p>Other triples can also the retrieved from the target data source, e.g. the DBpedia
abstract can become a skos:definition and geographical coordinates can be imported
and be used to display the location of a concept on the map, where appropriate. The
DBpedia category information may also be used to retrieve additional concepts of that
category as siblings of the concept in focus, in order to populate the thesaurus.</p>
      <p>PoolParty is not only capable of importing a SKOS thesaurus from a Linked Data
server, it may also receive updates to thesauri imported this way. This feature has
been implemented in the course of the KiWi10 project funded by the European
Commission. KiWi also contains SKOS thesauri and exposes them as LOD. Both
systems can read a thesaurus via the other’s LOD interfaces and may write it to their
own store. This is facilitated by special Linked Data URIs that return e.g. all the
topconcepts of a thesaurus, with pointers to the URIs of their narrower concepts, which
allow other systems to retrieve a complete thesaurus through iterative dereferencing
of concept URIs.</p>
      <sec id="sec-4-1">
        <title>9 http://dbpedia.org/ 10 http://kiwi-project.eu/</title>
        <p>Additionally KiWi and PoolParty publish lists of concepts created, modified,
merged or deleted within user specified time-frames. With this information the
systems can learn about updates to one of their thesauri in an external system. They
then can compare the versions of concepts in both stores and may write according
updates to their own store.</p>
        <p>This means each system decides autonomously which data it accepts and there is
no risk of a system pushing data that might lead to inconsistencies into an external
store. Data transfer and communication are achieved using REST/HTTP, no other
protocols or middleware are necessary. Also no rights management for each external
systems is needed, which otherwise would have to be configured separately for each
source.</p>
        <p>The synchronisation process via Linked Data will be improved in the ongoing
KiWi project. We will implement an update and conflict resolution dialogue through
which a user may decide which updates to concepts to accept and to consequently
write to the system’s store.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7 Personal Information Management utilizing Linked Open Data</title>
      <p>An example application that we are currently developing on top of PoolParty web
services is a Personal Information Manager (PIM) utilizing Open Data.</p>
      <p>Our goal is to enable users to create and utilize thesauri without requiring any
knowledge about thesauri. We aim to hide the complexity of thesauri and their
polyhierarchical structure and concentrate on presenting the user with listboxes filled with
terms from a thesaurus or LOD sources.</p>
      <p>The PIM will be a web based application that makes use of Linked Open Data in
order to assist the users with suggestions when they e.g. create categories. A movie
expert for example might want to create a knowledge model of filmmakers and the
countries they lived in. Upon the creation of a new project, the PIM asks the user to
specify a general domain it is about (people, places, things, events, organisations,
etc.). After the user selects "people", the system can use e.g. data about categories
from YAGO11, UMBEL12 or DBpedia that relate to "people" to help refine the user's
domain. The system might suggest popular categories or the user can start entering a
string like "fil" prompting the system’s autocomplete functionality to suggest the
DBpedia class "filmmaker". When the user confirms that this is one of the topics his
project is about and finishes the same process for the other topics (i.e. countries),
several links between the local model and the LOD cloud exist and more specific
information can be retrieved to assist the user's work on the thesaurus. The system
might suggest possible relations between painters and cities like "made movie in",
"was born in" or "lived in", and the user can specify which particular relations are of
interest to him.</p>
      <p>In a similar way the PIM will assist with creating instance data and suggest
particular filmmakers and countries (see mock-up in Fig. 2) from sources like
DBpedia that belong to the corresponding classes. In this way the PIM not only helps
11 http://www.mpi-inf.mpg.de/yago-naga/yago/
12 http://www.umbel.org/
with rapidly filling the model, but it automatically interlinks it with the LOD cloud in
one go.</p>
      <p>The user will also be able add his own classes and instances or use PoolParty's
natural language processing service to analyse web pages or documents to glean new
concepts for use in the model.</p>
      <p>Of course this PIM will not only consume LOD, but it can also publish the user
created knowledge models as part of the LOD cloud. In this way LOD can be
harnessed to enable the average internet user to create more Open Knowledge. There
will be an online version of this PIM that can be used free of charge.</p>
      <p>In the upcoming project LASSO funded by the Austrian Research Promotion
Agency (FFG)13 we will do research on algorithms that enable smart interlinking of
local data and LOD sources, which will be used for the PIM. Amongst the algorithmic
solutions we will pursue are graph based look-up services (e.g. querying LD sources
by taking context into account instead of just searching for keywords), network search
methods like Spreading Activation and statistical methods such as the Principal
Component Analysis.</p>
    </sec>
    <sec id="sec-6">
      <title>8 Final Remarks</title>
      <p>We have shown how Open Linked Data can help in various ways with easing the
creation of knowledge models like thesauri. At OKCon 2010 we will demonstrated
PoolParty and session visitors will learn how to manage a SKOS thesaurus and how
13 http://ffg.at/content.php
PoolParty supports the user in this process. The document analysis features will be
presented, showing how new concepts can be gleaned from text and integrated into a
thesaurus.</p>
      <p>It will be shown how to interlink local concepts from DBpedia, thereby enhancing
one’s thesaurus with triples from the LOD cloud. Finally the state of the PoolParty
PIM tool will be presented.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aitchison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gilchrist</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bawden</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Thesaurus Construction and Use: A Practical Manual</article-title>
          .
          <source>4th edn. Europa Publications</source>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Pastor-Sanchez</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martínez Mendez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rodríguez-Muñoz</surname>
            ,
            <given-names>J. V.</given-names>
          </string-name>
          :
          <article-title>Advantages of thesaurus representation using the Simple Knowledge Organization System (SKOS) compared with proposed alternatives</article-title>
          .
          <source>informationresarch</source>
          Vol
          <volume>14</volume>
          No.
          <issue>4</issue>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          <year>2009</year>
          . http://informationr.net/ir/14-4/paper422.html
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Viljanen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuominen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hyvönen</surname>
          </string-name>
          , E.:
          <article-title>Publishing and using ontologies as mashup services</article-title>
          .
          <source>In: Proceedings of the 4th Workshop on Scripting for the Semantic Web (SFSW</source>
          <year>2008</year>
          ),
          <source>5th European Semantic Web Conference</source>
          <year>2008</year>
          (
          <article-title>ESWC 2008), Tenerife, Spain (June 1-5</article-title>
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>DBpedia - A crystallization point for the Web of Data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web. Volume</source>
          <volume>7</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>3</given-names>
          </string-name>
          ,
          <string-name>
            <surname>September</surname>
            <given-names>2009</given-names>
          </string-name>
          , Pages
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>