<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Computing Recommendations for Long Term Data Accessibility basing on Open Knowledge and Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sergiu Gordea</string-name>
          <email>sergiu.gordea@ait.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew Lindley</string-name>
          <email>andrew.lindley@ait.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roman Graf</string-name>
          <email>roman.graf@ait.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AIT - Austrian Institute of, Technology GmbH</institution>
          ,
          <addr-line>Donau-City-Strasse 1, Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <fpage>51</fpage>
      <lpage>58</lpage>
      <abstract>
        <p>Digital access to our cultural heritage assets was facilitated through the rapid development of the digitization process and online publishing initiatives as Europeana or the Google books project. As Galleries, Libraries, Archiving institutions and Museums (GLAM) created digital representations of their masterpieces new concerns arise regarding the longterm accessibility of digitized and digitally born content. Repository managers of institutions need to take well documented decisions with regard to which digital object representations to use for archiving or long term access to their valuable collections. The digital preservation recommender system presented within this paper aims at reducing the complexity in the process of decision making by providing support for classification and the preservation risk analysis of digital objects. Technical information which is available as linked data in open knowledge sources facilitates the construction of the DiPRec's recommender knowledge base. This paper presents the DiPRec recommender system, a community approach on how to achieve the generation of well founded and trusted recommendations through open linked data and inferred knowledge in the domain of long-term information preservation for GLAM institutions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.7 [Information Systems Applications]: Digital
Libraries; M.8 [Knowledge Management]: Knowledge Reuse
Digital preservation, Recommender systems
Knowledge based recommender, open recommendations, linked
open data, preservation planning</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>Knowledge based recommender systems (KBRs) as
natural followers of expert systems are nowadays used for
supporting the decision making process in multiple application
areas as: e-commerce, financial services, tourism, etc. One
of the most important challenges of KBRs is the
construction of their underlying knowledge base. This is typically
composed by sets of factual knowledge, i.e. information
describing the application’s domain and business rules. Both
together enable the drawing of conclusions and support the
decisions making process when analyzing the utility of a
specific item in a given context as for example, analyzing the
effectiveness of digitizing and publishing Mircea Eliade’s book
”History of Religious Ideas” within Google books.</p>
      <p>Even though the world wide web has turned out to be
the largest knowledge base, information published lacks an
unified well-formed representation and mainly is intended
for human readers. The Linked Open Data (LOD)1 and
Open Knowledge2 initiatives address these weaknesses by
describing a method on how to provide structured data in a
well-defined and queriable format. By linking together and
inferring properties of di↵ erent independent and publically
available information sources like FreeBase3, DbPedia4 and
Pronom 5 within the specific context of a digital preservation
scenario we shortcut the well known challenge of KBRs, the
knowledge acquisition bottleneck.</p>
      <p>In this paper we present our work carried out in the
context of the Assets6 project with the aim of preparing the
ground for digital preservation within Europeana7. The
Europeana portal serves as a central point for the large public
to easily explore and research European cultural and
scientific heritage online. It aggregates and collects data on
digital resources from galleries, libraries, archives and
museums accross Europe and by now manages about 19 million
object descriptions collected from more than 15 hundred
institutions. Within this very heterogeneous context it is
easily understandable that digital objects are encoded in very
heterogeneous file formats and versions throughout various
di↵ erent hardware and software content repository systems.
Depending on the underlying use case it is likely that
mul1http://linkeddata.org/
2http://www.okfn.org/
3http://www.freebase.com
4http://dbpedia.org/
5http://www.nationalarchives.gov.uk/PRONOM/
6http://www.assets4europeana.eu/
7http://www.europeana.eu/portal/
tiple representations of the same ’physical’ object exist at
a time. For example in most cases it is useful to provide
access copies on demand which are easily distribuatable via
the web while the master record needs to adhere to di↵
erent requirements as for example the institution’s long-term
scenario and preservation policy.</p>
      <p>A key topic in preservation planning is the file formats
used for encoding the digital information. The Pronom
Unique Identifiers (PUIDs) registry provides persistent, unique
and unambiguous identifiers for file formats and therefore
takes a fundamental role in the process of managing
electronic records. Currently it lists information on about 820
di↵ erent PUIDs. While some of the formats are properly
documented, open-source and well supported, others may
be outdated, redeemed by software vendors and no longer
functional in modern operating systems. As always the the
binary file’s dependencies on the underlying platform, its
configuration (codecs, plugins, etc.) as well as the
rendering software are responsible on generating a concrete user
performance, it is vital to have a solid understanding on all
of them. This process is costly and requires a high degree
of engineering expertise. Many of the GLAM institutions
already outsource IT related activities and don’t have the
resources to keep track of the required level of complexity in
house.</p>
      <p>The Digital Preservation Recommender (DiPRec) system
addresses the topics of ’preservation watch’ and
’preservation policy recommendation’. It proposes a solution in the
domain of digital long-term preservation for making
documented recommendations based on risk scores, while the
underlying knowledge base is built through a linked data
approach. Information from FreeBase, DbPedia and Pronom
in the areas of file formats, file conversions tools, hardware
and software vendors is taken into account. The main
contribution of this paper consists in the integration of open
(general or domain specific) data when constructing
knowledge based recommendations. The ”knowledge acquisition
bottleneck” and the high costs of setting up and
maintaining KBRs are still an impediment for extensively adoption
by the industry. Recommendations provided by DiPRec are
meant to support GLAM institutions across Europe in the
process of analyzing their digital assets. The technical
foundation and the explanation of the DiPRec recommendations
are computed on top of shared and collaboratively built data
sources, trust in the area of LOD and digital preservation
is a key issue which has been left out for this paper due to
simplicity.</p>
      <p>The novelty of our work consists in combining expert tools
(as File, Droid or Fido) and automated object
identification processes, with structured information (e.g.
technical information on file formats) from open data
repositories. This information is use for infering new knowledge,
calculate preservation risks and finally for computing
recommendations on preservation actions in the domain of
digital long-term preservation. We present the rationale used
for the construction of the DiPRec recommender by
presenting concrete examples of a given content analysis which
was provided for the Assets project. The rest of the
paper is organized as follows; in Section 2 we present related
work carried out on recommender systems and in the field
of digital preservation. Section 3 highlights the architecture
of DiPRec by comparing it against the construction of
classical KBRs. The functionality provided by our system is
explained in detail through a concrete example on the TIFF
file format. The evaluation of our approach is presented in
Section 4 by analyzing the digital collections of the Assets
project. This is followed in the last Section of the paper (nr.
5) by the summarization of the concluding remarks for our
work.
2.</p>
      <p>RELATED WORK</p>
      <p>
        Knowledge Based Recommender systems gained broad
popularity in e-commerce and e-tourism [
        <xref ref-type="bibr" rid="ref11 ref19 ref24 ref7">7, 11, 24, 19</xref>
        ]
applications supporting customers in their decision making
processes. The two most popular use cases are guidance through
large and complex product o↵ ers (e.g. trip organization,
feature selection of technical equipment) as well as
accompanying the process of high cost decision making (e.g. financial
investments). When designing the DiPRec recommender we
took into consideration the Advisor Suite [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and Planets
Testbed infrastructure [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The main component of the
Advisor Suite is a multipurpose workbench which o↵ ers
support and advanced graphical user interfaces for constructing
knowledge based recommenders. Advisor Suite features
include the import of product catalogues, visual editing of a
recommendation workflow and the generation of a runtime
environment. The Planets8 project focused on constructing
practical services and tools for establishing empirical
evidence in the process of informed decision making in the area
of digital long-term preservation. A major achievement was
the definition of basic nouns and verbs for core
preservation operations. This allows to easily combine and swap
tools within a preservation workflow and lead to a
number of over fifty preservation services. Available services
were deployed and tested within the Planets Testbed [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], a
uniform environment for experimentation under well-defined
and controlled surroundings. It provides automated quality
assurance support for tools like DROID9, JHOVE10 and the
eXtensible Characterisation Languages11[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        A key topic in preservation planning is the process of
evaluating objectives under the limitation of well-known
constraints. A state of the art report on technical
requirements and standards as well as available tools to support
the analysis and planning of preservation actions is given in
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Strodl et al. present the Planets preservation planning
methodology Plato12 by an empirical evaluation of image
scenarios [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and demonstrate specific cases of
recommendations for image content in four major National Libraries in
Europe[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. After eliciting information regarding the
preservation scenario (user requirements) the Plato tool is able to
recommend specific preservation actions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for a given
scenario. The tool was specifically designed to work on
samples of the underlying data set and therefore is able to make
use of XCL or similar tools for automated quality assurance
and semi-automated evaluation of objectives. In contrast to
these scenario evaluations, DiPRec aims at collecting
information on a broader range from open linked data registries
and dynamic knowledge sources. It can evaluate more
general, even ’non-technical’ objectives (e.g. what is the risk
that no software vendor will support old formats like Word
8http://www.planets-project.eu/
9http://droid.sourceforge.net/
10http://hul.harvard.edu/jhove/
11http://planetarium.hki.uni-koeln.de/public/XCL/
12http://www.ifs.tuwien.ac.at/dp/plato/intro.html
      </p>
      <p>Open Domain</p>
      <p>Knowledge
Pronom</p>
      <p>UDFR
Domain Expert
Knowledge</p>
      <p>User requirements
3 documents? ). This is a significant improvement over the
Plato tool where all this information needs to provided by
domain experts.</p>
      <p>
        The Scape13 project is one of the major current initiatives
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] which is partially funded by the European Union’s FP7
on institutional preservation requirements. The project
addresses besides the issues of scalable preservation and
qualityassured preservation workflows also the topic of policy-based
preservation planning and watch.
      </p>
      <p>
        The paradigms of semantic Web and linked open data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
transform the web from a pool of information into a
valuable knowledge source of data according to the definitions
of a knowledge management theory [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The exploitation
of linked data as knowledge source for recommender
system started as research topic in the last few years and was
first applied to improve case-based and collaborative
filtering recommenders [
        <xref ref-type="bibr" rid="ref10 ref20 ref9">10, 9, 20</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] the authors present the
Talis Aspire system which is able to assists educational sta↵
in picking educational web resources. The employment of
linked data in collaborative filtering and case-based
reasoning was explored by Heitmann and Hayes in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
3. SYSTEM OVERVIEW
      </p>
      <p>Typically the creation of classic knowledge based
recommender systems consists of three main tasks. Dealing with
the collection of detailed descriptions of products o↵ ers is
followed by the process of constructing a recommendation
knowledge base (see section 3.2). At runtime user
requirements elicitation takes place and recommendations are
computed based on the underlying recommendation knowledge
base and the items that match the given user requirements.
DiPRec follows the same process but improves the way the
knowledge base is built in order to reduce the e↵ orts spent
on domain knowledge acquisition. This is especially relevant
for being exploited in GLAM preservation scenarios, where
the underlying knowledge base contains broader
information than the domain specific KBRs. Within the DiPRec
recommender the Domain Information Aggregation module
is responsible for collecting file format related information
(e.g. formats, vendors, applications, etc.) from the open
knowledge bases Pronom, DBPedia and Freebase.
Furthermore the Domain Knowledge Aggregation module combines
the outcome of a risk analysis process with the knowledge
manually provided by domain experts. Figure 1 compares
the process used by regular KBRs and the one presented by
DiPRec which enhances the process of building the
underlying knowledge base. In the following sections we present
extended details on how the knowledge base of DiPRrec is
built by using as example the Tagged Image File Format
(TIFF).</p>
      <p>
        The TIFF format is still very popular among the
publishing industry, as it is a very adaptable file format although
it did not have a major update since 1992. It was originally
created by Aldus and since 2009 it is now under control of
Adobe Systems. There are a number of extensions
available (e.g. TIFF/IT, TIFF-FX) which have been based on
the TIFF 6.0 specification, but not all of them are broadly
used. A standard and broadly accepted approach in the
archiving world is the migration of TIFF encoded content
to the JPEG2000 format. In [
        <xref ref-type="bibr" rid="ref2 ref4">4, 2</xref>
        ] one can find the context
in which several content providers took the decision to
perform this kind of content migration. However within these
scenarios, the context evaluation and the recommendation
were computed by domain experts and by expert systems.
      </p>
      <p>The DiPRec system, on the one hand applies to the
approach of well-documented and trackable decision making,
and at the same time it uses a semi-automatic approach
on domain knowledge acquisition. This reduces the human
e↵ ort invested by domain experts when providing
reservation recommendations, reduces the financial e↵ orts invested
in the context evaluations, and in the same time is able to
o↵ er good quality recommendations.</p>
      <sec id="sec-2-1">
        <title>FILE FORMAT DESCRIPTION</title>
        <p>Format Name Tagged Image File Format (P), Tagged Image File Format (D), Tagged Image File Format(F)</p>
      </sec>
      <sec id="sec-2-2">
        <title>Pronom Id fmt/10 (P)</title>
        <p>Mime Type /media type/image/ti↵ -fx, /media type/image/ti↵ (F), image/ti↵ (P)</p>
      </sec>
      <sec id="sec-2-3">
        <title>File Extensions .ti↵ , .tif (D)</title>
      </sec>
      <sec id="sec-2-4">
        <title>Current Version 6 (P)</title>
      </sec>
      <sec id="sec-2-5">
        <title>Current Version Release Date 03 Jun 1992 (P)</title>
      </sec>
      <sec id="sec-2-6">
        <title>Software License Proprietary software (D)</title>
        <p>Software QuickView Plus, Acrobat, AutoCAD, CorelDraw, Freemaker, GoLive, Illustrator, Photoshop,</p>
      </sec>
      <sec id="sec-2-7">
        <title>Powerpoint (P), SimpleText, Seashore, Imagine (D)</title>
        <p>Software Homepage http://adobe.com/photoshop(D)</p>
      </sec>
      <sec id="sec-2-8">
        <title>Operating System PC, Mac OS X, Microsoft Windows (D) Genre Image (Raster) (P), Image file format (I), SimpleText - Text editor, Adobe Photoshop - Raster graphics editor (D)</title>
      </sec>
      <sec id="sec-2-9">
        <title>Open Format none (P)</title>
        <p>Standards ISO 12639:2004 (W)
Vendors Aldus, Adobe Systems, Apple Computer, now Apple Inc., Microsoft (D), Adobe Systems</p>
      </sec>
      <sec id="sec-2-10">
        <title>Incorporated (P), Aldus Corporation (P)</title>
      </sec>
      <sec id="sec-2-11">
        <title>VENDOR DESCRIPTION</title>
      </sec>
      <sec id="sec-2-12">
        <title>Organization Name</title>
      </sec>
      <sec id="sec-2-13">
        <title>Country</title>
      </sec>
      <sec id="sec-2-14">
        <title>Foundation date</title>
      </sec>
      <sec id="sec-2-15">
        <title>Number of Employees</title>
      </sec>
      <sec id="sec-2-16">
        <title>Revenue</title>
      </sec>
      <sec id="sec-2-17">
        <title>Homepage</title>
      </sec>
      <sec id="sec-2-18">
        <title>Adobe Systems</title>
      </sec>
      <sec id="sec-2-19">
        <title>United States (P)</title>
        <p>Dec 1982 (F)
6068 (Jan 2007), 8660 (2009)(F), 9,117 (2010)(W)
3,579,890,000 US$ (Nov 28, 2008) (F)
http://adobe.com/photoshop(F)</p>
        <p>Di↵ erently to the e-commerce domain where KBRs import
detailed item descriptions from product catalogs there is no
such catalog for computer file formats. The Unified
Digital Format Registry (UDFR)14 project was started in 2009
by a group of Universities and GLAM institutions with the
aim of building a single, shared technical registry for file
formats based on a semantic web and linked data approach.
The project is based on the Pronom database which
provides basic information about a large number of file formats
and will be extended by data on migration pathways and
available software/tools. The registry should be available
from the beginning of 2012. As Pronom data is not rich
enough to build a recommendation and reasoning
mechanism for preservation scenarios of file formats on top, we
collect additional information sources and aggregate them
into a single homogeneous property representation in the
recommender’s knowledge base. DiPRec uses two types of
operations for aggregating domain information:
• data unification: the data representation retrieved from
di↵ erent knowledge bases is unified and combined
under the DiPRecs property model definition. For
example, the number of software tools supporting a given file
format is calculated over di↵ erent data sources. The
individual object’s namespace, the transformation
process of values, the query on how to extract a given
record, etc. are preserved and are part of the
property’s model representation.
• property composition: more abstract properties which
require a hierarchical composition are computed by
aggregating basic properties by weighted numbers. The
model on property definition is meant to be kept very
simple. For example ”supported by major vendors”
will check if at least one of the software companies is
considered to fulfill this requirement by combining the
properties like ”NUMBER OF EMPLOYEES”,
”VENDOR REVENUE”). See Table 2.</p>
        <p>
          When aggregating domain information we are
interrogating external knowledge sources like DbPedia and Freebase
which manage huge amounts of linked open data triples.
This allows us to extract fragmental descriptions on file
formats, software applications and vendors supporting given
file formats (see Table 1). DbPedia allows to post
sophisticated queries using SPARQL query and OWL ontology
languages [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] for retrieving data available in Wikipedia.
Freebase [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] is a practical, scalable semantic database for
structured knowledge and is mainly composed and
maintained by community members. Public read/write access to
Freebase is allowed through an graph-based query API using
the Metaweb Query Language (MQL) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. PRONOM data
is released as linked open data and is accessible through a
public SPARQL endpoint.
        </p>
      </sec>
      <sec id="sec-2-20">
        <title>AGGREGATED PROPERTIES</title>
      </sec>
      <sec id="sec-2-21">
        <title>File format related</title>
      </sec>
      <sec id="sec-2-22">
        <title>Is supported by major software vendors?</title>
      </sec>
      <sec id="sec-2-23">
        <title>Is an open file format?</title>
      </sec>
      <sec id="sec-2-24">
        <title>Is widely supported by current web browsers?</title>
      </sec>
      <sec id="sec-2-25">
        <title>Which versions o cially supported by vendor?</title>
      </sec>
      <sec id="sec-2-26">
        <title>Which versions are frequently used?</title>
      </sec>
      <sec id="sec-2-27">
        <title>Image file compression supported?</title>
      </sec>
      <sec id="sec-2-28">
        <title>Preservation related metadata</title>
      </sec>
      <sec id="sec-2-29">
        <title>Is creator information available?</title>
      </sec>
      <sec id="sec-2-30">
        <title>Is publisher information available?</title>
      </sec>
      <sec id="sec-2-31">
        <title>Is digital rights information available?</title>
      </sec>
      <sec id="sec-2-32">
        <title>Is file migration allowed?</title>
      </sec>
      <sec id="sec-2-33">
        <title>Object creation date?</title>
        <p>Is an object preview available?
yes
no
yes
6.0
6.0
yes
yes/no
yes/no
yes/no
yes/no
datetime
URL</p>
        <p>DIGITAL
PRESERVATION</p>
        <p>TOOLS
(Droid/Pronom)</p>
        <p>Property Sets
Domain Information</p>
        <p>Aggregation
FILE METADATA</p>
        <p>PRESERVATION</p>
        <p>RISK
Score / Report
Open Knowledge</p>
        <p>Sources
(DBPedia, Freebase,</p>
        <p>PRONOM )</p>
        <p>Pronom as presented before is a viable resource for
anyone requiring impartial and definitive information about the
file formats, software products and other related data.
Extremely valuable to the DiPRec recommender is the
information related to the file conversion tools based on a given
PUID. Therefore we employ the Droid 15 characterization
service for automatically extracting technical metadata and
identifying file formats from physical media files. This
metadata is then used in conjunction within the domain
knowledge aggregation process presented in the Fig. 2</p>
        <p>The risk analysis module is in charge of evaluating
information previously aggregated in the DiPRec knowledge
base for a given record at hand over following (exemplary)
dimensions of digital preservation:
• Web accessibility: Dissemination copies are published
and accessible on e.g. the content provider’s web
portal. There should be previews of objects (e.g.
thumbnails for images, video summaries, short intro for
audio files) and ’rich’ object descriptions to increase their
visibility and retrieval. The chosen file representations
should render in the latest browsers without plugin
support and cope with modern features (e.g. pseudo
streaming, progressive image display, HTML5, X3D,
etc). Content is made available through di↵ erent
exploitation channels.
• Archiving and costs: The decision of following a
specific institutional preservation policy for a given
technology is heavily influenced by given hardware and
budget constraints. Future exploitations on the costs
for content exploitation need to be predicted and taken
into account.</p>
        <p>Other scenarios may include:
• Provenance metadata
• Data exchange and collaborative data enrichment
• Publishing and digital rights management</p>
        <p>The definition of preservation dimensions is not
orthogonal and therefore certain properties might be involved more
than once when computing di↵ erent risk score. Due to
management and maintenance reasons properties are also
grouped by sets and a property may belong to one or more
property sets. The extent to which a property belongs to a
15http://sourceforge.net/projects/droid/
property set and consequently contributes to the risk
computation over a given dimension is modeled through the
introduction of specific weighting factors (see Equation 1).</p>
        <p>The value of the overall risk score for a given collection
of objects is computed as a weighted sum over all digital
preservation dimensions:</p>
        <p>Ri =</p>
        <p>X
ps2 P Si
wps,i ⇤</p>
        <p>X
p2 P ROPps
wp,ps ⇤ d(p, P F V (p))
(1)</p>
        <p>Where Ri represents the preservation risk computed over
the dimension i. ps represents the index of the current
property set within all sets associated to the dimension i. The
w(ps,i) is the weight of the contribution of the property set ps
to the dimension i. Similarly, p stands for the index of
current properties within the list of properties available in the
given property set P ROPps. wp,ps denotes the importance
of a property p for the property set ps. The distance
between the current property and the defined - ’preservation
conform’ - value for this property is represented through
d(p, P F V (p)).
3.3</p>
        <sec id="sec-2-33-1">
          <title>User requirements elicitation</title>
          <p>DiPRec is designed to work as a multi-purpose digital
preservation support tool which can be used in various
scenarios by di↵ erent types of customers. For examples the tool
may support content providers in analyzing the ’preservation
friendliness’ of their infrastructure, their archiving solutions
or the visibility of their artifacts published in the Europeana
portal. Recommendations are always to be seen in the
context in which the digital objects are used. Within the scope
of the Assets project there is the common interest to o↵ er
public access to digital assets through the Europeana portal
(i.e. web discovery), to provide advance search
functionality (i.e. description richness and preservation of provenance
information) as well as the topic of the data archiving
dimension.</p>
          <p>As a result of the requirements elicitation process user
profiles are created. A set of multiple choice questions is used to
distinguish the relevant dimensions of available preservation
objectives. According to di↵ erent levels of complexity, role
and required domain knowledge the system o↵ ers a subset
of questions which are well understood and the best
available choice for a user to express his needs. Fig. 3 presents
sample workflow which could be used to determine a given
user profile. For example a private user (ut = private
person) with a solid level of IT knowledge (itk = expert ) will
be asked about preferred encodings and compression types
of the digital content, while others would define attributes
about storage limitations and upload samples of a given
collection.
3.4</p>
        </sec>
        <sec id="sec-2-33-2">
          <title>Recommendation computation</title>
          <p>Di↵ erently to classic KBRs where the application’s scope
is very well delimited in terms of selecting the best
matching items in a list of known possibilities, the DiPRec
system relies on expressing an institutional preservation
context in form of user requirements that are combined with the
knowledge acquired about the long term accessibility
threatening. We employ tools to evaluate the content of a given
collection from a technical point of view and to generate fine
grained preservation risk scores. When records are identified
to have vulnerabilities on certain preservation dimensions a
rule based engine as JBoss Drools16 is used to propose
appropriate preservation actions. The set of available business
rules are defined by domain experts in form of simple
IFTHEN-ELSE rules. These rules are neither complete nor
meant to be non-overlapping. Unified tool access for
processing executable preservation plans is provided through
the Assets preservation normalisation framework which is
able to invoke the tools with exactly defined settings and
parameter configurations.</p>
          <p>IF ( rac &gt; 0.5 AND ia == true AND iwa == true AND
open f ormat == FALSE)
THEN migrate(preservation f ormat)
IF (content type == IMAGE)
THEN preservation format = (JPEG/2000:1, TIFF/6:0.8)
IF (f ile f ormat == TIFF/5 AND
preservation f ormat == JPEG/2000)</p>
          <p>T HEN migration tool = IM AGE MAGICK
(2)</p>
          <p>
            The preservation recommendations are computed using
the constraint solving problems (CSP) theory [
            <xref ref-type="bibr" rid="ref11 ref8">8, 11</xref>
            ].
Constraints are defined within the preservation actions
knowledge base, the CSP context is defined by user profiles and
the preservation risks are identified for the given data
collection. The recommendations are represented in form of
preservation actions. For example, the set of business rules
defined above combined with a user profile indicating
interest in the dimension of archiving and web accessibility will
lead to the following recommendation when analyzing a
collection of images in TIFF format:
migrate(T IF F/5, J P EG/2000, IM AGE M AGICK)
In free text translation, the recommendation will suggest
the migration of the files available in T IF F/5 format to
J P EG/2000 by using the IM AGE M AGICK software with
standard settings.
4.
          </p>
          <p>EVALUATION</p>
          <p>The evaluation of the first prototype of DiPRec was
conducted within the scope of the Assets project. Ten partners
of the project consortium provided metadata and binary
content (10 collections with a total size of 516GB contained
in 368067 media files) for supporting the development and
testing of services developed within the scope of the project.
The first step in the evaluation process was the identification
of file formats, definition of property sets and the
aggregation of the domain knowledge available in open knowledge
bases on these file formats.</p>
          <p>The Table 3 lists the distribution of file formats by content
type. Even the experimental data was taken from a small
number of content providers, we discovered a variety of 18
formats in 38 di↵ erent versions used for encoding the digital
content.</p>
          <p>The Digital Record Object Identification tool (DROID)
version 5, signature file 45 was executed through the
Assets preservation normalisation tool suite and was able to
successfully identify file formats in 95 percent of the cases
through its binary signature method except of the 3D model
objects which have not yet been collected by Pronom.
Appropriate information on all of the file formats was contained
in DbPedia and Freebase and the domain knowledge
acquisition process was completed by successfully computing the
preservation risk analysis scores.</p>
          <p>The second part of the evaluation consisted in
computing recommendations for the given content. Therefore, we
created a user profile for content providers that are
interested in making their content accesible through Europeana.
Within this context, the content providers manifest interest
for the web accessibility digital preservation dimension.</p>
          <p>The highest diversity of file formats was found in the
image collections. The recommendation to migrate these
files to the JPEG 2000 format didn not get a high priority
and will be performed within the next period of scheduled
storage migration. The Image Magick tool was the
recommended choise for performing this transformation action.
The whole audio content available in Assets was provided in
the mp3 format and no recommendation was made for
transforming audio collections. The most restrictive constraints
for web accessibility are defined for the video content. The
pseudostreaming protocol is an advanced technological
solution used for distributing information over the web. It allows
the user to interact with the media-player and to quickly
navigate within the content without the needed to
download the entire media file. This protocol is supported by
two file formats: flash video (FLV) and MPEG4 with H2.64
video encoding. It has native support in HTML5 and is used
in HTML4 with an adequate browser plugin. A part of the
Assets content is already available in FLV format and
another part is available in MPEG1 or MPEG2. The DiPRec
resulting recommendation is to migrate the content to FLV
by using the ↵ mpeg 17 tool.</p>
          <p>CONCLUSION</p>
          <p>Within this paper we introduced the DiPRec recommender
system, an expert support tool in the domain of digital
longterm preservation for GLAMs. An important contribution
of this papers is the exploitation of an open linked data
approach for constructing the recommender’s knowledge base
built upon open registries as DbPedia and Pronom. Since
the knowledge acquisition, aggregation and unification
process is fully automated it is easy to upgrade the
recommender’s knowledge base.</p>
          <p>We looked at preservation planning which is the process of
specifying clearly defined and relevant trees of objectives in a
defined preservation dimension and evaluating them within
a given (institutional) context to generate well-documented
decisions. DiPRec is able to advance the process with
inferred community knowledge and reduces the degree of
manual evaluation processes or require technical expertise in this
process.</p>
          <p>Am important concern related to the KBRs is the trust
in the provided recommendations. This is especially
relevant for the digital preservation domain where we deal with
a large amount of multimedia material and the execution
of the preservation actions is associated with considerable
costs. Within this paper we did not examine the
completeness, correctness and quality degree of the underlying data.
We however argue that data from open knowledge bases like
DbPedia or Freebase could protect from biases introduced
by the economical interests of professional companies by its
underlying community approach.</p>
          <p>
            The tool has been designed by reusing our past experience
in building knowledge based and case based recommender
systems [
            <xref ref-type="bibr" rid="ref23 ref8">23, 8</xref>
            ] and combining it with the expertise of
creation long-term preservation infrastructure and applications
[
            <xref ref-type="bibr" rid="ref1 ref14">14, 1</xref>
            ]. Based on this work the Assets normalisation tool
suite is able to automate the process of object identification
and characterisation and therefore directly integrates within
the property evaluation, risk analysis and recommendation
process for a given record. We presented a first evaluation
of digital content provided by national libraries and archives
through the Assets project where the underlying concepts of
the DiPRec approach were proven to work adequately.
          </p>
          <p>ACKNOWLEDGMENTS</p>
          <p>This work was partially supported by the EU project
”ASSETS - Advanced Search Services and Enhanced
Technological Solutions for the European Digital Library” (CIP-ICT
PSP-2009-3, Grant Agreement n. 250527).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Aitken</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Helwig</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicchiarelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ross</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The planets testbed: Science for digital preservation</article-title>
          .
          <source>Code4Lib</source>
          <volume>1</volume>
          (
          <issue>3</issue>
          ) (
          <year>2008</year>
          ), http://journal.code4lib.org/articles/83
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulovits</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guttenbrunner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strodl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofman</surname>
          </string-name>
          , H.:
          <article-title>Systematic planning for digital preservation: evaluating potential strategies and building preservation plans</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <fpage>133</fpage>
          -
          <lpage>157</lpage>
          (
          <year>2009</year>
          ), http://dblp.uni-trier.de/db/journals/jodl/ jodl10.html#BeckerKGSRH09
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulovits</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofman</surname>
          </string-name>
          , H.:
          <article-title>Plato: a service-oriented decision support system for preservation planning</article-title>
          .
          <source>In: Proceedings of the 8th ACM/IEEE-CS joint conference on Digital libraries</source>
          . pp.
          <fpage>367</fpage>
          -
          <lpage>370</lpage>
          . ACM, New York (
          <year>2008</year>
          ), http: //publik.tuwien.ac.at/files/PubDat_170832.pdf,
          <source>vortrag: 8th ACM/IEEE-CS joint conference on Digital libraries (JCDL</source>
          <year>2008</year>
          ), Pittsburgh, Pennsylvania;
          <fpage>2008</fpage>
          -06-
          <fpage>16</fpage>
          - 2008-06-20
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Four cases, three solutions: Preservation plans for images</article-title>
          .
          <source>Tech. rep.</source>
          , Vienna University of Technology, Vienna, Austria (April
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heydegger</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schnasse</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thaller</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A generic xml language for characterising objects to support digital preservation</article-title>
          .
          <source>In: SAC '08: Proceedings of the 2008 ACM symposium on Applied computing</source>
          . pp.
          <fpage>402</fpage>
          -
          <lpage>406</lpage>
          . ACM, New York, NY, USA (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linked data - the story so far</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst</source>
          .
          <volume>5</volume>
          (
          <issue>3</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Burke</surname>
          </string-name>
          , R.D.:
          <article-title>Hybrid web recommender systems</article-title>
          .
          <source>In: The Adaptive Web</source>
          . pp.
          <fpage>377</fpage>
          -
          <lpage>408</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Felfernig</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordea</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Ai technologies supporting e↵ ective development processes for knowledge-based recommender applications</article-title>
          .
          <source>In: SEKE</source>
          . pp.
          <fpage>372</fpage>
          -
          <lpage>379</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Heitmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
          </string-name>
          , C.: C.:
          <article-title>Using linked data to build open, collaborative recommender systems</article-title>
          .
          <source>In: In: AAAI Spring Symposium: Linked Data Meets Artificial IntelligenceSˇ</source>
          .
          <article-title>(2010</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Heitmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Enabling case-based reasoning on the web of data</article-title>
          .
          <source>In: The WebCBR Workshop on Reasoning from Experiences on the Web</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Jannach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuchs</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Constraint-based recommendation in tourism: A multiperspective case study</article-title>
          .
          <source>J. of IT &amp; Tourism</source>
          <volume>11</volume>
          (
          <issue>2</issue>
          ),
          <fpage>139</fpage>
          -
          <lpage>155</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Jannach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jessenitschnig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seidler</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Developing a conversational travel advisor with advisor suite</article-title>
          .
          <source>In: ENTER'07</source>
          . pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Jens</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Jo¨rg,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , So¨ren, A.:
          <article-title>Discovering unknown connections -the dbpedia relationship finder</article-title>
          .
          <source>In: Proceedings of the 1st Conference on Social Semantic Web (CSSW)</source>
          . vol. P-
          <volume>113</volume>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>109</lpage>
          . Gesellschaft fu¨r Informatik, Leipzig, Germany (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>King</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Steeg</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>The planets interoperability framework: An infrastructure for digital preservation actions</article-title>
          .
          <source>In: ECDL09 Proceedings of the 13th European conference on Research and advanced technology for digital libraries</source>
          . vol.
          <volume>5714</volume>
          /
          <year>2009</year>
          , pp.
          <fpage>425</fpage>
          -
          <lpage>428</lpage>
          . Springer-Verlag (
          <year>2009</year>
          ), http://dx.doi.org/10.1007/978-3-
          <fpage>642</fpage>
          -04346-8_
          <fpage>50</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Kurt</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Praveen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jamie</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In: SIGMOD '08 Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          . pp.
          <fpage>1247</fpage>
          -
          <lpage>1249</lpage>
          . ACM, New York, NY, USA (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Lindley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aitken</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A collaborative research environment for digital preservation - the planets testbed</article-title>
          .
          <source>Enabling Technologies, IEEE International Workshops on 0</source>
          ,
          <fpage>197</fpage>
          -
          <lpage>202</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Nonaka</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takeuchi</surname>
          </string-name>
          , H.:
          <article-title>The Knowledge-Creating Company: How Japanese Companies Create the Dynamics of Innovation</article-title>
          . Oxford University Press (May
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Orit</surname>
            <given-names>Edelstein</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Factor</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.K.T.R.E.S.P.T.</surname>
          </string-name>
          :
          <article-title>Evolving domains, problems and solutions for long term digital preservation</article-title>
          .
          <source>iPRES 2011 - 8th International Converence on Preservation of Digital Objects</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Werthner</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Case base querying for travel planning recommendation</article-title>
          .
          <source>Journal of IT &amp; Tourism</source>
          <volume>4</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>215</fpage>
          -
          <lpage>226</lpage>
          (
          <year>2001</year>
          ), http://dblp.uni-trier.de/ db/journals/jitt/jitt4.html#RicciW01
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Shabir</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarke</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Using linked data as a basis for a learning resource recommendation system</article-title>
          .
          <source>In: 1st International Workshop on Semantic Web Applications for Learning and Teaching Support in Higher Education (SemHE'09)</source>
          (
          <year>September 2009</year>
          ), http://eprints.ecs.soton.ac.uk/18053/
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Strodl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumayer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>How to choose a digital preservation strategy: evaluating a preservation planning procedure</article-title>
          .
          <source>In: JCDL '07: Proceedings of the 2007 conference on digital libraries</source>
          . pp.
          <fpage>29</fpage>
          -
          <lpage>38</lpage>
          . ACM, New York, NY, USA (
          <year>2007</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1255175.1255181
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Sven</surname>
            <given-names>Schlarb</given-names>
          </string-name>
          , Edith Michaelar,
          <string-name>
            <surname>M.K.A.L.B.A.S.R.A.J.:</surname>
          </string-name>
          <article-title>A case study on performing a complex file-format migration experiment using the planets testbed</article-title>
          .
          <source>IS&amp;T Archiving Conference</source>
          <volume>7</volume>
          ,
          <fpage>58</fpage>
          -
          <lpage>63</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Zanker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordea</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jessenitschnig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schnabl</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A hybrid similarity concept for browsing semi-structured product items</article-title>
          .
          <source>In: EC-Web</source>
          . pp.
          <fpage>21</fpage>
          -
          <lpage>30</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Zanker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jessenitschnig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jannach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordea</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Comparing recommendation strategies in a commercial context</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>22</volume>
          (
          <issue>3</issue>
          ),
          <fpage>69</fpage>
          -
          <lpage>73</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>