<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Linked Data Query Wizard: A Novel Interface for Accessing SPARQL Endpoints</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Patrick Hoefler</string-name>
          <email>er@know-center.at</email>
          <email>phoefler@know-center.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eduardo Veas</string-name>
          <email>eveas@know-center.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Granitzer</string-name>
          <email>michael.granitzer@uni-passau.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christin Seifert</string-name>
          <email>christin.seifert@uni-passau.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Know-Center GmbH</institution>
          ,
          <addr-line>Inffeldgasse 13, Graz</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Passau</institution>
          ,
          <addr-line>Innstraße 33a, Passau</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>8</volume>
      <issue>2014</issue>
      <fpage>2</fpage>
      <lpage>11</lpage>
      <abstract>
        <p>In an interconnected world, Linked Data is more important than ever before. However, it is still quite di cult to access this new wealth of semantic data directly without having in-depth knowledge about SPARQL and related semantic technologies. Also, most people are currently used to consuming data as 2-dimensional tables. Linked Data is by de nition always a graph, and not that many people are used to handle data in graph structures. Therefore we present the Linked Data Query Wizard, a web-based tool for displaying, accessing, ltering, exploring, and navigating Linked Data stored in SPARQL endpoints. The main innovation of the interface is that it turns the graph structure of Linked Data into a tabular interface and provides easy-to-use interaction possibilities by using metaphors and techniques from current search engines and spreadsheet applications that regular web users are already familiar with.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The amount of Linked Data available on the web keeps
growing, mainly due to an in ux of new data from research
and open government activities. At the time of writing, 886
Linked Open Datasets had been registered with datahub.io,
389 of those claiming to provide a SPARQL endpoint1.
However, it is still quite di cult to access this wealth of
semantically enriched data directly without having in-depth
knowledge of SPARQL and related semantic technologies.
1 http://datahub.io/dataset?tags=lod</p>
      <p>In this paper, we present the Linked Data Query Wizard2,
a novel way to explore the data contained in a SPARQL
endpoint using a tabular interface.</p>
      <p>The Linked Data Query Wizard is a web-based data
analysis tool that empowers regular web users to explore, lter,
and analyze Linked Data and should dramatically simplify
the process of accessing any kind of Linked Data contained
in SPARQL endpoints. The prototype currently o ers two
entry points: Users can either initiate a keyword search over
a given SPARQL endpoint, or they can select any of the
already available Linked Datasets represented as RDF Data
Cubes (which will be explained in more detail in Section 4.2).</p>
      <p>In both cases, the Linked Data Query Wizard presents
a table containing the results. The users can then choose
which columns they are interested in, and they can set
lters to narrow down the displayed data. Additionally, they
can explore the data by focusing on an entity, or they can
aggregate a dataset to get a quick overview of the data.</p>
      <p>This paper is structured as follows:</p>
      <p>In Section 2 we discuss the research context and the
requirements that de ned the parameters for the development
of the Linked Data Query Wizard.
2 http://code.know-center.tugraz.at/search
In Section 3 we take a quick look at which related
approaches already exist.</p>
      <p>In Section 4 we describe the Linked Data Query Wizard
and its functionality in more detail.</p>
      <p>In Section 5 we present and discuss the results of a user
study we conducted to nd out if the Linked Data Query
Wizard was actually usable by people who had no knowledge
about Semantic Web concepts and technologies.</p>
      <p>Finally in Section 6 we present our conclusion and
describe future work that could further enhance the Linked
Data Query Wizard.</p>
    </sec>
    <sec id="sec-2">
      <title>RESEARCH CONTEXT &amp; SYSTEM RE</title>
    </sec>
    <sec id="sec-3">
      <title>QUIREMENTS</title>
      <p>
        The Linked Data Query Wizard has been developed in
the context of the EU-funded CODE project3. As outlined
in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the vision of the CODE project has been
to establish a tool chain for the extraction of knowledge
encapsulated in scienti c research papers along with its release
as Linked Data, thereby facilitating the creation of new
insights. One of the project goals was the development of
a web-based visual analytics platform that enables regular
web users to easily perform exploration and analysis tasks
on Linked Data.
      </p>
      <p>With these prerequisites in mind, the following system
requirements had been de ned:</p>
      <p>R1: The system needed to be completely
webbased. It had to be usable without the need for any
client software other than an up-to-date web browser,
and it must not rely on any browser plug-ins or
extensions. Due to the potentially complex data analysis
tasks, support for mobile clients with limited screen
sizes was not a requirement.</p>
      <sec id="sec-3-1">
        <title>R2: The system needed to support data from</title>
        <p>any domain. It was clear from the start that the
system should automatically adapt to any kind of Linked
Data and not be tailored to a speci c domain or use
case.</p>
      </sec>
      <sec id="sec-3-2">
        <title>R3: The system should be based on Semantic</title>
      </sec>
      <sec id="sec-3-3">
        <title>Web standards as much as possible. Just as it</title>
        <p>should work with any kind of Linked Data, it should
also work with any SPARQL endpoint that complies
with the respective current W3C standards.</p>
      </sec>
      <sec id="sec-3-4">
        <title>R4: The system needed to be easy to use. Since</title>
        <p>the Linked Data Query Wizard is intended to be used
by regular web users, the interface had to be kept
simple. The end users should not know that they were
actually accessing the Semantic Web through SPARQL
queries.</p>
      </sec>
      <sec id="sec-3-5">
        <title>R5: The system should make use of what reg</title>
        <p>ular web users already know. In the context of
this prototype, this mainly meant how to use current
search engines and spreadsheet applications.</p>
      </sec>
      <sec id="sec-3-6">
        <title>R6: The system should use the semantic aspects of the data to the advantage of the users.</title>
        <p>This means that certain things should be easier or work
smarter compared to working with non-semantic data.</p>
        <sec id="sec-3-6-1">
          <title>3 http://code-research.eu</title>
        </sec>
      </sec>
      <sec id="sec-3-7">
        <title>R7: The system also needed to be useful for</title>
        <p>Semantic Web experts. While mainly supporting
regular web users, it should be possible to \peek
behind the curtain" and provide helpful functionality for
Semantic Web researchers and developers.</p>
        <p>The main idea for the Linked Data Query Wizard was
to make use of the prior knowledge the users already
possessed when it came to handling data, to make them feel as
comfortable as possible, and not to reinvent the wheel. The
main assumption was that the relevant target group |
people who are interested in looking up and handling data on
the web | already know how to use current search engines
and spreadsheet applications. Therefore the Linked Data
Query Wizard should make use of these concepts:
For getting started, using a simple search box known
from current search engines
For re ning search results, using a table and concepts
from current spreadsheet applications
Enhancing the interface with further functionality, made
possible through the semantic aspects of Linked Data
3.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>RELATED WORK</title>
      <p>The problem of easy-to-use interfaces for accessing Linked
Data is still largely unsolved.</p>
      <p>
        The majority of current tools are not aimed at regular
web users. As an example, Sindice [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a major Semantic
Web search engine, is practically unusable for ordinary web
users due to its complex search interface and results page.
      </p>
      <p>
        Moreover only very few web-based tools used tables for
representing Linked Data. One such example was Freebase
Parallax [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Although its main feature was the ability to
browse sets of related things, it also provided a table view
for these result sets.
      </p>
      <p>
        Another web-based tool that shared similarities with our
prototype was the Falcons Explorer [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Both tools featured
a search box as the main entry point | an idea that is also
central to our prototype. However, in both tools the table
view was not the central focus.
      </p>
      <p>Another tool that shares similarities with our prototype
is OpenRe ne4 (formerly known as Google Re ne and
Freebase Gridworks). It supports RDF, and there are also
extensions such as LODRe ne5 that focus on Linked Data.
OpenRe ne's main focus is cleaning up tabular data, and
it's also not available as a web service, even though its main
interface is browser-based.</p>
      <p>
        Another concept related to our approach is faceted search
and navigation as described e.g. in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and used in
OpenRe ne, SIMILE Exhibit [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or DBpedia's instance of
Virtuoso's Faceted Search &amp; Find feature6.
      </p>
      <p>Although the Linked Data Query Wizard incorporates
certain similarities, most interface elements and concepts
are actually much more similar to those found in current
spreadsheet applications than those used in faceted search
and navigation.</p>
      <sec id="sec-4-1">
        <title>4 http://openrefine.org 5 http://code.zemanta.com/sparkica 6 http://dbpedia.org/fct</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>THE LINKED DATA QUERY WIZARD</title>
      <p>The Linked Data Query Wizard is a completely web-based
tool for accessing Linked Data in SPARQL endpoints in an
innovative way.</p>
      <p>The front page of the Linked Data Query Wizard (see
Figure 1) currently consists of two areas.</p>
      <p>The top area is called \Search Linked Data". It looks
and works basically like current search engines: The central
user interface element is the search box where the users can
enter one or more search terms. These search terms are
then turned into a SPARQL query and handed over to the
chosen SPARQL endpoint. There, with the help of a
fulltext index, a search in all the rdfs:labels is performed, and
the rst 10 results are returned. Additionally, making use of
the COUNT feature of SPARQL 1.1, the SPARQL endpoint
is asked to return the total number of matching results to
be displayed in the user interface.</p>
      <p>The bottom area of the front page of the Linked Data
Query Wizard is called \Show Available Datasets". Here
the users can choose from several lists of preexisting Linked
Datasets that have been prepared and stored in the form of
RDF Data Cubes.</p>
      <p>In the remainder of this chapter we want to highlight
different aspects of the Linked Data Query Wizard. To
begin with, we will focus on the approach and hurdles of the
SPARQL full-text search and explain the concept of RDF
Data Cubes. Then we will showcase the tabular interface
in more detail. Finally we will present the integration with
other tools as well as advanced features for Semantic Web
researchers and developers.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>SPARQL Full-text Search</title>
      <p>As already stated before, one of the main assumptions for
the Linked Data Query Wizard was that the majority of
its target group is accustomed to searching for information
using one of the major search engines (Google, Bing or
Yahoo). Therefore it soon became clear that the main entry
point should be a simple search box that works and feels
similar to what users currently expect when they search for
information on the (non-semantic) web.</p>
      <p>
        The technical implementation of the search feature turned
out to be much more of a challenge: In the current version
of the SPARQL 1.1 Query Language speci cation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the
problem of performant full-text search is not addressed at
all. The only o cially speci ed way to search for something
in a SPARQL endpoint is to lter the results using a regular
expression. This approach, however, creates a potentially
massive performance issue:
      </p>
      <p>If the SPARQL query processor takes the speci cation
literally, the only o cial way to \search" in a SPARQL
endpoint is to rst look up all matching results and then lter
these results according to the regular expression. In the
worst case this means going through all triples before
starting to throw away the ones that do not match the lter
criteria. Apart from memory considerations, a runtime
performance of O(n) is simply not feasible in cases where the
triple store contains millions or even billions of triples.</p>
      <p>Due to this lack of speci cation, SPARQL endpoint
vendors have come up with their own querying mechanisms for
full-text search over SPARQL endpoints. Experiments in
the initial phase of the development of the Linked Data
Query Wizard showed that making use of these proprietary
full-text search mechanism was the only way to achieve a
system performance that regular web users had come to
expect from current search engines. This is also the reason
why the Linked Data Query Wizard currently only supports
Virtuoso7, OWLIM8 and Bigdata9 SPARQL endpoints in
the \Search Linked Data" mode, since all of these provide
integrated full-text search. The Linked Data Query
Wizard has been designed to use Semantic Web standards as
much as possible. Unfortunately, in the case of the full-text
search feature, slightly di erent SPARQL queries are needed
depending on the SPARQL endpoint software | sometimes
even for individual SPARQL endpoints.
4.2</p>
    </sec>
    <sec id="sec-7">
      <title>RDF Data Cubes</title>
      <p>7 http://virtuoso.openlinksw.com
8 http://ontotext.com/owlim
9 http://systap.com/bigdata.htm
therefore was a perfect t for our purposes. Any datasets
that comply with the RDF Data Cube standard and are
publicly available through a SPARQL endpoint can easily
be displayed, ltered, and explored using the Linked Data
Query Wizard (see Figure 2).</p>
      <p>The current version of the front page of the Linked Data
Query Wizard features automatically generated lists of RDF
Data Cubes for several publicly available SPARQL endpoints
(such as EU Open Data 10 or Vienna Linked Open Data11).</p>
      <p>Since the datasets are already pre-processed and mostly
of reasonable size, full-text search is not necessary in this
use case. This also means that the previously mentioned
limitation regarding the SPARQL endpoint vendors does not
apply when accessing RDF Data Cubes through the Linked
Data Query Wizard.</p>
      <p>Thanks to the underlying semantics of the Linked Data
and the aggregation features of SPARQL 1.1, the Linked
Data Query Wizard provides an easy interface to perform
custom aggregations over any given RDF Data Cube, as long
as it is saved in a publicly available SPARQL endpoint that
supports SPARQL 1.1 (see Figure 3).</p>
      <p>On the right side, next to the results table, users can add
more predicates by clicking on the \Add column . . . " button.
Users can then select the column from a list that displays
all available predicates that can provide additional data for
one or more of the currently displayed subjects.</p>
      <p>By default only the rst 10 results are displayed for any
given query. If there is more data available, users can load
all results, 10 more results or 100 more results via buttons
currently located below the results table.
4.4</p>
    </sec>
    <sec id="sec-8">
      <title>Interfaces to External Tools</title>
      <p>The Linked Data Query Wizard currently features 4
interfaces to other tools or services aimed at regular web users:
CODE Visualization Wizard
42-data
MindMeister</p>
      <p>Mendeley
4.4.1</p>
      <p>CODE Visualization Wizard</p>
      <p>
        Once the users are happy with their selected data, they
can visualize it using the CODE Visualization Wizard12 as
described in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The CODE Visualization Wizard is
basically the sibling of the Linked Data Query Wizard. It
enables visual analysis of Linked Data | in the form of
RDF Data Cubes | and supports the user by automating
the visualization process. This means that after analyzing
the structural and semantic characteristics of the provided
Linked Data, the CODE Visualization Wizard automatically
suggests any of the 10 currently available visualizations |
such as line charts, scatter plots, or parallel coordinates |
that are suitable for the provided data. Furthermore the Vis
Wizard automatically maps the data to the available visual
channels of the chosen visualization. If the users wish to
adjust the mapping, they can do so with a few simple clicks.
      </p>
      <p>Usually more than one visualization is suitable for any
given dataset. In that case, all of these visualizations can be
displayed side by side. When certain parts of the data are
selected in on of the visualizations, they are automatically
highlighted in all of the others as well. This can provide
quick insights into complicated data, taking advantage of
the powerful human visual perception system.
4.4.2</p>
      <p>42-data
42-data13 is is the central data marketplace and
integration hub of the CODE project. Users can integrate data
from the Linked Data Query Wizard into answers on
42data and thereby provide context for the data, making it
even more valuable.
4.4.3</p>
      <p>MindMeister</p>
      <p>The MindMeister14 mind mapping service is integrated
into the Linked Data Query Wizard to turn the results
table into a nicely looking mind map (see Figure 5). Each
main branch of the mind map represents a subject from the
results table, and the respective child branches represent the
available predicates and objects. This feature is especially
useful for getting a quick overview over a certain topic with
only a handful of results.
12 http://code.know-center.tugraz.at/vis
13 http://42-data.org
14 http://mindmeister.com</p>
    </sec>
    <sec id="sec-9">
      <title>Features for Semantic Web Researchers and Developers</title>
      <p>The main target group of the Linked Data Query Wizard
are regular web users. However, it can also provide helpful
resources for Semantic Web researchers and developers. The
respective functionalities are currently grouped under the
aptly named menu item \For the Geeks". Currently these
are:</p>
      <p>Cubify the results. For certain functionalities (e.g.
the CODE Visualization Wizard), \generic" RDF needs
to be turned into RDF Data Cubes rst. This is done
by the CODE Data Extractor16, developed at the
University of Passau. This conversion usually happens
automatically in the background, without the need for
any intervention by the users. However, via this menu
button, the \Expert Mode" of the CODE Data
Extractor can be activated, which o ers more exibility
in case of problems with the data.</p>
      <sec id="sec-9-1">
        <title>Display the SPARQL queries. As the name im</title>
        <p>plies, the Linked Data Query Wizard generates SPARQL
queries according to the query and re nements made
by the users. Regular web users are not interested in
SPARQL queries at all | this is one of the main
reasons why the Linked Data Query Wizard exists.
However, for Semantic Web researchers or developers, it
can be helpful to take a look or tweak the SPARQL
queries generated by the Linked Data Query Wizard
15 http://mendeley.com
16 http://zaire.dimis.fim.uni-passau.de:8080/
code-server/demo/dataextraction
(see Figure 6). Additionally, this feature can be used
for performance pro ling, since not only the SPARQL
queries are displayed, but also their respective runtime.</p>
      </sec>
      <sec id="sec-9-2">
        <title>Display the results as JSON-LD. Again, regular</title>
        <p>web users are probably not too interested in turning
their search results into JSON-LD. For Semantic Web
experts or programmers interested in getting started
with JSON-LD, this can nevertheless be a helpful
feature (see Figure 7).</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>EVALUATION</title>
      <p>During the development of the Linked Data Query Wizard
we followed the \release early, release often" principle. This
means that as soon as a feature was complete and ready for
testing, it immediately rolled out to our staging server and,
if no major problems were found, a short time later (usually
within hours, sometimes days) is was publicly available at
our production server. This also means that the Linked
Data Query Wizard has been under permanent scrutiny of
fellow researchers from the CODE project as well as other
interested colleagues for several months now. They regularly
provided valuable feedback on stability and usability issues
as well as helpful feature requests.</p>
      <p>Additionally, the Linked Data Query Wizard was used
in a workshop setting with around 20 students of the
Semantic Technologies course at Graz University of
Technologies. There it proved helpful in evaluating the quality of
the Linked Open Data the students had previously created
during the workshop.</p>
      <p>To conclude the rst development cycle of the Linked Data
Query Wizard, an in-depth evaluation was performed. In the
remainder of this section we will present the study design,
both the quantitative as well as the qualitative study results,
and a discussion of the ndings.
5.1</p>
    </sec>
    <sec id="sec-11">
      <title>Study Design</title>
      <p>
        Our study followed the principles of the Retrospective
Think Aloud protocol ([
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), combined with the NASA
Task Load Index ([
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
      <p>In the following we will describe the details of the study.</p>
      <p>In total, 14 people participated in this study, and 2 people
took part in the related pre-study.</p>
      <p>Each session started with a short explanation of the study
and the signing of the declaration of consent. The session
was then guided by a survey that was lled out by the
participants themselves. The rst page of the survey consisted
of background questions about the participant:</p>
      <p>How's your English? The study was conducted in
English with participants from di erent countries, but
no English native speakers. 10 participants declared
their English skills as \ uent", 3 as \okay" and 1 as
\basic".</p>
      <sec id="sec-11-1">
        <title>Have you used the Linked Data Query Wiz</title>
        <p>ard before? For this study, only participants with
no prior experience with the Linked Data Query
Wizard were selected. Accordingly, all 14 participants
answered with \no".</p>
        <p>How frequently do you use spreadsheet
applications? The Linked Data Query Wizard is mainly
intended for people that have prior experience with
spreadsheet applications. Although this was not checked
during the participant selection phase, all participants
indeed had at least some experience: 2 of them used
spreadsheet applications \every workday", 4 of them
\several times a week", 6 of them \several times a
month" and 2 of them \once a month or less often"</p>
      </sec>
      <sec id="sec-11-2">
        <title>How frequently do you look up information on</title>
        <p>the Internet? This question aimed to probe the level
of the participants' web experience, which turned out
to be quite high: 13 of the participants answered \every
workday", one of them \several times a week".</p>
      </sec>
      <sec id="sec-11-3">
        <title>How frequently do you write SPARQL queries?</title>
        <p>This question was intended to nd out if there were any
Semantic Web experts among the participants. Only
one of the 14 participants answered \once a month or
less often", whereas 6 of them answered \never" and 7
of them \What's SPARQL?".</p>
        <p>What's your age? The nal background question
provided information about the age ranges of the
participants: 4 of the participants were between 18 and 27
years old, 9 of them between 28 and 37, and 1 between
58 and 67.</p>
        <p>After the initial background questions, the participants
had to solve 4 tasks using the Linked Data Query Wizard.
These were as follows:</p>
      </sec>
      <sec id="sec-11-4">
        <title>Task 1: Service Data</title>
        <p>\There is an available dataset called `% of basic public
services for citizens, which are fully available online'
provided by EU Open Data. We are interested only in
the data from the year 2010, please lter it accordingly.
After that, please visualize the results.</p>
        <p>You have 3 minutes to complete this task."</p>
      </sec>
      <sec id="sec-11-5">
        <title>Task 2: Data Overview</title>
        <p>\This task deals with the same data as before, the
dataset called `% of basic public services for citizens,
which are fully available online' provided by EU Open
Data. However, this time we are interested in an overview
of the data. Therefore, please aggregate the dataset
and display the average values, grouped by year.
After that, please visualize the results.</p>
        <p>You have 3 minutes to complete this task."</p>
      </sec>
      <sec id="sec-11-6">
        <title>Task 3: Pulp Data</title>
        <p>\Before you start, please select the data source called
`Wikidata (CODE Edition)' (5th from the top) in the
`Search Linked Data' section.</p>
        <p>There is a lm called `Pulp Fiction'.
1. Was Bruce Willis a cast member of this lm?
2. Who was the director of this lm?
3. Are there any other lms by the same director where
Bruce Willis was a cast member? If so, which ones and
how many?
You have 5 minutes to complete this task."</p>
      </sec>
      <sec id="sec-11-7">
        <title>Task 4: More Data</title>
        <p>\Once again, before you start, please select the data
source called `Wikidata (CODE Edition)' (5th from
the top) in the `Search Linked Data' section.</p>
        <p>There is a music album that has the word `Antidote'
in it.
1. Who is the performer / musical artist of this album?
The one you are looking for starts with `Mor. . . '.
2. Which other albums has this artist released? Please
make sure that only albums (and no singles) are
displayed.
3. Make a MindMap containing all the information
that you just looked up.</p>
        <p>You have 5 minutes to complete this task."</p>
        <p>After each task was nished | either by the participant
successfully completing it, or by reaching the respective time
limit | the participants lled out a NASA Task Load Index
form and subjectively judged several aspects of the task they
had just worked on. The form consisted of the following
questions:</p>
        <p>Mental Demand. How mentally demanding was the
task?
Physical Demand. How physically demanding was
the task?
Temporal Demand. How hurried or rushed was the
pace of the task?
Performance. How successful were you in
accomplishing what you were asked to do?
E ort. How hard did you have to work to accomplish
your level of performance?
Frustration. How insecure, discouraged, irritated,
stressed, and annoyed were you?</p>
        <p>Additionally after each task the participants were asked
the following question: \Any comments? What was good
/ bad / unexpected / di cult?" Firstly they were asked
to write down what came to their minds. After that the
study conductor asked about speci c observations that he
had made during the task. After a usually short, sometimes
a little longer discussion, the participants added written
remarks that came up during the discussion.</p>
        <p>The nal page of the survey consisted of four questions
that gave the participants the opportunity to provide
additional qualitative feedback. These questions were:
What did you like about the Linked Data Query
Wizard?
What did you hate about the Linked Data Query
Wizard?
For which tasks would you personally use the Linked
Data Query Wizard?
If you could have solved the tasks with other tools of
your choice, which ones would you have used?</p>
        <p>After the participant had answered these nal questions,
the session was concluded.</p>
        <p>The four tasks that the participants had to solve were
basically divided into two groups:</p>
        <p>Tasks 1 and 2 concentrated on the \Show Available
Datasets" mode and were intentionally of moderate
complexity. Since all participants had, except for a
short guided tour, no prior experience with the Linked
Data Query Wizard, these tasks were intended to ease
them into the system and to discover potential
problems with the user experience at the same time.
Tasks 3 and 4 concentrated on the \Search Linked
Data" mode and were of signi cantly higher
complexity. By then the participants had become more familiar
with the Linked Data Query Wizard, since the learning
e ect should have already started to kick in.</p>
        <p>It was clear that the lack of randomization of the tasks
would lead to an uncompensated learning e ect. This did
not pose a problem under the circumstances, since it was
not the goal of this study to compare the di erent tasks
with each other, but rather to evaluate the general usefulness
of the system and nd its weak points with regard to user
experience.</p>
        <p>Apart from the quantitative feedback (NASA Task Load
Index) and the qualitative feedback (Retrospective Think
Aloud), the study conductor also measured the task
completion rate.</p>
        <p>Task completion time was not measured for several
reasons. For example, the study was conducted on the live
system and not on a lab setup, so uctuations in (external)
server response times were to be expected. Also, since the
Linked Data Query Wizard o ers a rather novel interface to
access Linked Data, the main goal of the study was to show
if users are able to use it at all and where potential problems
in understanding arise.</p>
        <p>Another idea was to go with a conventional Think Aloud
study instead of using the Retrospective Think Aloud
protocol. This would have had a negative impact on completion
time and, due to the given time limits for completing each
task, could have resulted in a lower task completion rate.</p>
        <p>Also, competing the Linked Data Query Wizard against
other tools was not really an option, since there are currently
no tools that could come close enough in functionality to
make a direct comparison feasible.</p>
        <p>Another possibility would have been to compare against
Google searches or SPARQL queries written by Semantic
Web experts. However, in the rst case, the test cases
could have been constructed in a way where the Linked
Data Query Wizard would have always won by a landslide,
which would have defeated the purpose of a direct
comparison. In the latter case, competing against manually created
SPARQL queries was not ideal either, since the focus of this
evaluation was on regular web users and not on Semantic
Web experts.
5.2</p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>Results and Discussion</title>
      <p>Due to the combination of the Restrospective Think Aloud
protocol with the NASA Task Load Index, it was possible
to generated four di erent kinds of results from the study:
Quantitative results (NASA Task Load Index)
Quantitative results (Task Completion Rate)
Comparison between participants with and without a
background in computer science</p>
      <p>Qualitative results (Retrospective Think Aloud)
5.2.1</p>
      <p>Quantitative Results (NASA Task Load Index)
The quantitative results of the NASA Task Load Index
can be seen in the box plots of Figure 8. The six plots
represent the results for the six di erent aspects of the NASA
Task Load Index. Those were mental demand, physical
demand, temporal demand, performance, e ort, and
frustration.</p>
      <p>The mental demand was rather low for the rst two
tasks and increased only slightly for the more complex
last two tasks. The variance between the participants
was quite high.</p>
      <p>The physical demand was, as expected, very low
throughout the study.</p>
      <p>The temporal demand | with respect to the time
limits of the tasks | basically corresponded with the
results from the mental demand, showing a generally
low demand with a high degree of variance between
the participants.</p>
      <p>The performance scores were very high with a median
of 10 out of 10 for all four tasks. Out of the 56 tasks
performed in total by the 14 participants, 49 were
successfully completed, 6 were not completed entirely in
time, and only 1 was not completed at all.</p>
      <p>The subjective e ort of the participants showed a high
variance between the participants, however it also showed
the learning e ect very nicely: The e ort necessary by
the participants decreased after the rst task, since the
second task was similar to the rst one. The third task
was completely di erent, which raised the level of
necessary e ort again. The fourth task was similar to the
third task, which again resulted in lower e ort.</p>
      <p>The frustration level was rather low throughout the
study, but again with a very high variance between
the participants.
5.2.2</p>
      <p>Quantitative Results (Task Completion Rate)
In addition to the subjective quantitative results
measured via the NASA Task Load Index, the task completion
rate was also measured objectively by the study conductor.
There was, however, no signi cant di erence between the
subjective performance as judged by the participants
themselves and the objectively measured task completion rate.
In detail, this means:
13 out of 14 participants were able to complete task
1 completely. 1 participant only received 2 out of 10
points.</p>
      <p>All 14 participants were able to solve task 2 completely
in time.</p>
      <p>Task 3 turned out to be the most di cult one: 10 out of
the 14 participants were able to solve it completely, the
other 4 participants only received 5 out of 10 points.
Task 4 was completely solved by 12 out of the 14
participants. 1 participant only received 5 out of 10
points, and 1 participant received 0 points.
5.2.3</p>
      <p>Background in Computer Science</p>
      <p>An interesting research question came up in the
preparation of the study: Would there be a signi cant di erence in
the results of participants with and participants without a
background in computer science? For this reason, 7 of the
study participants had a background in computer science,
whereas the other 7 did not.</p>
      <p>To determine if there was indeed a di erence between
these two groups, independent two-sample t-tests with equal
sample size were performed, comparing all 24 results of the
NASA Task Load Index (6 aspects * 4 tasks) as well as
the objectively measured task completion rates. The result
was that all calculated p-values were larger than 0.1. This
means that for our study, no signi cant di erence
regarding the results of the subjective NASA Task Load Index
or the objective task completion rate between participants
with and without a background in computer science could
be measured.
5.2.4</p>
      <p>Qualitative Results</p>
      <p>The qualitative results of the evaluation were based on
the statements of the participants collected during the
Retrospective Think Aloud phase of the study.</p>
      <p>Regarding task 1, the main problem that the participants
encountered concerned the ltering: To set a URI lter, the
participants had to click on the respective entity (in task 1,
it was the year 2010) and select \Add as lter" in its context
menu. However, 10 of the 14 participants had problems with</p>
      <p>Tasks
Performance</p>
      <p>Tasks
3
3
4
4</p>
      <p>Tasks
Effort
Tasks
3
3
4
4
8
6
4
2
0
8
6
4
2
0
10
1
1
2
2</p>
      <p>Tasks
Frustration</p>
      <p>Tasks
3
3
4
4
setting the lter because they expected it to be set through
a context menu item in the table header, as it is the case
for other lters (text, number and date) in the Linked Data
Query Wizard and in most current spreadsheet applications.</p>
      <p>Regarding task 2, the feedback was much more positive,
the use of the \Aggregate dataset" feature did not cause any
major problems.</p>
      <p>Regarding task 3, the combination of multiple lters turned
out to be quite a challenge for the participants. Also, the fact
that all cast members of the matching movies were displayed
even after \Bruce Willis" had been set as a lter confused
some users. Because all of the cast members were displayed
for each movie, this also meant that the rows became quite
high, which several participants found irritating.</p>
      <p>The majority of task 4 did not pose a problem for the
participants after having completed the similar third task.
The fact that the relevant search result did not appear on the
rst, but on the second result page, caused huge confusion
for almost all of the participants, even though the number
of total results and the \Load all results" button were visible
to all participants all of the time.</p>
      <p>When asked about what they liked about the Linked Data
Query Wizard, they general opinion was that once they had
worked out how the ltering worked, the interface was easy
enough to use. Additionally they liked how they could create
a useful list of results from a huge database with only a few
simple steps.</p>
      <p>When asked about what they didn't like about the Linked
Data Query Wizard, there was no general theme. Four of
the participants mentioned that it would have been hard for
them to choose a data source if they had not been told which
one to use.</p>
      <p>When asked about what they would personally use the
Linked Data Query Wizard for, no clear trend could be
recognized. The answers ranged from \don't know yet" and
\statistical data" to \newspaper entries" and \exploration of
data sources".</p>
      <p>Finally, when asked about which other tools they would
have used to solve similar tasks if the Linked Data Query
Wizard had not been available, it became very clear what
the direct competitors for the Linked Data Query Wizard
were: Almost every participant immediately mentioned that
they would use Google to search for data or information.
The majority of participants mentioned that they would also
use specialized portals to look for certain information, e.g.
IMDB for data about movies. When it came to working with
data, analyzing and visualizing it, almost all participants
mentioned that they would use Microsoft Excel and that
they would probably manually collect and copy the data
into the spreadsheet.
6.</p>
    </sec>
    <sec id="sec-13">
      <title>CONCLUSION &amp; FUTURE WORK</title>
      <p>In this paper we introduced the Linked Data Query
Wizard, a novel interface for accessing Linked Data in SPARQL
endpoints, either through a keyword search or by selecting
available Linked Datasets represented as RDF Data Cubes.</p>
      <p>The results of the conducted user study showed that the
tool had a few weak spots that could be improved, but was
in general very usable, both for people with and without a
background in computer science.</p>
      <p>In the future we plan to address the main challenges that
came up during the user study, mainly the possibility to add
all types of lters through the table header.</p>
      <p>Also the total number of results could be displayed more
prominently, and the implementation of an \in nite scroll"
mechanism that automatically displays more data as soon as
the users scroll to the bottom of the screen could circumvent
the problem that users need to load more data manually in
order to nd what they are looking for.</p>
      <p>
        Another important point for improvement is the current
limitation that users can only search one SPARQL endpoint
at a time, which they need to select beforehand. The
integration of services like Balloon Fusion [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] could help in
this regard, providing SPARQL rewriting based on collected
co-reference information combined with automatic endpoint
discovery, resulting in an intelligent query federation.
      </p>
    </sec>
    <sec id="sec-14">
      <title>ACKNOWLEDGEMENTS</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] G. Cheng, H. Wu,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ge</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qu</surname>
          </string-name>
          . Falcons Explorer:
          <article-title>Tabular and Relational End-user Programming for the Web of Data</article-title>
          .
          <source>In Semantic Web Challenge</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Reynolds</surname>
          </string-name>
          .
          <source>The RDF Data Cube Vocabulary</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O.</given-names>
            <surname>Erling</surname>
          </string-name>
          .
          <article-title>Faceted Views over Large-Scale Linked Data</article-title>
          .
          <source>Linked Data on the Web (LDOW)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cuddihy</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Ramey.</surname>
          </string-name>
          <article-title>The validity of the stimulated retrospective think-aloud method as measured by eye tracking</article-title>
          .
          <source>In Proceedings of the SIGCHI conference on Human Factors in computing systems - CHI '06, page 1253</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Harris</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Seaborne</surname>
          </string-name>
          .
          <source>SPARQL 1</source>
          .1
          <string-name>
            <given-names>Query</given-names>
            <surname>Language</surname>
          </string-name>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Hart</surname>
          </string-name>
          .
          <article-title>Nasa-task load index (nasa-tlx); 20 years later</article-title>
          .
          <source>In Proceedings of the Human Factors and Ergonomics Society Annual Meeting</source>
          , volume
          <volume>50</volume>
          , pages
          <fpage>904</fpage>
          {
          <fpage>908</fpage>
          .
          <string-name>
            <surname>Sage</surname>
            <given-names>Publications</given-names>
          </string-name>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Huynh</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Karger</surname>
          </string-name>
          .
          <article-title>Parallax and companion: Set-based browsing for the data web</article-title>
          .
          <source>WWW Conference</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Huynh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Karger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Exhibit: lightweight structured data publishing</article-title>
          .
          <source>Proceedings of the 16th international conference on World Wide Web, Ban</source>
          , Alb:
          <volume>737</volume>
          {
          <fpage>746</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mutlu</surname>
          </string-name>
          , P. Hoe er, G. Tschinkel,
          <string-name>
            <given-names>E.</given-names>
            <surname>Veas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sabol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stegmaier</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Granitzer</surname>
          </string-name>
          .
          <article-title>Suggesting Visualisations for Published Data</article-title>
          .
          <source>In Proceedings of IVAPP</source>
          <year>2014</year>
          , Lisbon, Portugal,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Oren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Delbru</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          .
          <article-title>Extending faceted navigation for RDF data</article-title>
          .
          <source>In The Semantic Web ISWC 2006 5th International Semantic Web Conference ISWC 2006 Athens GA USA November 59 2006 Proceedings</source>
          , volume
          <volume>4273</volume>
          , pages
          <fpage>559</fpage>
          {
          <fpage>572</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Schlegel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stegmaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Granitzer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Kosch</surname>
          </string-name>
          .
          <source>Balloon Fusion: SPARQL Rewriting Based on Uni ed Co-Reference Information. In 5th International Workshop on Data Engineering Meets the Semantic Web, co-located with the 30th IEEE International Conference on Data Engineering</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Seifert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Granitzer</surname>
          </string-name>
          , P. Ho er,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mutlu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sabol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Schlegel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stegmaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zwicklbauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Kern</surname>
          </string-name>
          .
          <article-title>Crowdsourcing Fact Extraction from Scienti c Literature</article-title>
          .
          <source>In Workshop on Human-Computer Interaction and Knowledge Discovery (SouthCHI)</source>
          , volume
          <volume>7947</volume>
          <source>of LNCS</source>
          , Maribor, Slovenia,
          <year>2013</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Stegmaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Seifert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Patrick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Granitzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lindstaedt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mutlu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sabol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Schlegel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Zwicklbauer</surname>
          </string-name>
          .
          <source>Unleashing Semantics of Research Data. In The Second Workshop on Big Data Benchmarking WBDB2012in</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tummarello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Delbru</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Oren. Sindice.</surname>
          </string-name>
          <article-title>com: Weaving the open linked data</article-title>
          .
          <source>Lecture Notes in Computer Science</source>
          ,
          <volume>4825</volume>
          :
          <fpage>552</fpage>
          {
          <fpage>565</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>M. Van Den Haak</surname>
            , M. De Jong, and
            <given-names>P. Jan</given-names>
          </string-name>
          <string-name>
            <surname>Schellens</surname>
          </string-name>
          .
          <article-title>Retrospective vs. concurrent think-aloud protocols: testing the usability of an online library catalogue</article-title>
          .
          <source>Behaviour &amp; Information Technology</source>
          ,
          <volume>22</volume>
          (
          <issue>5</issue>
          ):
          <volume>339</volume>
          {
          <fpage>351</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>