<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluating Semantic Search Systems to Identify Future Directions of Research?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Khadija Elbedweihy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stuart N. Wrigley</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Ciravegna</string-name>
          <email>f.ciravegnag@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dorothee Reinhard</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abraham Bernstein</string-name>
          <email>bernsteing@ifi.uzh.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Proceedings of the Second International Workshop on Evaluation of Semantic Technologies</institution>
          ,
          <addr-line>IWEST 2012</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Project SEALS, Semantic Evaluation at Large Scale</institution>
          ,
          <addr-line>FP7-238975</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of She eld</institution>
          ,
          <addr-line>Regent Court, 211 Portobello, She eld</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Zurich</institution>
          ,
          <addr-line>Binzmuhlestrasse 14, CH-8050 Zurich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <fpage>25</fpage>
      <lpage>36</lpage>
      <abstract>
        <p>Recent work on searching the Semantic Web has yielded a wide range of approaches with respect to the style of input, the underlying search mechanisms and the manner in which results are presented. Each approach has an impact upon the quality of the information retrieved and the user's experience of the search process. This highlights the need for formalised and consistent evaluation to benchmark the coverage, applicability and usability of existing tools and provide indications of future directions for advancement of the state-of-the-art. In this paper, we describe a comprehensive evaluation methodology which addresses both the underlying performance and the subjective usability of a tool. We present the key outcomes of a recently completed international evaluation campaign which adopted this approach and thus identify a number of new requirements for semantic search tools from both the perspective of the underlying technology as well as the user experience.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        the fourth employed a formal query approach which was hidden from the end user
by a graphical query interface. Recently, evaluating semantic search approaches
gained more attention both in IR { within its most established evaluation
conference TREC { [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as well as in the Semantic Web community (SemSearch [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
and QALD3 challenges).
      </p>
      <p>
        The above evaluations are all based upon the Cran eld methodology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]4:
using a test collection, a set of tasks and a set of relevance judgments. This
leaves aside aspects of user-oriented evaluations concerned with the usability of
the evaluated systems and the user experience which is as important as assessing
the performance of the systems. Additionally, the above attempts are separate
e orts lacking standardised evaluation approaches and measures. Indeed, Halpin
et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] note that \the lack of standardised evaluation has become a serious
bottleneck to further progress in this eld".
      </p>
      <p>The rst part of this paper describes an evaluation methodology for assessing
and comparing the strengths and weaknesses of user-focussed Semantic Search
approaches. We describe the dataset and questions used in the evaluation and
discuss the results of the usability study. The analysis and feedback from this
evaluation are described. The second part of the paper identi es a number of new
requirements for search approaches based upon the outcomes of the evaluation
and analysis of the current state-of-the-art.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Evaluation Design</title>
      <p>
        In the Semantic Web community, semantic search is widely used to refer to a
number of di erent categories of systems:
{ gateways (e.g., Sindice [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Watson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) locating ontologies and documents
{ approaches reasoning over data and information located within documents
and ontologies (PowerAqua [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and Freya [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ])
{ view-based interfaces allowing users to explore the search space while
formulating their queries (Semantic Crystal [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], K-Search [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Smeagol [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ])
{ mashups integrating data from di erent sources to provide rich descriptions
about Semantic Web objects (Sig.ma [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>The evaluation described here focuses on user-centric semantic search tools (e.g.
query given as keywords or natural language or using a form or a graph) querying
a repository of semantic data and returning answers extracted from them. The
tools' results presentation is not limited to a speci c style (e.g., list of entity
URIs or a visualisation of the results). However, the results returned must be
answers rather than documents matching the given query.</p>
      <p>Search is a user-centric activity; therefore, it is important to emphasise the
users' experience. An important aspect of this is the formal gathering of feedback
from the participants which should be achieved using standard questionnaires.
Furthermore, the use of an additional demographics questionnaire allows more</p>
      <sec id="sec-2-1">
        <title>3 http://www.sc.cit-ec.uni-bielefeld.de/qald-1</title>
        <p>4 http://www.sigir.org/museum/pdfs/ASLIB%20CRANFIELD%20RESEARCH%
20PROJECT-1960/pdfs/
in-depth ndings to be identi ed (e.g., if a particular type of user prefers a
particular search approach).
2.1</p>
        <sec id="sec-2-1-1">
          <title>Datasets and Questions</title>
          <p>
            Subjects are asked to reformulate a set of questions using a tool's interface.
Thus, it is important that the data set would be from an understandable and
well-known domain (and hence, easily understandable by non-expert users) and,
preferably, already have a set of questions and associated groundtruths. The
geographical dataset from the Mooney Natural Language Learning Data5 satis es
these requirements and has been used in a number of usability studies [
            <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
            ].
Although the Mooney dataset is di erent from ones currently found on the
Semantic Web such as DBpedia in terms of size, heterogeneity and quality, the
assessment of the tools ability to handle these aspects is not the focus of this
phase but rather the usability of the tools and the user experience.
          </p>
          <p>
            The questions [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] used in the rst evaluation campaign were generated based
on the existing templates within the Mooney dataset. These contained questions
with varying complexity and assessing di erent features. For instance, they
contained simple with only 1 unknown concept such as \Give me all the capitals
of the USA?" and comparative questions such as \Which rivers in Arkansas are
longer than Aleghany river".
2.2
          </p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Criteria and Analyses</title>
          <p>
            Usability Di erent input styles (e.g., form-based, NL, etc.) can be compared
with respect to the input query language's expressiveness and usability. These
concepts are assessed by capturing feedback regarding the user experience and
the usefulness of the query language in supporting users to express their
information needs and formulate searches [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. Additionally, the expressive power of
a query language speci es what queries a user is able to pose [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. The
usability is further assessed with respect to results presentation and suitability of the
returned answers (data) to the casual users as perceived by them. The datasets
and associated questions were designed to fully investigate these issues.
Performance Users are familiar with the performance of commercial search
engines (e.g., Google) in which results are returned within fractions of a second;
therefore, it is a core criterion to measure the tool's performance with respect
to the speed of execution.
          </p>
          <p>Analyses The experiment was controlled using custom-written software which
allowed each experiment run to be orchestrated and timings and results to be
captured. The results included the actual result set returned by a tool for a
query, the time required to execute a query, the number of attempts required</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>5 http://www.cs.utexas.edu/users/ml/nldata.html</title>
        <p>
          by a user to obtain a satisfactory answer as well as the time required to
formulate the query. We used post-search questionnaires to collect data regarding the
user experience and satisfaction with the tool. Three di erent types of online
questionnaires were used which serve di erent purposes. The rst is the System
Usability Scale (SUS) questionnaire [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The test consists of ten normalized
questions and covers a variety of usability aspects, such as the need for support,
training, and complexity and has proven to be very useful when investigating
interface usability [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. We developed a second, extended, questionnaire which
includes further questions regarding the satisfaction of the users. This
encompasses the design of the tool, the input query language, the tool's feedback, and
the user's emotional state during the work with the tool. Finally, a demographics
questionnaire collected information regarding the participants.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Execution and Results</title>
      <p>
        The evaluation consisted of tools from form-based, controlled-NL-based and
freeNL-based approaches. Each tool was evaluated with 10 subjects (except K-Search
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] which had 8) totalling 38 subjects (26 males, 12 females) aged between 20
and 35 years old. They consisted of 28 students and 10 researchers drawn from
the University population. Subjects rated their knowledge of the Semantic Web
with 6 reporting their knowledge to be advanced, 5 good, 9 average, 10 little
and 8 having no experience. In addition, their knowledge of query languages was
recorded, with 5 stating their knowledge to be advanced, 12 good, 8 average, 6
little and 7 having no experience.
      </p>
      <p>Firstly, the subjects were presented with a short introduction to the
experiment itself such as why the experiment is taking place, what is being tested, how
the experiment will be executed, etc. Then the tool itself was explained to the
subjects; they learnt about the type and the functionality of the tool and how to
apply it's speci c query language to answer the given tasks. The users were then
given sample tasks to test their understanding of the previous phases. After that,
the subjects did the actual experiment: using the tool's interface to formulate
each question and get the answers. Having nished all the questions, they were
presented with the three questionnaires (Section 2.2). Finally, the subjects had
the chance to talk about important and open questions and give more feedback
and input to their satisfaction or problems with the system being tested.</p>
      <p>Table 1 shows the results for the four tools participating in this phase. The
mean number of attempts shows how many times the user had to reformulate
their query in order to obtain answers with which they were satis ed (or indicated
that they were con dent a suitable answer could not be found). This latter
distinction between nding the appropriate answer and the user `giving up' after
a number of attempts is shown by the mean answer found rate. Input time refers
to the amount of time the subject spent formulating their query using the tool's
interface, which acts as a core indicator of the tool's usability.</p>
      <p>
        According to the ratings of SUS scores [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], none of the four
participating tools fell in either the best or worst category. Only one of the tools
(PowerAqua [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) had a `Good' rating with a SUS score of 72.25, other two tools
Mean experiment time (s)
Mean SUS (%)
Mean ext. questionnaire (%)
Mean number of attempts
Mean answer found rate
Mean execution time (s)
Mean input time (s)
      </p>
      <p>K-Search
Form-based</p>
      <p>
        NLP-Reduce PowerAqua
NL-based NL-based
(Ginseng [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and K-Search [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) fell in the `Poor' rating while the last one
(NlpReduce [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]) was classi ed as `Awful'. The results of the questionnaires were
con rmed by the recorded usability measures. Subjects using the tool with the
lowest SUS score (Nlp-Reduce) required more than twice the number of attempts
of the other tools before they were satis ed with the answer or moved on.
Similarly, subjects using the two tools with the highest SUS and extended scores
(PowerAqua and K-Search) found satisfying answers to their queries twice the
times as for the other tools. Altogether, this con rms the reliability of the results
and the feedback of the users and also the conclusions based on them.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Usability Feedback and Analysis</title>
      <p>This section discusses the results and feedback collected from the subjects of the
usability study. Figure 1 summarises the features most liked and disliked based
on their feedback.
4.1</p>
      <sec id="sec-4-1">
        <title>Input Style</title>
        <p>
          On the one hand, Uren et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] state that forms can be helpful to explore the
search space when it is unknown to the users. Additionally, Corese [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] { which
uses a form-based interface to allow users to build their queries { received very
positive comments from its users among which was an appreciation for its
formbased interface. On the other hand, Lei et al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] see this exploration as a burden
on users that requires them to be (or become) familiar with the underlying
ontology and semantic data. The results of our evaluation and the feedback
from the users support both arguments. Additionally, we found that form-based
interfaces allow users to build more complex queries than the natural language
interfaces. However, building queries by exploring the search space is usually time
consuming especially as the ontology gets larger or the query gets more complex.
This was shown by Kaufmann et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] in their usability study which found that
users spent the most time when working with the graph-based system Semantic
Crystal. Our evaluation supports this general conclusion: subjects using the
formbased approach took between two to three times the time taken by users of
natural language approaches. However, our analysis suggests a more nuanced
behaviour. While freeform natural language interfaces are generally faster in
terms of query formulation, we found this did not hold for approaches employing
Input  Style  
 
 
 
 
 
Query  
Execu9on  
Results  
Presenta9on  
        </p>
        <p>Liked/Required  
View  search  
domain  </p>
        <p>Auto-­‐  
compleEon  </p>
        <p>Build  complex  
queries  (AND,  </p>
        <p>OR,…  )  
Easy  &amp;  fast  </p>
        <p>input  
Natural  &amp;  familiar  </p>
        <p> language  
Feedback  during  
query  execuEon  
Merging  
results  </p>
        <p>Show  
provenance  of  
results  </p>
        <p>Input  format  
complexity  </p>
        <p>No    support  for  
superlaEves  or  
comparaEves  in  </p>
        <p>queries  
Slow  response  </p>
        <p>Disliked  
Restricted  
language  
model  </p>
        <p>Requires  
knowledge  </p>
        <p>of  
ontologies  
AbstracEon  of  
search  domain    
No  incremental  </p>
        <p>results  
Not  suitable  for  
casual  users  
No  storing/
reuse  of  query  
results  </p>
        <p>No  sorEng,  
grouping,  or  
filtering  of  results  
a very restricted language model. For instance, query formulation took longer
using Ginseng (restricted natural language) than K-Search (form-based). This
is further supported by user feedback in which it was reported that they would
prefer typing a natural language query because it is faster than forms or graphs.</p>
        <p>
          Kaufmann et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] also showed that a natural language interface was judged
by users to be the most useful and best liked. Their conclusion, that this was
because users can communicate their information needs far more e ectively when
using a familiar and natural input style, is supported by our ndings. The same
study found that people can express more semantics when they use full sentences
as opposed to simply keywords. Similarly, Demidova et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] state that natural
language queries o er users more expressivity to describe their information needs
than keywords { a nding also con rmed by the user feedback from our study.
        </p>
        <p>However, natural language approaches su er from both syntactic as well as
semantic ambiguities. This makes the overall performance of such approaches
heavily dependent upon the performance of the underlying natural language
processing techniques responsible for parsing and analysing the users' natural
language sentences. This was shown by the feedback we received from users of
one of the natural language-based tools, one of which was \the response is very
dependent on the use of the correct terms in the query". This was also con rmed
by that approach achieving the lowest precision. Another limitation faced by the
natural language approach is the lack of knowledge of the underlying ontology
terms and relations by the users due to the high abstraction of the search domain.
The e ect of this is that any keywords or terms used by users are likely to be very
di erent from the semantically-corresponding terms in the ontology. This in turn
increases the di culty of parsing the user query and a ects the performance.</p>
        <p>
          Using a restricted grammar as employed by Ginseng is an approach to limit
the impact of both of these problems. The `autocompletion' provided by the
system based on the underlying grammar attempts to bridge the domain
abstraction gap and also resembles the form-based approach in helping the user to
better understand the search space. Although it provides the user with
knowledge regarding which concepts, relations and instances are found in the search
space and hence can be used to build valid queries, it still lacks the power of
visualising the structure of the used ontology. The impact of this
`intermediate' functionality can be observed in the user feedback with a lower degree of
dissatisfaction regarding the ability to conceptualise the underlying data but
still not completely eliminated. The restricted language model also prevents
unacceptable/invalid queries in the used grammar by employing a guided input
natural language approach. However, only accepting speci c concepts and
relations { found in the grammar { limits the exibility and expressiveness of the
user queries. User coercion into following prede ned sentence structures proves
to be frustrating and too complicated [
          <xref ref-type="bibr" rid="ref1 ref24">1, 24</xref>
          ].
        </p>
        <p>The feedback from the questionnaires showed that using superlatives or
comparatives in the user queries (e.g.: highest point, longer than) was not supported
by any of the participating tools; an issue raised by 8 subjects in the answer of
the SUS question \What didn't you like about the system and why?" and by
others in the open feedback after the experiment. Only one provided a feature
similar to this functionality: the ability to specify a range of values for numeric
datatypes. A comparative such as less than 5000 could then be translated to
the range 0 to 5000. However, this was deemed to be both confusing (since
the user had to decide what to use as the non-speci ed bound) and, when the
non-speci ed bounds were incorrect, having a negative impact on the results.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Query Execution and Response Time</title>
        <p>Speed of response is an important factor for users since they are used to the
performance of commercial search engines (e.g., Google) in which results are
returned within fractions of a second. Many users in our study were
expecting similar performance from the semantic search tools. Although the average
response time of three of the tools (K-Search, NLP-Reduce, Ginseng) is less
than a second (44ms, 51ms, and 51ms respectively), users reported their
dissatisfaction with these timings especially the ones who evaluated PowerAqua with
response time of 11 seconds on average. The lack of feedback on the status of
the execution process only served to increase the sense of dissatisfaction: no tool
indicated the execution progress or whether a problem had occurred in the
system. This lack of feedback resulted in users suspecting that something had gone
wrong with the system { even if the search was still progressing{ and start a
new search. Furthermore, some tools made it impossible to distinguish between
an empty result set, a problem with the query formulation or a problem with
the search. This not only a ected the users experience and satisfaction but also
the approach's measured performance.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Results Presentation</title>
        <p>Semantic Search tools are di erent from Semantic Web gateways or entry points
such as Watson and Sindice. The latter are not intended for casual users but
for other applications or the Semantic Web community to locate Semantic Web
resources such as ontologies or Semantic Web documents and are usually
presented as a set of URIs. For example, Sindice shows the URIs of documents and,
for every document, it additionally presents the triples contained within the
document, an RDF graph of the triples, and the used ontologies. Semantic Search
tools are, on the other hand, used by casual users (i.e., users who may be experts
in the domain of the underlying data but may have no knowledge of semantic
technologies). Such users usually have di erent requirements and expectations
of what and how results should be presented to them.</p>
        <p>In contrast to these `casual user' requirements, a number of the search tools
did not present their results in a user-friendly manner and this was re ected in
the feedback. Two approaches presented the full URIs together with the concepts
in the ontology that were matched with the terms in the user query. Another used
the instance labels to provide a natural language presentation; however, such
labels (e.g., `montgomeryAl') were not necessarily suitable for direct inclusion
into a natural language phrase. Indeed, the tool also displayed the ontologies
used as well as the mappings that were found between the ontology and the
query terms. Although potentially useful to an expert in the semantic web eld,
this was not helpful to casual users.</p>
        <p>The other commonly reported limitation of all the tools was the degree to
which a query's results could be stored or reused. A number of the questions used
in the evaluation had a high complexity level and needed to be split into two
or more sub-queries. For instance, for the question \Which rivers in Arkansas
are longer than the Allegheny river?", the users were rst querying the data for
the length of the Allegheny river and then performing a second query to nd
the rivers in Arkansas which are longer than the answer they got. Therefore,
users often wanted to use previous results as the basis of a further query or
to temporarily store the results in order to perform an intersection or union
operation with the current result set. Unfortunately, this was not supported by
any of the participating tools. However, this shows that users have very high
expectations of the usability and functionalities o ered by a semantic search
tool as this requirement is not provided even by traditional search systems (e.g.,
Google and Yahoo). Another means of managing the results that users requested
was the ability to lter results according to some suitable criteria and checking
the provenance of the results; only one tool provided the latter. Indeed, even basic
manipulations such as sorting were requested { a feature of particular importance
for tools which did not allow query formulations to include superlatives.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Future Directions</title>
      <p>This section identi es a number of areas for improvement for semantic search
tools from the perspective of the underlying technology and the user experience.
5.1
Usability The feedback shows that it's very helpful for users { especially those
who are unfamiliar with the underlying data { to explore the search space while
building their queries using view-based interfaces which expose the structure of
the ontology in a graphical manner. It gives users a much better understanding
of what information is available and what queries are supported by the tool.
In contrast, the feedback also shows that, when creating their queries, users
prefer natural language interfaces because they are quick and easy. Clearly both
approaches have their advantages; however, they su er from various limitations
when used separately as discussed in Sec. 4.1. Therefore, we believe that the
combination of both approaches would help get the best of both worlds.</p>
      <p>Users not familiar with the search domain can use a form-based or natural
language-based interface to build their queries. Simultaneously, the tool should
dynamically generate a visual representation of the user's query based upon the
structure of the ontology. Indeed, the user should be able to move from one
query formulation style to another { at will { with each being updated to re ect
changes made in the other. This `dual' query formulation would ensure a casual
user correctly formulates their intended query. Expert users, or those who nd
it laborious to use the visual approach, would simply use the natural language
input facility provided by the tool. An additional feature for natural language
input would be an optional `auto-completion' feature which could guide the user
to query completion given knowledge of the underlying ontology.
Expressiveness The feedback also shows that the evaluated tools had
difculties with supporting complex queries such as the ones containing logical
operators (e.g, \AND"). Allowing the user to input more than one query and
combining them with their chosen logical operator from a list included in the
interface would reduce the impact of this limitation. The tool would merge the
results according to the used operator (e.g., \intersection" for \AND"). For
instance, a query such as \What are the rivers that pass through California and
Arizona?" would be constructed as two subqueries: \What are the rivers that
pass through California?" and \What are the rivers that pass through Arizona?"
with the nal results being the intersection of both result sets.</p>
      <p>
        Furthermore, the evaluated tools faced similar di culties with supporting
superlatives and comparatives in users' queries. Freya [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] deals with this problem
by asking the user to identify the correct choice from a list of suggestions. To
illustrate this we'll use the query \Which city has the largest population in
California?". If the system captures a concept in the user query that is a datatype
property of type number, it generates maximum, minimum and sum functions.
The user can then choose the correct superlative or comparative depending on
their needs. A similar approach can be used to allow the use of superlatives
and comparatives in natural language interfaces and form-based interface. In
the case of the latter, whenever a datatype property is selected by the user, the
tool can allow them to select from a list of functions that cover superlatives and
comparatives (e.g., `maximum', `minimum', `more than', `less than').
      </p>
      <sec id="sec-5-1">
        <title>5.2 Query Execution and Response Time</title>
        <p>
          Several users reported dissatisfaction with the tools' response time to some of
their queries. Users appreciated the fact that the tools returned more accurate
answers than they would get from traditional search engines, however this did
not remove the e ect of the delay in response { even if it was relatively small.
Additionally, the study found that the use of feedback reduces the e ect of the
delay; users showed greater willingness to wait if they were informed that the
search is still being performed and that the delay is not due to a failure in the
system. The presentation of intermediate, or partially complete, results reduces
the perceived delay associated with the complete result set (e.g., Sig.ma [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]).
Although only partial results are available initially, it provides both feedback
that the search is executing properly and allows the user to start thinking about
the content of the results before the complete set is ready. However, it ought
to be noted that this approach may induce confusion in the user as the screen
content changes rapidly for a number of seconds. Adequate feedback is essential
even for tools which exhibit high performance and good response times. Delays
may occur at a number of points in the search process and may be the result of
in uences beyond the developer's control (e.g., network communication delays).
5.3
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Results Presentation</title>
        <p>Most of the users were frustrated by the fact that they didn't understand the
results presented to them, feeling that too much technical knowledge was
assumed. The evaluation shows that the tools underestimated the e ect of this on
the users' experience and satisfaction.</p>
        <p>Query answers ought to be presented to users in an accessible and attractive
manner. Indeed, the tool should go a step further and augment the direct answer
with associated information in order to provide a `richer' experience for the
user. This approach is adopted by WolframAlpha6; for example, in response to
`What is the capital of Alabama? ' WolframAlpha includes the natural language
presentation of the answer as well as various population statistics, a map showing
the location of the city, and other related information such as the current local
time, weather and nearby cities.</p>
        <p>
          An interesting requirement found by our study was the ability to store the
result set of a query to use in subsequent queries. This would allow more
complex questions to be answered which, in turn, improves the tools' performance.
QuiKey [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] provides a functionality similar to this. QuiKey is an interaction
approach that o ers interactive ne grained access to structured information
sources in a lightweight user interface. It allows a query to be saved which can
later be used for building other queries. More complex queries can be constructed
by combining saved queries with logical operators such as `AND' and `OR'.
        </p>
        <p>
          Result management was also identi ed as being of importance to users with
commonly requested functionality included sorting, ltering and more complex
activities such as establishing the provenance and trustworthiness of certain
results. For example, Sig.ma [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] creates information aggregates called Entity
Proles and provides users with various capabilities to organise, use and establish
the provenance of the results. Users can see all the sources contributing to a
speci c pro le and approve or reject certain ones thus ltering the results. They can
6 http://www.wolframalpha.com/
also check which values in the pro le are given by a speci c source thus checking
provenance of the results. Sig.ma also supports the aspect of merging separate
results by allowing users to view ones returned only from selected sources.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>We have presented a exible and comprehensive methodology for evaluating
di erent semantic search approaches; we have also highlighted a number of
empirical ndings from an international semantic search evaluation campaign based
upon this methodology. Finally, based upon analysis of the evaluation outcomes,
we have described a number of additional requirements for current and future
semantic search solutions.</p>
      <p>
        In contrast to other benchmarking e orts, we emphasised the need for an
evaluation methodology which addressed both performance and usability [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
We presented the core criteria that must be evaluated together with a discussion
of the main outcomes. This analysis identi ed two core ndings which impact
upon semantic search tool requirements.
      </p>
      <p>Firstly, we found that an intelligent combination of natural language and
view-based input styles would provide a signi cant increase in search e
ectiveness and user satisfaction. Such a `dual' query formulation approach would
combine the ease with which a view-based approach can be used to explore and
learn the structure of the underlying data whilst still being able to exploit the
e ciency and simplicity of a natural language interface.</p>
      <p>Secondly, (and perhaps of greatest interest to users) was the need for more
sophisticated results presentation and management. Results should allow a large
degree of customisability (sorting, ltering, saving of intermediate results,
augmenting, etc). Indeed, it would also be bene cial to provide data which is
supplementary to the original query to increase `richness'. Furthermore, users expect
to be able to have immediate access to provenance information.</p>
      <p>In summary, this paper has presented a number of important ndings which
are of interest both to semantic search tool developers but also designers of
interactive search evaluations. Such evaluations (and the associated analyses as
presented here) provide the impetus to improve search solutions and enhance
the user experience.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , E.:
          <article-title>Talking to the Semantic Web | Natural Language Query Interfaces for Casual End-Users</article-title>
          .
          <source>PhD thesis</source>
          , University of Zurich (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Balog</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serdyukov</surname>
          </string-name>
          , P.,
          <string-name>
            <surname>de Vries</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <article-title>Overview of the TREC 2011 Entity Track</article-title>
          . In: TREC 2011 Working Notes
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Halpin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herzig</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mika</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blanco</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pound</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>H.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
          </string-name>
          , D.T.:
          <article-title>Evaluating Ad-Hoc Object Retrieval</article-title>
          .
          <source>In: Proc. IWEST 2010 Workshop</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cleverdon</surname>
            ,
            <given-names>C.W.:</given-names>
          </string-name>
          <article-title>Report on the rst stage of an investigation onto the comparative e ciency of indexing systems</article-title>
          .
          <source>Technical report</source>
          , The College of Aeronautics, Cran eld,
          <source>England</source>
          (
          <year>1960</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Tummarello</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oren</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delbru</surname>
          </string-name>
          , R.: Sindice.
          <article-title>com: Weaving the Open Linked Data</article-title>
          .
          <source>In: Proc. ISWC/ASWC 2007</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>d'Aquin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldassarre</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gridinoc</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angeletou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>Characterizing Knowledge on the Semantic Web with Watson</article-title>
          .
          <source>In: EON</source>
          . (
          <year>2007</year>
          )
          <volume>1</volume>
          {
          <fpage>10</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uren</surname>
          </string-name>
          , V.:
          <article-title>PowerAqua: Fishing the Semantic Web</article-title>
          .
          <source>In: The Semantic Web: Research and Applications</source>
          . (
          <year>2006</year>
          )
          <volume>393</volume>
          {
          <fpage>410</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Damljanovic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agatonovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
          </string-name>
          , H.:
          <article-title>Natural Language Interface to Ontologies: combining syntactic analysis and ontology-based lookup through the user interaction</article-title>
          .
          <source>In: Proc. ESWC</source>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaufmann</surname>
            , E., Gohring,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiefer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Querying Ontologies: A Controlled English Interface for End-users</article-title>
          .
          <source>In: Proc. ISWC</source>
          <year>2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Bhagdev</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lanfranchi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>In: Proc.</article-title>
          .
          <source>ESWC 2008</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Clemmer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davies</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Smeagol: A speci c-to-general semantic web query interface paradigm for novices</article-title>
          .
          <source>In: Proc. DEXA</source>
          <year>2011</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Tummarello</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catasta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danielczyk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delbru</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Sig.ma:
          <article-title>live views on the web of data</article-title>
          .
          <source>In: Proc. WWW</source>
          <year>2010</year>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Wrigley</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbedweihy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reinhard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>D13</source>
          .
          <article-title>3 Results of the rst evaluation of semantic search tools</article-title>
          .
          <source>Technical report, SEALS Consortium</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Uren</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giordanino</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The usability of semantic search tools: a review</article-title>
          .
          <source>The Knowledge Eng. Rev</source>
          .
          <volume>22</volume>
          (
          <year>2007</year>
          )
          <volume>361</volume>
          {
          <fpage>377</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Angles</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutierrez</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The Expressive Power of SPARQL</article-title>
          .
          <source>In: Proc. ISWC</source>
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Brooke</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>SUS: a quick and dirty usability scale</article-title>
          . In: Usability Evaluation in Industry. (
          <year>1996</year>
          )
          <volume>189</volume>
          {
          <fpage>194</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bangor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kortum</surname>
            ,
            <given-names>P.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          :
          <article-title>An Empirical Evaluation of the System Usability Scale</article-title>
          .
          <source>Int't J. Human-Computer Interaction</source>
          <volume>24</volume>
          (
          <issue>6</issue>
          ) (
          <year>2008</year>
          )
          <volume>574</volume>
          {
          <fpage>594</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Bangor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kortum</surname>
            ,
            <given-names>P.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          :
          <article-title>Determining what individual SUS scores mean: Adding an adjective rating scale</article-title>
          .
          <source>J. Usability Studies</source>
          <volume>4</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <volume>114</volume>
          {
          <fpage>123</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Querying the Semantic Web with Ginseng: A Guided Input Natural Language Search Engine</article-title>
          .
          <source>In: Proc. WITS 2005 Workshop</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <string-name>
            <surname>NLP-Reduce</surname>
          </string-name>
          :
          <article-title>A \nave" but Domainindependent Natural Language Interface for Querying Ontologies</article-title>
          .
          <source>In: Proc. ESWC</source>
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Corby</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dieng-Kuntz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faron-Zucker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gandon</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>Searching the Semantic Web: Approximate Query Processing Based on Ontologies. IEEE Intelligent Systems</source>
          <volume>21</volume>
          (
          <year>2006</year>
          )
          <volume>20</volume>
          {
          <fpage>27</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Lei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uren</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>SemSearch: A Search Engine for the Semantic Web</article-title>
          .
          <source>In: Proc. EKAW2006</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Demidova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nejdl</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Usability and Expressiveness in Database Keyword Search : Bridging the Gap</article-title>
          .
          <source>In: Proc. VLDB PhD Workshop</source>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wrigley</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbedweihy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reinhard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Evaluating Semantic Search Tools using the SEALS platform</article-title>
          .
          <source>In: Proc. IWEST 2010 Workshop</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Haller</surname>
          </string-name>
          , H.:
          <article-title>QuiKey - An E cient Semantic Command Line</article-title>
          .
          <source>In: Proc. EKAW</source>
          <year>2010</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>