<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Entity-based Data Source Contextualization for Searching the Web of Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreas Wagnery</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Haasez</string-name>
          <email>peter.haase@fluidops.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Achim Rettingery</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Holger Lammz</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>To allow search on the Web of data, systems have to combine data from multiple sources. However, to e ectively ful ll user information needs, systems must be able to \look beyond" exactly matching data sources and o er information from additional/contextual sources (data source contextualization). For this, users should be involved in the source selection process { choosing which sources contribute to their search results. Previous work, however, solely aims at source contextualization for \Web tables", while relying on schema information and simple relational entities. Addressing these shortcomings, we exploit work from the eld of data mining and show how to enable Web data source contextualization. Based on a real-world use case, we built a prototype contextualization engine, which we integrated in a system for searching the Web of data. We empirically validated the e ectiveness of our approach { achieving performance gains of up to 29% over the state-of-the-art.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>e s : data / t e c 0 0 0 1
r d f : type qb : O b s e rv a t i o n ;
es prop : geo es d i c :DE;
es prop : u n i t es d i c : MIO EUR ;
es prop : i n d i c n a es d i c : B11 ;
sd : time "2010 01 01" ^^ xs : date ;
sm : obsValue " 2 4 9 6 2 0 0 . 0 " ^^ xs :
double .</p>
      <p>e s : data / g o v q g g d e b t
r d f : type qb : O b s e rv a t i on ;
es prop : geo es d i c :DE;
es prop : u n i t es d i c : MIO EUR ;
es prop : i n d i c n a es d i c : F2 ;
sd : time "2010 01 01" ^^ xs : date ;
sm : obsValue " 1 7 8 6 8 8 2 . 0 " ^^ xs :</p>
      <p>d e c i m a l .</p>
      <p>Src. 3. NY.GDP.MKTP.CN (Worldbank).
wbi :NY.GDP.MKTP.CN
r d f : type qb : O b s e rv a t i on ;
sd : r e f A r e a wbi : c l a s s i f i c a t i o n / country /DE;
sd : r e f P e r i o d "2010 01 01" ^^ xs : date ;
sm : obsValue " 2 5 0 0 0 9 0 . 5 " ^^ xs : double ;
wbprop : i n d i c a t o r wbi : c l a s s i f i c a t i o n / i n d i c a t o r /NY.GDP.MKTP.CN .</p>
      <p>Note, contextual sources are actually not relevant to the user's query, but
relevant to her information need. Thus, integration of these \additional" sources
provides a user with broader results in terms of result dimensions (schema
complement) and result entities (entity complement). See our example in Fig. 1.</p>
      <p>For enabling systems to identify and integrate sources for contextualization,
we argue that user involvement during source selection is a key factor. That is,
starting with an initial search result (obtained via, e.g., a SPARQL or keyword
query), a user should be able to choose and change sources, which are used
for result computation. In particular, users should be recommended contextual
sources at each step of the search process. After modifying the selected sources,
the results may be reevaluated and/or the query expanded.</p>
      <p>
        Unfortunately, recent work on data source contextualization focuses on Web
tables [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], while using top-level schema such as Freebase. Further, the authors
restrict data to a simple relational form. We argue that such a solution is not a
good t for the schemaless, heterogeneous Web of data.
      </p>
      <p>
        Contributions. This paper continues our work in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. More speci cally,
we provide the following contributions: (1) Previous work on Web table
contextualization su ers from inherent drawbacks. Most notably, it is restricted to
xed-structured relational data and requires external schemata to provide
additional information. Omitting these shortcomings, we present an entity-based
solution for data source contextualization in the Web of data. Our approach is
based on well-known data mining strategies and does not require schema
information or data adhering to a particular form. (2) We implemented our system,
the data-portal, based on a real-world use case, thereby showing its practical
relevance and feasibility. A prototype version of this portal is publicly available
and is currently tested by a pilot customer.3 (3) We conducted two user studies
to empirically validate the e ectiveness of the proposed approach: our system
outperforms the state-of-the-art with up to 29%.
      </p>
      <sec id="sec-1-1">
        <title>3 http://data.fluidops.net/</title>
        <p>Outline. In Sect. 2, we present our use case. We give preliminaries in Sect. 3,
outline the approach in Sect. 4, and present its implementation in Sect. 5. We
discuss the evaluation in Sect. 6, related work in Sect. 7, and conclude in Sect. 8.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Use Case Scenario</title>
      <p>In this section, we introduce a real-world use case to illustrate challenges w.r.t.
data source selection during searching the Web of data. The scenario is provided
by a pilot user in the nancial industry and available online.4</p>
      <p>In their daily work, nancial researchers heavily rely on a variety of open
and closed Web data sources, in order to provide prognoses of future trends. A
typical example is the analysis of government debt. During the nancial crisis
in 2008-2009, most European countries made high debts. To lower doubts about
repaying these debts, most countries set up a plan to reduce their public
budget de cits. To analyze such plans, a nancial researcher requires an overview of
public revenue and expenditure in relation to the gross domestic product (GDP).
To measure this, she needs information about the de cit target, the
revenue/expenditure/de cit, and GDP estimates. This information is publicly available,
provided by catalogs like Eurostat and Worldbank. However, it is spread across a
huge space of sources. That is, there is no single source satisfying her information
needs { instead data from multiple sources has to be identi ed and combined.
To start her search process, a researcher may give \gross domestic product" as
keyword query. The result is GDP data from a large number of sources. At this
point, data source selection is \hidden" from the researcher, i.e., sources are
solely ranked via number and quality of keyword hits. However, knowing where
her information comes from is critical. In particular, she may want to restrict
and/or know the following meta-data:
{ General information about the data source, e.g., the name of the author and
a short description of the data source contents.
{ Information about entities contained in the data source, e.g., the single
countries of the European Union.
{ Description about the dimensions of the observations, e.g., the covered time
range or the data unit of the observations.</p>
      <p>By means of faceted search, the researcher nally restricts her data source
to tec00001 (Eurostat, Fig. 1) featuring \Gross domestic product at market
prices". However, searching the data source space in such a manner requires
extensive knowledge. Further, the researcher was not only interested in plain GDP
data { she was also looking for additional information.</p>
      <p>For this, a system should suggest data sources that might be of interest, based
on sources known to be relevant. These contextual sources may feature related,
additional information w.r.t. current search results/sources. For instance, data
sources containing information about the GDP of further countries or with a
di erent temporal range. This way, the researcher may discover new sources
more easily, as one source of interest links to another { allowing her to explore
the space of sources.</p>
      <sec id="sec-2-1">
        <title>4 http://data.fluidops.net/resource/Demo_GDP_Germany</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Preliminaries</title>
      <sec id="sec-3-1">
        <title>Data Model. Data in this work is represented as RDF:</title>
        <p>De nition 1 (RDF Triple, RDF Graph). Given a set of URIs U , blank
nodes B, and a set of literals L, t = hs; p; oi 2 U [ B U (U [ L [ B) is a RDF
triple. A RDF graph G = (V; E ) is de ned by vertices V = U [ L [ B and a set
of triples as edges E = fhs; p; oig.</p>
      </sec>
      <sec id="sec-3-2">
        <title>A data source may contain one or more RDF graphs:</title>
        <p>De nition 2 (RDF Data Source). A data source Di 2 D is a set of n RDF
graphs, i.e., Di = fG1i ; : : : ; Gnig, with D as set of all sources.</p>
        <p>Notice, the above de nition abstracts from the data access, e.g., via HTTP
GET requests. In particular, Def. 2 also covers Linked Data sources, see Fig. 1.</p>
        <p>
          Entity Model. Given a data source Di, an entity e is a subject that is
identi ed with a URI de in Gji. Entity e is described via a connected subgraph of
Gji containing de (called Ge). Subgraphs Ge, however, may be de ned in di erent
ways [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. For our work, we used the concise bound description, where all triples
with subject de are comprised in Ge. Further, if an object is a blank node, all
triples with that blank node as subject are also included and so on [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Example. We have 3 subjects in Fig. 1 and each of them stands for an entity.
Every entity description, Ge, is a one-hop graph. For instance, the description
for entity es:data/tec0001 comprises all triples in Src. 1.</p>
        <p>
          Kernel Functions. We compare di erent entities by comparing their
descriptions, Ge. For this, we make use of kernel functions [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]:
De nition 3 (Kernel function). Let : X X 7! R denote a kernel function
such that (x1; x2) = h'(x1); '(x2)i, where ' : X 7! H projects a data space X
to a feature space H and h ; i refers to the scalar product.
        </p>
        <p>Note, ' is not restricted, i.e., a kernel is constructed without prior knowledge
about '. We will give suitable kernels for entity descriptions, Ge, in Sect. 4.</p>
        <p>
          Clustering. We use k-means [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] as a simple and scalable algorithm for
discovering clusters, Ci, of entities in data sources D. It works as follows [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]:
(1) Choose k initial cluster centers, mi. (2) Based on a dissimilarity function,
dis, an indicator function is given as: 1(e; Ci) is 1 if dis(e; Ci) &lt; dis(e; Cj ); 8j 6= i
and 0 otherwise. That is, 1(e; Ci) assigns each entity e to its \closest" cluster Ci.
(3) Update cluster centers mi and reassign, if necessary, entities to new clusters.
(4) Stop if convergence threshold is reached, e.g., no (or minimal) reassignments
occurred. Otherwise go back to (2). A key problem is de ning a dissimilarity
function for entities in D. Using kernels we may de ne such a function { as we
will show later, cf. Sect. 4.1.
        </p>
        <p>Problem. We address the problem of nding contextual data sources for
a given source. Contextual sources should: (a) Add new entities that refer to
the same real world object as given ones (entity complement). (b) Add entities
with same URI identi er, but have new properties in their description (schema
complement). Note, (a) and (b) are not disjoint, i.e., some entities may be new
and add additional properties.</p>
        <p>Example. In Fig. 1 Src. 3 contextualizes Src. 1 in terms of both, entity as
well as schema complement. That is, Src. 3 adds entity wbi:NY.GDP.MKTP.CN,
while providing new properties and a di erent GDP value. In fact, even Src. 2
contextualizes Src. 1, as entities in both sources refer to the real world object
\Germany" { one source captures the GDP, while the other describes the debt.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Entity-based Data Source Contextualisation</title>
      <p>Approach. The intuition behind our approach is simple: if data sources contain
similar entities, they are somehow related. In other words, we rely on (clusters
of ) entities to capture the \latent" semantics of data sources.</p>
      <p>
        More precisely, we start by extracting entities from given data sources. In a
second step, we apply clustering techniques, to mine for entity groups. Notice,
for the k-means clustering we employ a kernel function as similarity measure.
This way, we abstract from the actual entity representation, i.e., RDF graphs,
and use a high-dimensional space (feature space) to increase data
comparability/separability [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Last, we rely on entity clusters for relating data sources to
each other and compute a contextualization score based on these clusters. Note,
only the last step is done online { all other steps/computations are o ine.
      </p>
      <p>
        Discussion. We argue our approach to allow for some key advantages: (1) We
decouple representation of source content and source similarity, by relying on
entities for capturing the overall source semantics. (2) Our solution is highly
exible as it allows to \plug-in" application-speci c entity de nitions/extraction
strategies, entity similarity measures, and contextualization heuristics. That is,
we solely require an entity to be described as a subgraph, Ge, contained in its
data source D. Further, similarity measures may be based on any valid kernel
function. Last, various heuristics proposed in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] could be adapted for our
approach. (3) Exploiting clusters of entities allows for a scalable and maintainable
approach, as we will outline in the following.
4.1
      </p>
      <p>Related Entities
In order to compare data sources, we rst compare their entities with each other.
That is, we extract entities, measure similarity between them and nally cluster
them (Fig. 2). All of these procedures are o ine.</p>
      <p>
        Entity Similarity. We start by extracting entities from each source Di 2 D:
First, for scalability reasons, we go over all subjects in RDF graphs in Di and
collect an entity sample, with every entity e having the same probability of
being selected. Then, we crawl the concise bound description [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for each entity
e in the sample. For cleaning Ge, we apply standard data cleansing strategies
to x, e.g., missing or wrong data types. Having extracted entities, we de ne a
similarity measure relating pairs of entities. For this, we use two kinds of kernels:
(1) kernels for structural similarities s and (2) those for literal similarities l.
      </p>
      <p>
        With regard to the former, we measure structural \overlaps" between entity
descriptions, Ge0 and Ge00, using graph intersections [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]:
De nition 4 (Graph Intersection G0 \ G00). Let the intersection between two
graphs, denoted as G = G0 \ G00, be given by V = V0 \ V00 and E = E 0 \ E 00.
tec0001
Src. 1
      </p>
      <p>gov_q_ggdebt
Src. 2</p>
      <p>NY.GDP.MKTP.CN
Src. 3</p>
      <p>Cluster C1
Size = 2
Cluster C2
Size = 1</p>
      <p>
        We aim at connected structures in Ge0 \ Ge00. Thus, we de ne a path [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]:
De nition 5 (Path). Let a path in a graph G be de ned as a sequence of
vertices and triples v1; hv1; p1; v2i; v2; : : : ; vn, with hvi; pi; vi+1i 2 E , having no
cycles. The path length is given by the number of contained triples. The set of all
paths, up to length l, in G is denoted as pathl(G).
      </p>
      <sec id="sec-4-1">
        <title>The corresponding path kernel is [16]:</title>
        <p>De nition 6 (Path Kernel s). A path kernel is ls; (G1; G2) = Pli=1
p 2 pathsi(G1 \ G2)g j, with &gt; 0 as discount factor for path length.
i j fp j</p>
        <p>
          Note, [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] introduced further kernels, however, we found path kernels to be
simple and perform well in our experiments, cf. Sect. 6.
        </p>
        <p>Example. Extracted entities from sources in Fig. 1 are given in Fig. 2-a. In
Fig. 2-b, we compare the structure of tec0001 (short: e1) and gov q ggdebt
(short: e2). For this, we compute an intersection: Ge1 \ Ge2. This yields a set of
4 paths, each with length 0. The unnormalized kernel value is 0 4.</p>
        <p>
          For literal similarities, one can use di erent kernels l on, e.g., strings or
numbers [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. For space reasons, we restrict presentation to the string subsequence
kernel, ls [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. A numerical kernel, ln, is outlined in our extended report [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
De nition 7 (String Subsequence Kernel ls). Let denote a vocabulary
for strings, with each string s as nite sequence of characters in . Let s[i :
j] denote a substring si; : : : ; sj of s. Further, let u be a subsequence of s, if
indices i = (i1; : : : ; ijuj) exist with 1 i1 ijuj jsj such that u = s[i].
The length l(i) of subsequence u is ijuj i1 + 1. Then, a kernel function l is
de ned as sum over all common, weighted subsequences for strings s, t: l (s; t) =
Pu Pi:u=s[i] Pj:u=t[j] l(i) l(j), with as decay factor.
        </p>
        <p>Example. For instance, strings \MI" and \MIO EUR" share a common
subsequence \MI" with i = (1; 2). Thus, the unnormalized kernel is 2 + 2.</p>
        <p>
          As literal kernels, l, are only de ned for two literals, we sample over every
possible literal pair (with the same data type) for two given entities and aggregate
the needed kernels for each pair. Finally, we aggregate structure kernel, s and
literal kernels, l, resulting in one single kernel [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]:
        </p>
        <p>
          Entity Clustering. Using this dissimilarity measure, we may learn clusters
of entities. Notice, our algorithm does not separate the input data (graphs Ge),
but instead its representation in the feature space. Based on [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], cluster center
mi in the feature space is: mi = jC1ij P 1('(e); Ci)'(e). Distance between a
projected entity '(e) and mi is given by [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]:
dis2('(e); mi) = k'(e)
        </p>
        <p>2
f (e; Ci) :=</p>
        <p>mik2 = (e; e) + f (e; Ci) + g(Ci), with</p>
        <p>X 1('(ej ); Ci) (e; ej )
jCij j
g(Ci) :=</p>
        <p>1
jCij2 j</p>
        <p>X X 1('(ej ); Ci)1('(el); Ci) (ej ; el)</p>
        <p>l
Now, k-means can be applied as introduced in Sect. 3.</p>
        <p>Example. In Fig. 2-c, we found two entity clusters. Here, structural similarity
was higher for entities e1 vs. e2 than for e1 vs. e3. However, numerical
similarity between e1 vs. e3 was stronger: \2496200.0" was closer to \2500090.5" as
\1786882". As we weighted literal similarity to be more important than structural
similarity, this leads to e1 and e3 forming one cluster. In fact, such a clustering
is intuitive, as e1 and e3 is about GDP, while e2 is concerned with debt.
4.2</p>
        <p>
          Related Data Sources
Given a data source D0, we score its contextualization w.r.t. another source D00,
using entities contained in D0 and D00. Note, in contrast to previous work [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ],
we do not rely on any kind of \external" information, such as top-level schema.
Instead, we solely exploit semantics as captured by entities.
        </p>
        <p>
          Contextualisation Score. Similar to [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], we compute two scores, ec(D00 j
D0) and sc(D00 j D0), for a data source D00, given another source D0. The former
is an indicator for the entity complement of D00 w.r.t. D0. That is, how many
new, similar entities does D00 contribute to given entities in D0. The latter score
judges how many new \dimensions" are added by D00, to those present in D0
(schema complement). Both scores are aggregated to a contextualization score
for data source D00 given D0.
        </p>
        <p>Let us rst de ne an entity complement score ec : D D 7! [0; 1]. We may
measure ec simply by counting the overlapping clusters between both sources:</p>
        <p>Example. Assuming Src. 1 is given, cs(D2 j D1) = cs(D3 j D1) = 1=2, because
ec(D3 j D1) = 1 and sc(D3 j D1) = 0 (the other way around for D2 j D1). These
scores are meaningful, because D2 adds additional properties (debt), while D3
complements D1 with entities (GDP). See also Fig. 2-c.</p>
        <p>Runtime Behavior and Scalability. Regarding online performance, i.e.,
computation of contextualization score cs, given the o ine learned clusters, we
aimed at simple and lightweight heuristics. For ec only an assignment of data
sources to clusters (function cluster()) and cluster size jCj is needed. Further,
measure sc only requires an additional mapping of clusters to \contained"
properties (function props()). All necessary statistics are easily kept in memory.</p>
        <p>
          With regard to the o ine clustering behavior, we expect our approach to
perform well, as existing work on kernel k-means clustering showed such approaches
to scale to large data sets [
          <xref ref-type="bibr" rid="ref1 ref22 ref5">1, 5, 22</xref>
          ].
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Searching the Web of Data via Source Contextualization</title>
      <p>
        Based on our real-world use case (Sect. 2), we show how a source
contextualization engine may be integrated in a real-world query processing system for the Web
of data. Notice, an extended system description is included in our report [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>Overview. Towards an active involvement of users in the source selection
process, we implemented a data portal, which o ers a data source space
exploration as well as a distributed query processing service.</p>
      <p>Using the former, users may explore the space of sources, i.e., search and
discover data sources of interest. Here, the contextualization engine fosters
discovery of relevant sources during exploration. The query processing service, on
the other hand, allows queries to be federated over multiple sources.</p>
      <p>Data Source</p>
      <p>Search</p>
      <p>Data Source</p>
      <p>Visualization</p>
      <p>Data Source</p>
      <p>Contextualization Engine
Data
Access</p>
      <p>Data Source
Meta-Data
Meta-Data Updates</p>
      <p>by Providers
Worldbank</p>
      <p>Provider
Eurostat
Provider</p>
      <p>Entity
Clusters</p>
      <p>Inspect Source
Contributing to
Current Result</p>
      <p>Data Source
gov_q_ggdebt
Offline Entity Extraction
and Clustering</p>
      <p>Data Sources
Eurostat</p>
      <p>Federation Layer</p>
      <p>Data Source</p>
      <p>tec00001</p>
      <p>Data Loader
Worldbank</p>
      <p>Query &amp;
Visualize
Results</p>
      <p>Processing
SPARQL Queries
against the</p>
      <p>Federation
Data Source</p>
      <p>NY.GDP.MKTP.CN
Loading and Populating
of SPARQL-Endpoints</p>
      <p>FluidOps Data Portal
Data Source Exploration</p>
      <p>Query Processing
Select/Remove Source
for Query Federation</p>
      <p>SPARQL
Query</p>
      <p>Result
Visualization</p>
      <p>Interaction between both services is tight and user-driven. In particular,
sources discovered during source exploration may be used for answering queries.
On the other hand, sources employed for result computation may be inspected
and other relevant sources may be found via contextualization.</p>
      <p>
        The data portal is based on the Information Workbench [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and a running
prototype is available.5 Following our use-case (Sect. 2), we populated the system
with statistical data sources from Eurostat and Worldbank. This population
involved an extraction of meta-data, which is used to load the sources locally as
well as give users insights into the source. Every data source is stored in a triple
store { accessible via a SPARQL endpoint. See Fig. 3 for an overview.
      </p>
      <p>Source Exploration and Selection. A typical search process starts with
looking for \the right" sources. That is, a user begins with exploration of the
data source space. For instance, she may issue a keyword query \gross domestic
product", yielding sources with matching words in their meta-data. If this query
does not lead to sources suitable for her information need, a faceted search
interface or a tag-cloud may be used. For instance, she re nes her sources via entity
\Germany" in a faceted search, Fig.4-a. Once the user discovered a source of
interest, its structure as well as entity information is shown. For example, a source
description for GDP (current US$) is given in Fig.4-b/c. Note, entities used here
have been extracted by our approach and are visualized by means of a map. Using</p>
      <sec id="sec-5-1">
        <title>5 http://data.fluidops.net/</title>
        <p>(a)
(b)
(c)
(d)
these rich source descriptions, a user can get to know the data and data sources
before issuing queries. Further, for every source a ranked list of contextualization
sources is given. For GDP (current US$), e.g., source GDP at Market Prices
is recommended, Fig.4-d. This way, the user is guided from one source of interest
to another. At any point, she may select a particular source for querying.
Eventually, she not only knows her relevant sources, but has also gained rst insights
into data schema and entities.</p>
        <p>Processing Queries over Selected Sources. Say, a user has chosen GDP
(current US$) as well as its contextualization source GDP at Market Prices
(Fig. 4-d). Due to her previous exploration, she knows that the former provides
the German GDP from 2000 - 2010, while the second source features GDP from
years 2011 and 2012. Thus, she may issue a SPARQL query over both sources
to visualize the GDP over the years 2000 - 2012, cf. Fig. 5.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Evaluation</title>
      <p>We now discuss evaluation results to analyze the e ectiveness of our approach.
We conducted two experiments: (1) E ectiveness w.r.t. a gold standard, thereby
measuring the accuracy of the contextualization score. (2) E ectiveness w.r.t.
data source search result augmentation. In other words, we ask: How useful are
top-ranked contextualization sources in terms of additional information?</p>
      <p>Participants. Experiment (1) and (2) were performed by two di erent
groups, each comprising 14 users. Most participants had an IT background
and/or were CS students. Three users came from the nance sector. The users
vary in age, 22 to 37 years and gender. All experiments were unsupervised.</p>
      <p>Data Sources and Clustering. Following our use case in Sect. 2, we
employed Linked Data sources from Eurostat6 and Worldbank7. Both provide a
6 http://eurostat.linked-statistics.org/
7 http://worldbank.270a.info
g
? obs1 a qb : Observation ; ? obs2 a qb : Observation ;
wb p r o p e r t y : i n d i c a t o r qb : d a t a s e t es data : t e c 0 0 0 0 1 ;
wbi c i :NY.GDP.MKTP.CN; es p r o p e r t y : geo es d i c : geo#DE;
sdmx dimension : r e f A r e a sdmx dimension : t i m e P e r i o d ? year ;
wbi cc :DE; sdmx measure : obsValue ? gdp .
sdmx dimension : r e f P e r i o d ? year ; FILTER( ? year &gt; "2010 01 01" ^^ xs : date )
sdmx measure : obsValue ? gdp . g</p>
      <p>UNION
f
g
large set of data sources comprising statistical information: Eurostat holds 5K
sources (8000M triples), while Worldbank has 76K sources (167M triples). Note,
for Eurostat we simply used instances of void:Dataset as sources and for
Worldbank we considered available dumps as sources. Further, sources varied strongly
in their size and # entities. While small sources contained 10 entities, large
sources featured 1K entities. For extracting/sampling of entities, we restricted
attention to instances of qb:Observation. We employed the k-means algorithm
with k = 30K (chosen based on experiments with di erent values for k) and
initiated the cluster centers via random entities.</p>
      <p>
        Queries. Based on evaluation queries in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], domain experts constructed 15
queries. A query listing is depicted in Table 1. This query load was designed to
be \broad" w.r.t. the following dimensions: (1) We cover a wide range of topics,
which may be answered by Eurostat and Worldbank. (2) \Best" results for each
query are achieved by using multiple sources from both datasets. (3) # Sources
and # entities relevant for a particular query vary.
      </p>
      <p>
        Systems. We implemented entity-based source contextualization (EC) as
described in Sect. 4. In particular, we made use of three kernels for capturing entity
similarity: a path kernel, a substring kernel, and a numerical kernel [
        <xref ref-type="bibr" rid="ref16 ref19">16, 19</xref>
        ]. As
baselines we used two approaches: a schema-based contextualization SC [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and
a keyword-based contextualization KC. SC maps entities in data sources to
toplevel schemata and nds complementary sources based on schema as well as
label similarities of mapped entities. More precisely, the baseline is split in two
approaches: SCec and SCsc. SCec aims at entity complements, i.e., data sources
with similar schema, but di erent entities. SCsc targets data sources having
complementary schema, but holding the same entities. Last, the KC approach treats
every data source as a bag-of-words and judges the relevance of a source w.r.t.
a query by the number and quality of its matching keywords.
      </p>
      <p>E ectiveness of Contextualisation Score
Gold Standard. We calculated a simple gold standard that ranks pairs of
sources based on their contextualization. That is, for each query in Table 1,
we randomly selected one \given" source and 5 \additional" sources { all of
which contained the query keywords. Then, 14 users were presented that given
source and had to rank how well each of the 5 additional sources contextualizes
it (scale on 0 - 5, where higher is better). To judge the sources' contents, users
were given a schema description and a short source extract. We aggregated the
user rankings, thereby obtaining a gold standard ranking.</p>
      <p>
        Metric. We applied the footrule distance [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] as rank-distance indicator,
which measures the accuracy of the evaluation systems vs. gold standard. Footrule
distance is: k1 Pi=1;:::;k jranka(i) rank(i)j, with ranka(i) and rank(i) as
approximated and gold standard rank for the \additional" source Di.
      </p>
      <p>Results. Fig. 6-a/d give an overview over results of EC and SC. Note, we
excluded the KC baseline, because it was too simplistic. That is, reducing sources
to bags-of-words, one may only intersect these bags for two given sources. This
yielded, however, no meaningful contextualization scores/ranking.</p>
      <p>We noticed EC to lead to a good and stable performance over all queries
(Fig. 6-a) as well as data source sizes (Fig. 6-d). Overall, EC could outperform
SCsc by 5:7% and SCec by 29%. Further, as shown in Fig. 6-d, while we observed
a decrease in rank distance for all systems in source and entity sample size, EC
yielded the best results. We explain these results with our ne-grained source
semantics captured by entity clusters. Note, most of our employed
contextualization heuristics are very similar to those from SC. However, we observed SC
performance to strongly vary with the quality of its schema mappings. Given an
accurate classi cation of entities contained in a particular source, SC was able
to e ectively relate that source with others. However, if entities were mapped to
\wrong" concepts, it greatly a ected computed scores. In contrast, our approach
relied on clusters learned from instance data, thereby achieving a \more reliable"
mapping from sources to their semantics (clusters).</p>
      <p>On the other hand, we observed EC to result in equal or, for 2 outlier queries
with many \small" sources, worse performance than SCsc (cf. Fig. 6-d). In
particular, given query Q2, SCsc could achieve a better ranking distance by 20%.
We explain such problematic queries with our simplistic entity sampling. Given</p>
      <p>(b)
3.17
small sources, only few entities were included in a sample, which unfortunately
pointed to \misleading" clusters. Note, we currently only use a uniform random
sample for selecting entities from sources. We expect better results with more
re ned techniques for discovering important entities for a particular source.</p>
      <p>Last, we observed SCec to lead to much less accurate rankings as SCsc and
EC. This is due to exact matching of entity labels: SCec did not capture common
substrings (as EC did), but solely relied on a boolean similarity matching.</p>
      <p>
        E ectiveness of Augmented Data Source Results
Metric. Following [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we employ a reordering of data source search results as
metric: (1) We obtain top 100 data sources via the KC baseline for every query.
Here, the order is determined solely by number and quality of keyword hits in
the sources' bags-of-words. (2) Via SC and EC we reorder the rst 10 results:
      </p>
      <p>For each data source Di in the top-10 results, we search the top-100 sources
for a Dj that contextualizes Di best. If we nd a Dj with contextualization score
higher than a threshold, we move Dj after Di in the top-10 results.</p>
      <p>This reordering procedure yields three di erent top-10 source rankings: one
via KC and two (reordered ones) from SC/EC. For every top-10 ranking, users
provided a relevance feedback for each source w.r.t. a given query, using a scale
0 - 5 (higher means \more relevant"). For this, users had a schema description
and short source extract for every source. Last, we aggregated these relevancy
judgments for each top-10 ranking and query.</p>
      <p>Results. User relevance scores are depicted in Fig. 6-b/e. Overall, EC yielded
an improved ranking, i.e., it ranked more relevant sources w.r.t. a given query.
More precisely, EC outperforms SCsc with 2:5%, SCec with 2:8% and KC with
6:2%, cf. Fig. 6-b. Furthermore, our EC approach led to stable results over varying
source sizes (Fig. 6-e). However, for 2 queries having large sources in their top-100
results, we observed misleading a clustering of sampled entities. Such associated
clusters, in turn, led to a bad contextualization, see Fig. 6-e.</p>
      <p>
        Notice that all result reorderings, either by EC or by SC, yielded an
improvement (up to 6:2%) over the plain KC baseline, cf. Fig. 6-b/e. These ndings
con rm results reported in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and clearly show the potential of data source
contextualization techniques. In other words, such observations demonstrate the
usefulness of contextualization { participants wanted to have
additional/contextualized sources, as provided by EC or SC.
      </p>
      <p>In Fig. 6-c we show the # sources returned for contextualization of the
\original" top-10 sources (obtained via KC). EC returned 19% less contextualization
sources than SCsc and 28% less than SCec. Unfortunately, as contextualization is
a \fuzzy" criteria, we could not determine the total number of relevant sources
to be discovered in the top-100. Thus, no recall factor can be computed.
However, combined with relevance scores in Fig. 6-b, # sources gives a precision-like
indicator. That is, we can compute the average relevance gain per
contextualization source: 1:13 EC, 0:89 SCsc, and 0:77 SCsc, see Fig. 6-b/c. Again, we explain
the good \precision" of EC with the ne-grained clustering, providing better and
more accurate data source semantics.</p>
      <p>Last, it is important to note that, while Exp. 1 (Sect. 6.1) and Exp. 2 (Sect. 6.2)
di er greatly in terms of their setting, our observations were still similar: EC
outperformed both baselines due to its ne-grained entity clusters.</p>
    </sec>
    <sec id="sec-7">
      <title>7 Related Work</title>
      <p>
        Closest to our approach is work on contextual Web tables [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] focuses
on at entities in tables, i.e., entities adhere to a simple and xed relational
structure. In contrast, we consider entities to be subgraphs contained in data
sources. Further, we do not require any kind of \external" information. Most
notably, we do not use top-level schemata.
      </p>
      <p>
        Another line of work is concerned with query processing over distributed RDF
data, e.g., [
        <xref ref-type="bibr" rid="ref10 ref13 ref18 ref6 ref9">6, 9, 10, 13, 18</xref>
        ]. During source selection, these approaches frequently
exploit indexes, link-traversal, or source meta-data, for mapping queries/query
fragments to sources. Our approach is complementary, as it enables systems
to involve users during source selection. We outlined such an extension of the
traditional search process as well as its bene ts throughout the paper.
      </p>
      <p>
        Last, data integration for Web (data) search has received much attention [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
In fact, some works aim at recommendations for sources to be integrated, e.g., [
        <xref ref-type="bibr" rid="ref14 ref17 ref3">3,
14, 17</xref>
        ]. Here, the goal is to give recommendations to identify (integrate) identical
entities and schema elements across di erent data sources.
      </p>
      <p>In contrast, we target a \fuzzy" form of integration, i.e., we do not give exact
mappings of entities or schema elements, but merely measure whether or not
sources contain entities that might be \somehow" related. In other words, our
contextualization score indicates whether sources might refer to similar entities
and may provide contextual information. Furthermore, our approach does not
require data sources to adhere to a known schema { instead, we exploit data
mining strategies on instance data (entities).</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>We presented a novel approach for Web data source contextualization. For this,
we adapted well-known techniques from the eld of data mining. More precisely,
we provide a framework for source contextualization, to be instantiated in an
application-speci c manner. By means of a real-world use-case and prototype,
we show how source contextualization allows for user involvement during source
selection. We empirically validated our approach via two user studies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Chitta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jin</surname>
          </string-name>
          , T. C.
          <article-title>Havens, and</article-title>
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Jain</surname>
          </string-name>
          .
          <article-title>Approximate kernel k-means: solution to large scale kernel clustering</article-title>
          .
          <source>In SIGKDD</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>A. Das Sarma</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Xin</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Finding related tables</article-title>
          .
          <source>In SIGMOD</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>3. H. R. de Oliveira</surname>
            ,
            <given-names>A. T.</given-names>
          </string-name>
          <string-name>
            <surname>Tavares</surname>
            , and
            <given-names>B. F.</given-names>
          </string-name>
          <string-name>
            <surname>Loscio</surname>
          </string-name>
          .
          <article-title>Feedback-based Data Set Recommendation for Building Linked Data Applications</article-title>
          . In I-SEMANTICS,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Halevy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z. G.</given-names>
            <surname>Ives</surname>
          </string-name>
          .
          <article-title>Principles of Data Integration</article-title>
          . Morgan Kaufmann,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Fausser</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Schwenker</surname>
          </string-name>
          .
          <article-title>Clustering large datasets with kernel methods</article-title>
          .
          <source>In ICPR</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>O.</given-names>
            <surname>Go</surname>
          </string-name>
          <article-title>rlitz and S. Staab. SPLENDID: SPARQL Endpoint Federation Exploiting VOID Descriptions</article-title>
          .
          <source>In COLD Workshop</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Grimnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Preece</surname>
          </string-name>
          .
          <article-title>Instance based clustering of semantic web resources</article-title>
          .
          <source>In ESWC</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>P.</given-names>
            <surname>Haase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Schwarte</surname>
          </string-name>
          .
          <article-title>The Information Workbench as a SelfService Platform for Linked Data Applications</article-title>
          .
          <source>In COLD Workshop</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karnstedt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sattler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          .
          <article-title>Data summaries for on-demand queries over linked data</article-title>
          .
          <source>In WWW</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>O.</given-names>
            <surname>Hartig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Freytag</surname>
          </string-name>
          .
          <article-title>Executing SPARQL Queries over the Web of Linked Data</article-title>
          .
          <source>In ISWC</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>A. K. Jain</surname>
            ,
            <given-names>M. N.</given-names>
          </string-name>
          <string-name>
            <surname>Murty</surname>
            , and
            <given-names>P. J.</given-names>
          </string-name>
          <string-name>
            <surname>Flynn</surname>
          </string-name>
          .
          <article-title>Data clustering: a review</article-title>
          .
          <source>ACM Computing Surveys</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>M.</given-names>
            <surname>Kendall</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Gibbons</surname>
          </string-name>
          .
          <source>Rank Correlation Methods</source>
          .
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. G. Ladwig and
          <string-name>
            <given-names>T.</given-names>
            <surname>Tran</surname>
          </string-name>
          .
          <article-title>Linked Data Query Processing Strategies</article-title>
          . In ISWC,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>L. A. P. P.</given-names>
            <surname>Leme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Casanova</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Dietze</surname>
          </string-name>
          .
          <article-title>Identifying Candidate Datasets for Data Interlinking</article-title>
          .
          <source>In ICWE</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. H.
          <string-name>
            <surname>Lodhi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Shawe-Taylor</surname>
            , N. Cristianini, and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Watkins</surname>
          </string-name>
          .
          <article-title>Text classi cation using string kernels</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Losch et al.
          <article-title>Graph kernels for RDF data</article-title>
          .
          <source>In ESWC</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>A.</given-names>
            <surname>Nikolov</surname>
          </string-name>
          and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>d'Aquin. Identifying Relevant Sources for Data Linking using a Semantic Web Index</article-title>
          .
          <source>In LDOW Workshop</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>A.</given-names>
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schwarte</surname>
          </string-name>
          , and
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Hutter. FedSearch: E ciently Combining Structured Queries and Full-Text Search in a SPARQL Federation</article-title>
          . In ISWC.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. J.
          <string-name>
            <surname>Shawe-Taylor</surname>
            and
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Cristianini</surname>
          </string-name>
          .
          <article-title>Kernel Methods for Pattern Analysis</article-title>
          .
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>A.</given-names>
            <surname>Wagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Haase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rettinger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Lamm</surname>
          </string-name>
          .
          <article-title>Discovering related data sources in data-portals</article-title>
          . In First International Workshop on Semantic Statistics,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>A.</given-names>
            <surname>Wagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Haase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rettinger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Lamm</surname>
          </string-name>
          .
          <article-title>Entity-based Data Source Contextualization for Searching the Web of Data</article-title>
          . Techreport,
          <year>2013</year>
          . http://www. aifb.kit.edu/web/Techreport3043.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Rudnicky</surname>
          </string-name>
          .
          <article-title>A large scale clustering scheme for kernel K-Means</article-title>
          .
          <source>In Pattern Recognition</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>