<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards Dataset Dynamics: Change Frequency of Linked Open Data Sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>J u¨rgen Umbrich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Hausenblas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aidan Hogan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel Polleres</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Decker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Enterprise Research Institute (DERI) National University of Ireland</institution>
          ,
          <addr-line>Galway IDA Business Park, Lower Dangan</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <volume>27</volume>
      <issue>2010</issue>
      <abstract>
        <p>Datasets in the LOD cloud are far from being static in their nature and how they are exposed. As resources are added and new links are set, applications consuming the data should be able to deal with these changes. In this paper we investigate how LOD datasets change and what sensible measures there are to accommodate dataset dynamics. We compare our findings with traditional, document-centric studies concerning the “freshness” of the document collections and propose metrics for LOD datasets.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The Linked Open Data (LOD) movement has gained
remarkable momentum over the past years. At the time of
writing, well over one hundred datasets – including UK
governmental data, the New York Times dataset, and
LinkedGeoData – have been published, providing several billion
RDF triples interlinked by hundreds of millions of RDF
links. Some datasets, such as DBpedia, have been available
from the very beginning of the LOD movement and
regularly undergo changes on both the instance level and the
schema level. New resources are added and old resources
are removed; new links are set to other datasets, and old
links are removed as the target has vanished. We should
hence assume that datasets in the LOD cloud are dynamic
in their very nature. Dataset dynamics is a term we
recently coined [1], essentially addressing content and
interlinking changes in Linked Data sources.</p>
      <p>
        Our main contributions herein are: (i) define dataset
dynamics characteristics and how to measure them, and (ii)
compare the dataset dynamics of the LOD cloud to the
traditional Web (Web of HTML Documents). The motivating
use-case for our study of dataset dynamics is to gain
insights into – and hopefully improve – concurrent work on an
efficient system for performing live queries over the Linked
Open Data Web [
        <xref ref-type="bibr" rid="ref12">13</xref>
        ]. However, aside from this use-case
having knowledge about dataset dynamics is essential for a
number of tasks:
• web crawling and caching [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ];
• distributed query optimisation [
        <xref ref-type="bibr" rid="ref12">13</xref>
        ];
• maintaining link integrity [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ];
• servicing of continuous queries [
        <xref ref-type="bibr" rid="ref21">22</xref>
        ];
• replication and synchronisation [
        <xref ref-type="bibr" rid="ref23">24</xref>
        ].
      </p>
      <p>We begin in Section 2 by reviewing existing work, and
continue in Section 3 by discussing and contrasting document
vs. entity centric perspectives concerning dynamics.
Thereafter, in Section 4 we present the background of our
analysis, in Section 5 we describe our methodology for analysing
dataset dynamics, and in Section 6 we discuss the results of
our analysis. Finally, in Section 7, we conclude and render
future work.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        As motivated above, the study of changes in documents
and data sets is very relevant for a broad range of application
domains. Earlier work discussed analysis of the dynamics of
the Web circa. 2008, leveraging their findings for
optimisation of re-indexing techniques [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ]. The work of Cho et.
al. provides a comprehensive study regarding the change
frequency of Web documents: earlier work focussed on how
to integrate the knowledge for an incremental crawler [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ];
further work provided a detailed discussion for estimators
of the frequency of changes given incomplete history [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ].
Other research has focused on, for example, investigating
the dynamics of Wikipedia articles [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ] and the evolution of
database schema over time [
        <xref ref-type="bibr" rid="ref20">21</xref>
        ].
      </p>
      <p>
        With respect to the Semantic Web, some research
regarding dynamics has been conducted with respect to analysing
the evolution of ontologies in the life science community [
        <xref ref-type="bibr" rid="ref14">15</xref>
        ].
In [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ] the authors reported on their work concerning
DSNotify, a system for detecting and fixing broken links in LOD
datasets.
      </p>
      <p>However – and to the best of our knowledge – we are
not aware of any published studies more generally regarding
the change frequency of resources on the Linked Open Data
Web, and thus deem the work herein to be novel.
3.</p>
    </sec>
    <sec id="sec-3">
      <title>DOCUMENTS VS. ENTITIES: DIFFER</title>
    </sec>
    <sec id="sec-4">
      <title>ENT PERSPECTIVES ON LINKED DATA</title>
      <p>There are various aspects of dataset dynamics which must
be considered in order to achieve a comprehensive overview
of how Linked Open Data changes and evolves on the Web.
Firstly, the change frequency of data on the Web can vary
significantly across datasets, from rather static sources –
such as archives – to high-frequently changing sources – for
example in the micro-blogging domain. Also, the change
volume can range from small-scale updates – in our case,
updates involving a low number of triples – to bulk updates,
which potentially affect many resources. One must also pay
attention to the perspective one takes on resources: that
is, whether we are interested in local changes of particular
datasets, or are interested in global changes with respect to
what is said about a URI in all accessible linked datasets.</p>
      <p>
        Before we continue, however, we must first provide some
preliminaries. Firstly, our notion of a ‘document’ refers to
an atomic Web ‘container’ in which Linked Data is typically
exposed: these include RDF/XML, (X)HTML+RDFa
documents, etc. Secondly, we often refer to an ‘entity’ by which
we intuitively mean anything identified by a URI in Linked
Data, including classes, properties, and the “real-world
artefacts” described.1 Following from both, we can now
distinguish the following perspectives in dataset dynamics:
1. A document-centric perspective, which focuses on
datasets and is motivated by the “traditional” Web as well
as the REST community [
        <xref ref-type="bibr" rid="ref1 ref11">12, 2</xref>
        ]
2. An entity-centric perspective, which focuses on entities
as described in the Linked Open Data Web [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] – we
further separate the entity-centric perspective into:
(a) An entity-per-document perspective which takes
into account occurrences of an entity with respect
to a specific document
(b) A global entity perspective which takes into
account all appearances of an entity across the Web
In particular, the entity-centric perspectives are more
LODspecific than the document-centric perspective prevalent in
more traditional views on dataset dynamics. Many
applications operating on the LOD cloud assume an entity-centric
view where entities become the unit of knowledge and data
on such entities are aggregated from multiple documents.
Also, LOD documents may be dynamically served by an
entity-centric index (e.g., a SPARQL endpoint), whereby a
change in one entity may entail changes in many documents.
Thus, we believe the distinction between the document- and
entity-centric perspectives to be important for our purposes
herein.
      </p>
      <p>In fact, the global entity perspective may be infeasible to
monitor as arbitrary new sources can publish data about any
entities. For this reason – and despite formally discussing
2b herein – note that in the present work we will focus on
the analysis of 1 and 2a, and leave approximative techniques
for analysis of 2b as part of our future research (discussed
in Section 7).</p>
      <p>
        Despite the two distinct perspectives, both are somehow
related: there is naturally a relation between entities and
their appearances in different containers. Along these lines,
Figure 1 depicts a typical distribution of entities per
document in the LOD cloud. As we have already shown
elsewhere [
        <xref ref-type="bibr" rid="ref16">17</xref>
        ], this distribution follows a power law.
      </p>
      <p>
        In order to formalise what we mean by these different
perspectives, let R = {r1, ..., rn} be the set of all resources
as of the Architecture of the World Wide Web [
        <xref ref-type="bibr" rid="ref18">19</xref>
        ]: that
is, HTTP entities and documents. Further, we define D =
{d1, ..., dn}, D ⊂ R as the set of all documents (i.e.,
dereferenceable entities that point to RDF data) and E = {e1, ..., en},
E ⊂ R as the set of all entities. A document di can mention
1Note that in this paper, we currently overlook entities
‘identified’ by blank-nodes; concretely, blank-node entities
do not have consistent naming which has adverse
consequences on the analysis presented in Section 5.3
10000
s
t
n
e
um1000
c
o
d
f
roe 100
b
m
u
n 10
1
1
10
      </p>
      <p>100 1000
number of entities per document
10000
various entities; thus, we denote the set of entities mentioned
in document d as E(d) ⊆ E and likewise the set of all
documents mentioning e as D(e) ⊆ D. Further, let ver(d, t)
be the state of document d at time-point t – i.e., the RDF
graph served by d at time t. It is clear, that different use
cases require specific state functions ver(d, t) and equality
measures; e.g. a state function could be the hash value of the
RDF graph, a set of RDF statements or the set of inferable
new statements.</p>
      <p>Then, the document change function of document d from
time t to t0 (where t &lt; t0) is defined as follows:</p>
      <p>Definition 1. Document Change Function</p>
      <p>Cd(t, t0) =
 0
1
if ver(d, t) = ver(d, t0)
otherwise
Likewise, we define the entity-per-document change function
as follows:</p>
      <p>Definition 2. Entity-per-document Change Function
Cde(t, t0) =
 0
1
if ver(d, t) ∩ e = ver(d, t0) ∩ e
otherwise
where by G∩u we denote all triples in graph G mentioning u.
Finally, the entity change function can be defined as follows:
Definition 3. Entity Change Function</p>
      <p>Ce(t, t0) = max∀d∈D(e)(Cd(t, t0))</p>
      <p>Please note that we pursue a purely ‘syntactic’ notion of
change, and do not consider more advanced notions
relating to ‘semantic’ change: for example, we would consider
a change in a datatype literal if the syntax of that literal
changes even though the semantic interpretation does not
– this change would then propagate to the respective
entity/document despite no real change on the semantic level.
Further, we do not consider any forms of reasoning in the
changes – e.g., we do not propagate changes in a class
definition as changes to it’s member entities. We leave further
discussion and related analysis of ’semantic vs. syntactic
change’ for future work.</p>
      <p>
        It may also be interesting to consider more closely the
relationship between documents and the entities they contain,
examining separately the change function of entities which
are considered ‘local’ with respect to the document they
appear in. To this end, we introduce the term local entity,
meaning an entity in a document whose pay-level domain
(PLD) is the same as the document’s PLD: here, a PLD is
defined as any domain that requires payment at a
[top-leveldomain] (TLD) or country-code TLD registrar [
        <xref ref-type="bibr" rid="ref19">20</xref>
        ]. Taking
an example, let P LD(uri) be the PLD extraction function;
then:
      </p>
      <p>P LD(http : //www.deri.ie/) = deri.ie
We can now define a local entity as follows:</p>
      <p>Definition 4. Local Entity
We define the set of local entities Elocal(d) of document d as</p>
      <p>Elocal(d) = {e ∈ E(d) | P LD(e) = P LD(d)}</p>
      <p>
        Definition 4 is closely related to a similar notion defined
in [
        <xref ref-type="bibr" rid="ref6">7</xref>
        ], which defines locality based on the correspondence of
hostnames. Note that, according to this definition, an entity
may be local to several documents, which may not always
be desirable. Alternatively, one could focus on the
authoritative relationship between entities and documents whereby
the document an entity redirects to is the authoritative
document for that entity [
        <xref ref-type="bibr" rid="ref17">18</xref>
        ]. In this paper, we currently only
consider the locality relationship between documents and
entities and plan to investigate stronger notions such as
authoritativeness in future work.
      </p>
    </sec>
    <sec id="sec-5">
      <title>CHANGE DETECTION MECHANISM</title>
      <p>
        So far, we have focused on identifying and formalising
different notions of change – particularly change functions
– as a foundational aspect of dataset dynamics. We now
discuss how such changes can be detected; one can group
change detection mechanisms as follows:
• HTTP-metadata monitoring: analysis of HTTP
response headers – including datestamp and ETag [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ] –
to detect whether something has changed;
• content monitoring: fetching the entire content and
determining locally what has changed;
• notification: active notification by a data source that
something has changed (ideally what has changed) [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ].
      </p>
      <p>
        The Table 1 summarises aspects of the the
aforementioned change detection mechanisms. The aspects –
motivated by [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ] – are as follows: (i) availability, meaning
if the respective solution is available out-of-the-box in
currently deployed systems on the Web; (ii) reliability, referring
to the ability to correctly capture all changes; (iii) costs,
referring to the resources needed for the approach (in terms of
band-width, storage, etc.); and (iv) scalability with respect
to the number of involved data publishers (in terms of
infrastructure) and consumers (concerning, for example number
of concurrent “subscribers” in a notification system).
Further, we have included two Linked Data specific aspects in
Table 1: (v) support for document-centric change detection,
and (vi) support for entity-centric change detection.
      </p>
      <p>
        Both content and HTTP metadata monitoring
mechanisms are well studied and discussion about those is available
elsewhere (cf. [
        <xref ref-type="bibr" rid="ref10 ref9">10, 11</xref>
        ]). The characteristics of Web-scale
notification mechanisms – especially concerning reliability,
costs, and scalability are subject to research at time of
writing. However, there are some remarkable implementation
availability
reliability
costs
scalability
      </p>
      <p>
        Notification
+
+
high
high
yes
no
± [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ]
± [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ]
low
high
yes
partially
      </p>
      <p>
        ± [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ]
unknown
unknown
unknown
yes
yes
and standardisation efforts ongoing, including but not
limited to:
• online services;2
• earlier efforts for a lightweight notification standard:
for instance the Event Notification Protocol (ESN) (see
“Requirements for Event Notification Protocol” [
        <xref ref-type="bibr" rid="ref22">23</xref>
        ]);
• pubsubhubbub: a simple, open, server-to-server
webhook-based pubsub (publish/subscribe) protocol as an
extension to Atom and RSS.3
5.
      </p>
    </sec>
    <sec id="sec-6">
      <title>METHODOLOGY</title>
      <p>
        To the best of our knowledge, this is the first study
regarding the dynamics of documents and entities of the Linked
Open Data Web. Hence, the methodologies used in our
evaluation are inspired by legacy related work for Web
documents. Specifically, we applied similar evaluation methods
– and indeed try to answer similar questions – as presented
in [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ]. The experiments require a large data set which is
constantly monitored over a long timespan to conclude
significant findings: we are not aware of any significant,
heterogeneous and publicly available data-set of Linked Open
Data resources which includes a complete history of changes.
Nevertheless, we have access to such a dataset collected for
an extended period in early 2009; although the dataset was
originally collected for a different purpose – and thus, as we
will see is not as suitable for our analysis as a bespoke corpus
might be – we can derive some illustrative statistics which
give some early insights into the dynamic nature of Linked
Data on the Web.4 Next, we describe how this dataset was
monitored and which methods we use for our evaluation.
5.1
      </p>
    </sec>
    <sec id="sec-7">
      <title>Monitoring</title>
      <p>
        To gain first insights about the dynamics of resources of
the Linked Open Data Web we analyse 24 data dumps
collected by weekly snapshots of the 7 hop neighborhood of
Tim Berners-Lee’s FOAF file5. The weekly snapshots were
collected using the MultiCrawler framework [
        <xref ref-type="bibr" rid="ref13">14</xref>
        ] with the
following steps applied in each crawl cycle:
1. gathering the content of a list of URIs;
2. parsing of RDF/XML content;
2http://www.changedetection.com/
3http://code.google.com/p/pubsubhubbub/
4Notably, this dataset was already studied by Biessmann
et. al. [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ] w.r.t. to dependency dynamics between people
described in the data set.
5http://www.w3.org/People/Berners-Lee/card
3. extracting of all URIs at the subject and object
position of a triple;
      </p>
      <sec id="sec-7-1">
        <title>4. shuffling list of extracted URIs;</title>
        <p>5. applying a per-domain limit for the URIs (5000 URIs
per PLD).</p>
        <p>
          Please note that steps 4) and 5) were done for politeness
reasons to prevent too many parallel HTTP requests to one
server: these steps introduce a non-deterministic element
into our crawl and thus, we did not monitor a fixed list of
URIs every week. Indeed, this passive monitoring makes
change frequency analysis more challenging [
          <xref ref-type="bibr" rid="ref8">9</xref>
          ]. We have to
deal with an incomplete history of sources, wherein it is very
likely that many sources appear only once in the snapshot –
thus, we sometimes present statistics which use only a small
subset of the total dataset: the subset derived from sources
that were available in more than 20 of the 24 snapshots.
5.2
        </p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Data Corpus</title>
      <p>The data collection was performed over 24 weeks starting
from the 2nd of November 2008 and contains 550K RDF/XML
documents with a total of 3.3M unique subjects (∼6
entities appearing in the subject position per source) with 2.8M
locally defined entities per our definition 4.
5.3</p>
    </sec>
    <sec id="sec-9">
      <title>Change detection function</title>
      <p>
        The change detection of a document Cd(t, t‘) or entity
Cde(t, t‘) between two snapshots t, t‘ is a trivial task as long as
the statements of the resource do not contain blank nodes [
        <xref ref-type="bibr" rid="ref23">24</xref>
        ].
For our preliminary evaluation, we used a simple change
detection algorithm – based on a merge-sort scan over the
weekly snapshots – as follows:
1. skolemise blank nodes within a document;
2. sort all relevant statements for the change detection of
an document or entity by their syntactic natural order
(subject-predicate-object-[context]);
3. perform pairwise comparison of the statements by
scanning two snapshots in linear time;
4. trigger a detection of change (either w.r.t. a document
or entity) as soon as the order of the statements
differs between two snapshots (e.g. new statements were
added or removed).
5.4
      </p>
    </sec>
    <sec id="sec-10">
      <title>Evaluation</title>
      <p>In this subsection, we describe in detail the evaluation we
performed on the data set.</p>
      <p>Document-centric evaluation. Firstly – and as a baseline
– we performed a document-centric evaluation which allows
us to compare our results with earlier studies about HTML
documents. For this study, we compute the changes of a
document as defined in Definition 1.</p>
      <p>Entity-centric evaluation. Secondly, we studied the change
frequency of entities from an entity-per-document
perspective as defined in Definition 2. In fact, more accurately
we analysed the change frequency from a
local-entity-perdocument perspective – a notion which follows intuitively
from Definitions 2 and 4: to detect a change in an entity
Celocal (t, t‘), we compare only the statements which 1) are
d
contained in documents whose URIs matches on the PLD
level with the entity URI and 2) in which the entity URI
appears in the statement. Thus, we consider only the changes
from documents in the locality of the entity as defined in
Definition 4.
5.5</p>
    </sec>
    <sec id="sec-11">
      <title>Change Process - A Poisson Process</title>
      <p>
        Finally, for the purposes of comparison, we use an
established model for changes of Web documents. Previously
published studies [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ] report that changes in Web documents
can be modeled as a Poisson process (Equation 1). Poisson
processes are used – for example – to model arrival times of
customers, the times of radioactive emissions or the number
of sharks appearing on a beach in a given year. The model
allows to calculate the probability of a number of events
occurring in a fixed period of time given that (i) the events are
independent of the time elapsed since the last event and (ii)
the events occur with a known average frequency rate λ. The
parameter λ is the expected ‘events’ or ‘arrivals’ that occur
per the required unit-of-time (in our case, a week). Further,
let N(t + τ ) − N(t) be the number of changes in an interval
(t, t + τ ] with τ given as the number of weeks – to take an
example in our scenario, if an entity e changed five times in
a window of the last 10 weeks, then Ne(14 + 10) − Ne(14) = 5.
Finally, let k be the number of occurrences of a document or
entity in the total monitoring time (24 weeks in our case).
Then, according to the Poisson process, the probability of
an event occurring within a given interval (t, t + τ ] is given
as:
(λτ )k
      </p>
      <p>k!
P [N(t + τ ) − N(t) = k] =
exp(−λτ ) for k = 1, 2... (1)
6.</p>
    </sec>
    <sec id="sec-12">
      <title>FINDINGS</title>
      <p>In this section we present several early findings about the
change frequency of resources on the Linked Data Web.</p>
      <p>
        Firstly, we examine the usage of Etag and Last-Modified
HTTP header fields, followed by an analysis of the various
dynamic aspects which are aligned to the studies of the
traditional Web in [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ].
6.1
      </p>
      <p>Usage of Etag and Last-Modified</p>
      <p>One way to detect changes is to use the information
contained in HTTP response headers as discussed in Table 1.
The HTTP protocol offers two header fields to indicate a
change of a document, viz: the Etag and Last-Modified
fields. Using such methods of change detection is more
economical in that it avoids the need for content sniffing.</p>
      <p>We verified the usage (or lack thereof) of these two fields
for all the documents in our corpus; Table 2 summarises the
findings:</p>
      <sec id="sec-12-1">
        <title>Header field</title>
        <p>only Etag
only Last-Modified
Both
None</p>
      </sec>
      <sec id="sec-12-2">
        <title>Fraction</title>
        <p>7.12%
8.18%
16.75%
67.95%</p>
        <p>
          Similarly to studies about the usage of these two fields
for HTML documents [
          <xref ref-type="bibr" rid="ref9">10</xref>
          ], we found that 67.95% of the
550K documents did not report either of these two fields.
Both fields were available by 16.75% of all the documents.
Thus, we have to rely on actively monitoring of documents
to detect their changes.
6.2
        </p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>Access and lifespan distribution</title>
      <p>
        We move now to analysis involving the content of data in
our corpus. Firstly, we are interested in characterising the
distribution of the number of accesses (i.e., appearances)
and the lifespan (i.e., the time interval between the first
and last appearance) of documents and entities respectively.
This is a slightly different computation from [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ], where, for
example, the authors estimated the lifespan of a document
by doubling the time the document was seen in the
monitoring window if the document occurred at the beginning of
the experiment but not at the end. Figure 2 contains the
plots of the frequency and lifespan distribution for the
documents (crosses) and entities (circles); we observe that the
distributions follow approximately an “80-20” law.
      </p>
      <p>From this figure, we can also conclude that only a fraction
of the documents appeared frequently in the different
snapshots – considering the importance of having as much
information as possible to apply and verify our change frequency
model (Section 5.5) – and thus to gain a good overview
about their dynamics – going forward, we will give special
consideration to the subset of our corpus derived from
documents that appear in at least 20 weekly snapshots and ignore
missing observations when considering changes. Again, this
is necessitated by the non-deterministic factor in our
incidentally crawled snapshots.</p>
      <p>1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24
number of weeks</p>
    </sec>
    <sec id="sec-14">
      <title>How often do the resources change?</title>
      <p>Next, we will analyse the average change frequency of a
resource. For the purposes of this analysis, we only consider
the subset of the corpus which features resources that
appear in more than 20 snapshots. Let us assume a document d
changed 12 times during our monitoring interval of 24 weeks.
In this case we can estimate the average change frequency
of d to be 24 weeks/12 = 2 weeks. Following this example,
the results for average change frequency of documents and
1000000
1t00000
n
u
o
c
10000
1000
documents (accesses)
documents (lifespan)
entities (accesses)
entities (lifespan)
80%
70%
60%
50%
n
o
it
rca40%
f
30%
20%
10%
1000000
100000
iitts10000
e
n
foe1000
r
e
b
um 100
n
10
1
entities are summarised in Figure 3. The left side of the
diagram shows the percentage of all resources that were not
observed to change (static resources). The right side of the
diagram shows the percentage of non-static resources that
were observed to have an average change frequency within
the given interval. An interesting finding is that 62% of the
total documents did not change at all, along with 68% of the
entities. Further, we see that the fraction of documents is
increasing with bigger change intervals, whereas for entities
it is quite the opposite: by inspecting the data closer, we
figured out that 51% of the entities with a change frequency of
less than 1 week appear in more than one ‘local’ document.
Thus, for example, one document may change the
description of many local entities: along these lines, Figure 4 shows
the distribution of the number of frequency of entities
appearing in a given number of documents, where again, we
can observe a power law distribution.</p>
      <p>62.12%</p>
      <p>68.88%
static
52%</p>
      <p>24% 23%
9%
9%</p>
      <p>10%
&lt;1 week</p>
      <p>≤&gt;11wmeoenkth ≤&gt;13mmoonntthh
average change frequency
59%</p>
      <p>Continuing, we now study how quickly and what fraction
of the documents and entities changes over time. Along
these lines, we count how many documents – and
respectively entities – changed after a certain time period.
Figure 5 presents the cumulative change function for documents
(circles) and entities (squares). The graph cumulatively
shows how many documents and entities had changed after
X weeks. The plot contains the cumulative change function
for all resources (appearing at least once), and for resources
that appeared in at least 20 snapshots. Again, the plot
correlates with Figure 3 in that for the subset of the corpus with
more than 20 observations, we can also see a large amount
of entities changing after the first week, with a more
gradual increase in observed document changes. An interesting
observation is that the entities with more than 20
observations show a higher propensity to change; one could assume
that such entities are better linked (and thus appear more
often in our crawl) and so are reused in more documents (cf.
Figure 4).</p>
      <p>documents
documents (#occ &gt; 20 weeks)
entities
entities (#occ &gt; 20 weeks)
30 %
n
ito20 %
c
a
fr
10 %
0 %
2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24
number of weeks</p>
    </sec>
    <sec id="sec-15">
      <title>Change process - A mathematical model?</title>
      <p>Next, we analyse whether we can apply the Poisson model
presented in Section 5.5 to the changes of documents and
entities detected in our analysis. Therefore, we must compute
the average change rate λ for each document d and entity e.
We group the documents and entities with the same change
rate and plot their distribution of successive change
intervals; e.g., a document which changed in week 2 and 6 has
a successive change interval of 4. If the changes can be
modeled as a Poisson process, the resulting graph should be
distributed exponentially.</p>
      <p>For illustration, we selectively present the graph for
documents with an average change frequency of 4 weeks
(Figure 6) and the graph for entities with an average change
frequency of 4 weeks (Figure 7). We performed a
Poisson regression ( log-linear regression) and use the maximum
likelihood method to estimate the parameters. The
predicated poisson process is plotted in the graphs as the line and
describes the observed data quite well, despite some small
variations. Similar effects are observed for around half of
the other plots. However, we also spotted several graphs
for documents and entities in which the Poisson model does
not well describe the observed data points: The main reason
for this observation is that there are not enough available
sample points. As a conclusion of the findings: we currently
cannot accept or reject the described change model with
statistical significancy. Further studies with more data samples
are required.</p>
      <p>We found that in 90% of all documents less than 10% of
the entities changed, as depicted in Figure 8, which shows
the distribution of the average fraction of entities that changed
per document. It is hence safe to assume that – in the
context of Linked Data – the finer-grained entity-centric
perspective for changes is superior, compared to the more
traditional document-centric point of view.</p>
      <p>Drawing towards a conclusion to our analysis, we now
discuss the observed changes for the documents over time.
Therefore, we defined the following three main change
categories:
• Update (U) – that is, between two snapshots of a
document, the entities described were the same but the
information about the entities changed: new statements
were added and/or removed;
• Add (A) – that is, between two snapshots of a
document, new entities were added;
• Del (D) – that is, between two snapshots of a
document, entities were deleted;
• Combination of the three categories mentioned above:</p>
      <p>UA, UD, AD; UAD.</p>
      <p>Table 3 lists the fraction of documents which encountered
such a change (or combination thereof) for each of the seven
categories. We can see that 76% of the documents have
only entity updates as changes, whereas in 9.46% of the
documents new entities were added.</p>
      <p>U
A
D</p>
    </sec>
    <sec id="sec-16">
      <title>CONCLUSION</title>
      <p>We motivated this work by highlighting the importance
of a fundamental understanding of dataset dynamics with
respect to Linked Open Data sources; we further claim that
such knowledge can be leveraged to optimise existing
systems and algorithms, such as making incremental index
updates techniques more efficient. Further, we discussed in
detail the differences between document-centric and
entitycentric dynamics together with possible approaches for change
detection: content-monitoring, HTTP header monitoring,
and active notifications.</p>
      <p>The findings we gained from weekly snapshots of the
neighborhood graph of Tim Berners-Lee FOAF file are the
following:
• less than 35% of the monitored documents contained
Etag and Last-Modified HTTP header fields in the
response;
• a surprisingly small amount (∼ 35%) of the monitored
resources changed over the time interval of 24 weeks;
• half of the documents that changed had a change
frequency of more than 3 months – in contrary, on a
entity-centric level, half of the entities had a change
frequency of less than a week applying our definition
of local entities (based on PLD correspondences
between document and entity);
• comparing our results to previous published studies we
cannot verify that the change frequency of the
documents and entities follow entirely the change model of
a Poisson process.</p>
      <p>We should perhaps look at these early findings with a
critical eye in that we did not actively monitor a fixed set
of sources. This work is very much an early attempt in this
field, and needs further exploration and research to fully
understand and exploit the change frequency of resources in
the Linked Data Web.
7.1</p>
    </sec>
    <sec id="sec-17">
      <title>Future Work</title>
      <p>
        Large scale experiment To verify our early findings and
derive statistical significant results, we plan to expend and
run our evaluation for a larger dataset which is monitored
over a longer time period. Further, we plan to study in more
detail the dynamics on a entity-centric level; e.g. studying
the dynamics of only authoritative entities as defined in [
        <xref ref-type="bibr" rid="ref17">18</xref>
        ]
or the dynamics of the global entities as defined in Section 3.
      </p>
      <p>Active monitoring. A major drawback of the current study
is the monitoring method used for our data set. To overcome
the problem of an incomplete change history, we will actively
survey a selected set of documents over a long time period,
thus creating a tailored corpus for our analysis. In addition
to active monitoring, we plan to study how we can
dynamically adapt the monitoring interval based on the estimated
change frequency of a resource.</p>
      <p>
        Fine-grained analysis of changes on a entity-centric level
Finally, the findings of this work will be integrated into an
existing system which aims to execute live queries over the
LOD Web, which uses efficient data summary approaches [
        <xref ref-type="bibr" rid="ref12">13</xref>
        ].
Thus, using our analytics, we would hope to discern
documents which are highly dynamic and those which are more
static: highly dynamic documents would thus be better
suited to direct-lookup approaches, whereas static data would
be more suited to index summaries (or indeed, full-blown
data warehousing approaches) for query-answering.
Similarly, we could also investigate what kinds of statements for
an entity in a document changes; e.g. a rdf:type statement
should be rather very static, whereas a statement describing
the values of sensor data is rather very dynamic.
8.
      </p>
    </sec>
    <sec id="sec-18">
      <title>ACKNOWLEDGEMENTS</title>
      <p>Our work has partly been supported by the European
Commission under Grant No. SFI/08/CE/I1380 (Lion-2),
and under Grant No. 231335, FP7/ICT- 2007.4.4 iMP project.
Further, this work has greatly benefited from the feedback
of Bernhard Haslhofer and Nico Popitsch as well as from
discussions with Andreas Harth.
9. REFERENCES
[1] Dataset dynamics (esw wiki).</p>
      <p>http://esw.w3.org/topic/DatasetDynamics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[2] REST and RDF Granularity</article-title>
          . http://dret.typepad.com/dretblog/2009/05/ rest-and
          <article-title>-rdf-granularity</article-title>
          .html, May
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mozafari</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Cho</surname>
          </string-name>
          .
          <article-title>On the evolution of wikipedia</article-title>
          .
          <source>In Int. Conf. on Weblogs and Social Media</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Biessmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          .
          <article-title>Analysing dependency dynamics in web data</article-title>
          .
          <source>In Linked AI: AAAI Spring Symposium ”Linked Data Meets Artificial Intelligence”</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Linked Data-The Story So Far</article-title>
          . Special Issue on Linked Data,
          <source>International Journal on Semantic Web and Information Systems (IJSWIS)</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. E.</given-names>
            <surname>Brewington</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Cybenko</surname>
          </string-name>
          .
          <article-title>How dynamic is the web? Comput</article-title>
          . Netw.,
          <volume>33</volume>
          (
          <issue>1-6</issue>
          ):
          <fpage>257</fpage>
          -
          <lpage>276</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7] G. Cheng and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qu</surname>
          </string-name>
          .
          <article-title>Term dependence on the semantic web</article-title>
          .
          <source>In ISWC '08: Proceedings of the 7th International Conference on The Semantic Web</source>
          , pages
          <fpage>665</fpage>
          -
          <lpage>680</lpage>
          , Berlin, Heidelberg,
          <year>2008</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cho</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          .
          <article-title>The evolution of the web and implications for an incremental crawler</article-title>
          .
          <source>In VLDB</source>
          , pages
          <fpage>200</fpage>
          -
          <lpage>209</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cho</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          .
          <article-title>Estimating frequency of change</article-title>
          .
          <source>ACM Trans. Internet Techn.</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>256</fpage>
          -
          <lpage>290</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Clausen</surname>
          </string-name>
          .
          <article-title>Concerning Etags and Datestamps</article-title>
          .
          <source>In Proceedings of the 4th International Web Archiving Workshop</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fielding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gettys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mogul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Frystyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Masinter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lafon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nottingham</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Reschke</surname>
          </string-name>
          .
          <source>HTTP/1</source>
          .1, part 6:
          <string-name>
            <surname>Caching. Internet Draft</surname>
          </string-name>
          ,
          <source>Expires: April</source>
          <volume>29</volume>
          ,
          <year>2010</year>
          , IETF HTTPbis Working Group,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fielding</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Taylor</surname>
          </string-name>
          .
          <article-title>Principled design of the modern Web architecture</article-title>
          .
          <source>ACM Trans. Internet Technol.</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <fpage>115</fpage>
          -
          <lpage>150</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karnstedt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          , K.-U. Sattler, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          .
          <article-title>Data summaries for on-demand queries over linked data</article-title>
          .
          <source>In Proceedings of the 19th World Wide Web Conference (WWW2010)</source>
          , Raleigh,
          <string-name>
            <surname>NC</surname>
          </string-name>
          , USA, Apr.
          <year>2010</year>
          . ACM Press.
          <article-title>accepted for publication.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          .
          <article-title>Multicrawler: A pipelined architecture for crawling and indexing semantic web data</article-title>
          .
          <source>In International Semantic Web Conference</source>
          , pages
          <fpage>258</fpage>
          -
          <lpage>271</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hartung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kirsten</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <article-title>Analyzing the evolution of life science ontologies and mappings</article-title>
          .
          <source>In DILS '08: Proceedings of the 5th international workshop on Data Integration in the Life Sciences</source>
          , pages
          <fpage>11</fpage>
          -
          <lpage>27</lpage>
          , Berlin, Heidelberg,
          <year>2008</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Haslhofer</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Popitsch</surname>
          </string-name>
          .
          <article-title>DSNnotify - detecting and fixing broken links in linked data sets</article-title>
          .
          <source>In Proceedings of the 8th International Workshop on Web Semantics (WebS 09)</source>
          , co-located
          <source>with DEXA</source>
          <year>2009</year>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hausenblas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Halb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Raimond</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          .
          <article-title>What is the Size of the Semantic Web? In I-Semantics 2008:</article-title>
          <source>International Conference on Semantic Systems</source>
          , Graz, Austria,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          .
          <article-title>Scalable Authoritative OWL Reasoning for the Web</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>I.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <source>Architecture of the World Wide Web</source>
          , Volume One.
          <source>W3C Recommendation 15 December</source>
          <year>2004</year>
          ,
          <source>W3C Technical Architecture Group (TAG)</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>H.-T.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Leonard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Loguinov</surname>
          </string-name>
          . Irlbot:
          <article-title>Scaling to 6 billion pages and beyond</article-title>
          .
          <source>ACM Trans. Web</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B. S.</given-names>
            <surname>Lerner</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Habermann</surname>
          </string-name>
          .
          <article-title>Beyond schema evolution to database reorganization</article-title>
          .
          <source>In OOPSLA/ECOOP '90: Proceedings of the European conference on object-oriented programming on Object-oriented programming systems, languages, and applications</source>
          , pages
          <fpage>67</fpage>
          -
          <lpage>76</lpage>
          , New York, NY, USA,
          <year>1990</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ramamritham</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Chakrabarti</surname>
          </string-name>
          .
          <article-title>Monitoring the dynamic web to respond to continuous queries</article-title>
          .
          <source>In WWW '03: Proceedings of the 12th international conference on World Wide Web</source>
          , pages
          <fpage>659</fpage>
          -
          <lpage>668</lpage>
          , New York, NY, USA,
          <year>2003</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Reddy</surname>
          </string-name>
          and M. Fisher.
          <article-title>Requirements for Event Notification Protocol</article-title>
          .
          <source>Internet Draft, May</source>
          <volume>1</volume>
          ,
          <year>1998</year>
          , IETF WEBDAV Working Group,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tummarello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Morbidoni</surname>
          </string-name>
          , R. Bachmann-Gmu¨r, and
          <string-name>
            <given-names>O.</given-names>
            <surname>Erling</surname>
          </string-name>
          . Rdfsync:
          <article-title>Efficient remote synchronization of rdf models</article-title>
          .
          <source>In ISWC/ASWC</source>
          , pages
          <fpage>537</fpage>
          -
          <lpage>551</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>