<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reliability Analyses of Open Government Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Ceolin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luc Moreau</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kieron O'Hara</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guus Schreiber</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alistair Sackley</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wan Fokkink</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Willem Robert van Hage</string-name>
          <email>willem.van.hage@synerscope.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nigel Shadbolt</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hampshire County Council</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SynerScope B.V.</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>VU University Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Public authorities are increasingly sharing sets of open data. These data are often preprocessed (e.g. smoothened, aggregated) to avoid to expose sensible data, while trying to preserve their reliability. We present two procedures for tackling the lack of methods for measuring the open data reliability. The rst procedure is based on a comparison between open and closed data, and the second derives reliability estimates from the analysis of open data only. We evaluate these two procedures over data from the data.police.uk website and from the Hampshire Police Constabulary in the UK. With the rst procedure we show that the open data reliability is high despite preprocessing, while with the second one we show how it is possible to achieve interesting results concerning the open data reliability estimation when analyzing open data alone.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Open Government Data are valuable for boosting the economy, enhancing the
transparency of public administration and empowering the citizens. These data
are often sensitive and so need to be preprocessed for privacy reasons. In the
paper, we refer to the public Open Government Data as \open data" and to the
original data as \closed data".</p>
      <p>
        Di erent sources expose open data in di erent manners. For example, Crime
Reports [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and data.police.uk [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] both publish UK crime data, but in di
erent format (maps vs. CSV les), level of aggregation, smoothing and timeliness
(daily vs. monthly update), which all represent possible reasons for reliability
variations. For di erent stakeholders it is important to understand how reliable
di erent sources are. The police, who can access the closed data, needs to know
if open data are reliable enough e.g. to be used in projects involving the citizens.
The citizens wish to know the reliability of the di erent datasets to understand
the reasons for di erences between authoritative sources. We present two
procedures to cope with the lack of methods to analyze these data: one for computing
the reliability of open data by comparing them with the closed data, and one to
estimate variations in the reliability of the open data by relying only on these.
      </p>
      <p>
        The analysis of open data is spreading, led by the Open Data Institute
(http://www.theodi.org) and others. For instance, Koch-Weser [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presents
an interesting analysis of the reliability of China's Economic Data, thus analyzing
the same aspect as we are interested in, on a di erent typology of dataset. Tools
for the quality estimation of open data are being developed (e.g. Talend Open
Studio for Data Quality [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], Data Cleaner [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]), but their goal is less targeted
than ours, since they aim at quantifying the quality of open data in general as to
provide a substrate for a more comprehensive open data analysis infrastructure.
Relevant for this work is also a paper from Ceolin et al. that uses a statistical
approach to model categorical Web data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and one that uses provenance to
estimate reliability [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We plan to adopt the approach proposed by Ebden et
al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to measure the impact of di erent processes on the data.
      </p>
      <p>The rest of this paper is structured as follows: Section 2 describes a procedure
for measuring the reliability of open data given closed data and a case study
implementation; Section 3 presents a procedure for analyzing open data and a
case study; lastly, Section 4 provides nal discussion.
2</p>
      <p>
        Procedure for Comparing Closed and Open Data
The UK Police Home O ce aggregates (i.e., presents coarsely) and smoothens
(introduces some small error) the open data for privacy reasons. We represent
the open data provenance with the PROV Ontology [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] as in Fig. 1. In general,
a faulty aggregation process or aggregating data coming from heterogeneous
sources not properly manipulated might unexpectedly a ect the resulting data
reliability, while smoothing should a ect it explicitly but in a limited and
controlled manner. The following procedure aims at capturing such variation:
rdf:type
prov:Entity
prov:Activity
      </p>
      <p>Aggregation</p>
      <p>Smoothing
Closed Data</p>
      <p>Open Data</p>
      <p>Select the relevant data Closed data might be spurious, so we select the data
items that are relevant for our analyses. The selection of the data might
involve the temporal aspect (i.e. only data referring to the relevant period
are considered), their geographical location (select only the data regarding
the area of interest), or other constraints and their combination;
Roll up categorical data There exists a hierarchy of categories because each
level is available to a di erent audience: open data are presented coarsely to
the citizens, while closed data are ne grained. We bring the categorization
to the same level, hence bringing the closed data to the same level as the
open data.</p>
      <p>
        Compare the corresponding counts Di erent measures are possible, because
the di erence between datasets can be considered from di erent points of
view: relative, absolute, etc.. For instance, the ratio of the correct items over
the total amount or the Wilcoxon signed-rank test [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>Case Study 1 We compare a set of crime counts per categories grouped per
neighbourhood and month from data.police.uk with a limited set (30,436
relevant entries) of corresponding closed data from the Hampshire Constabulary
by implementing the procedure above as follows:
Data Selection Select the data for the relevant months and geographical area.</p>
      <p>
        In this latter case, we load the KML le describing the Hampshire
Constabulary area using the maptools library [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in the R environment [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and check
if the crimes coordinates occur therein using the SDMTools library [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ];
Data Aggregation We apply two kinds of aggregation: temporal, to group
together data about the same month and geographical, to aggregate per
neighbourhood. The closed data items report the address of occurrence of
the crimes, while the open data are aggregated per police neighbourhood.
We match the zip code of the addresses and the neighbourhoods using the
MapIt API [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Data Comparison We average the result of the Wilcoxon signed-rank test
applied per neighbourhood, to compare open and aggregated closed data.</p>
      <p>For each neighbourhood we compute a Wilcoxon signed-rank test to check the
signi cance of the di erence between open and closed data and we average the
outcomes (see Table 1a). We compute the test on the di erences of the two counts
(open and closed data) to check whether the estimated average of the distribution
of the di erences is zero (that is, the two distributions are statistically equivalent)
or not.</p>
      <p>
        The results at our disposal are limited, since we could analyze only two
complete months. Still, we can say that smoothing, in these datasets, introduces
a small but signi cant error. The highest error average (2.75) occurs with the
entry with the highest error variance: this suggests that the higher error is due
to a few, sparse elements, and not to the majority of the items. To prove this,
we checked the error distribution among the entities and we reported the results
in Table 1b. A 2 test [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] at 95% con dence level con rms that the two error
distributions do not di er in a statistically signi cant manner.
3
      </p>
      <p>Procedure for Analyzing Open Data
We propose here a procedure for analyzing open data alone, to be used when
closed data are not available, which provides weaker but still useful results,
(b) Percentage of items in each open dataset
presenting a relative error of at most 0%,
25%, 50%, 75% and 100% with respect to the
corresponding closed data item.</p>
      <p>Month % of Entries per Relative Error</p>
      <p>0% 25% 50% 75% 100%
month 1 35% 44% 65% 74% 96%
month 2 34% 43% 57% 65% 91%
compared to the previous one. It compares each dataset with the consecutive
one, measures their similarity and pinpoints the occurrence of possible
reliability changes based on variations of similarity over time. We use a new similarity
measure for comparing datasets, that aggregates di erent similarity \tests"
performed on couples of datasets. Given two datasets d1 and d2, their similarity is
computed as follows:</p>
      <p>sim(d1; d2) = avg(t1(d1; d2); : : : ; tn(d1; d2))
where avg aggregates the results of n similarity tests ti, with i 2 f1 : : : ng. We
propose the following families of tests, although we are not restricted to them:
Statistical test Check with a statistical test (e.g. Wilcoxon signed-rank test)
if the data are drawn from signi cantly di erent distributions.</p>
      <p>
        Model Comparison test Build a model (e.g. linear regression [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] or Support
Vector Machines [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) on one of the two datasets and evaluate its
performance (precision, recall) over the other dataset. These models represent an
abstraction over the rst dataset and by evaluating them over the other one,
we check, according to such a model, how similar the two datasets are.
The tests can be aggregated, for instance, by averaging them or by merging them
in a \subjective opinion" [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which is a construct of a probabilistic logic that
is equivalent to a Beta probability distribution about the correct value for the
similarity. The expected value of the Beta is close to the arithmetical average,
but the variance represents the uncertainty in our calculation, since it reduces as
long as we consider more tests. The similarity measure alone does not stand for
reliability: there can be many reasons for a similarity variation (e.g. a new law or
a particular event that makes the crime rate rise) without implying a reliability
change. Also, a similarity value alone might be di cult to interpret in terms of
reliability, when a gold standard is not available. So we analyze the similarity of
consecutive datasets to pinpoint items that possibly present reliability variations:
if the similarity between datasets remains similar for a period of time, and then
a variation occurs, one of the possible reasons for such a variation is a change
in the data reliability. Unfortunately, we can not discriminate between this and
other causes, unless we have additional information at our disposal.
Case Study 2 We analyze the police open data for the Hampshire Constabulary
from data.police.uk, that consist of crime counts, aggregated per
neighbourhood from April 2011 to December 2012. We know that in this period open
data creation policy changes occurred. These might have a ected the datasets
reliability. We compare the distribution of the crime counts among the crime
categories, and we represent the similarity between two datasets as the
percentage of neighbourhoods that are statistically similar (according to a Wilcoxon
signed-rank test). The results of the comparison are reported in Figure 2, where
each point represents the similarity between two datasets, in sequence. At the
twelfth comparison the similarity trend breaks and then starts a new one. That
is likely to be a point where the reliability diverges as the similarity variation
possibly hints, and it actually coincides with a policy change (the number of
neighbourhoods varies from 248 to 232), and since the area divided by these
neighbourhoods is the same, this possibly introduces a variation in the impact
of the smoothing error, but we do not have at our disposal a con rmation of such
impact. As we stressed earlier, the procedure allows us only to pinpoint possibly
problematic data, but without additional information, our analysis cannot be
precise, that is, we cannot be certain about the reason of the similarity change.
y
itr
a
il
m
i
S
We presented two procedures for the computation of the reliability of open data:
one based on the comparison between open and closed data, the other one based
on open data alone. Both procedures have been evaluated using data from the
data.police.uk website and from the Hampshire Police Constabulary in the
UK. The rst procedure allows us to estimate the reliability of open data, and
shows that smoothing procedures, although introducing some error, preserve a
high data reliability. The second procedure is useful to grasp indications about
the data reliability, although more weakly than the rst one, since it allows
only to pinpoint possible reliability variations in the data. Despite the fact that
open data are exposed by authoritative institutions, these procedures allow us
to enrich the open data with information about their reliability, to increase the
con dence of both the insider specialist and the common citizen who use them
and to help in understanding possible discrepancies between data exposed by
di erent authorities. We plan to extend the range of analyses applied and of
datasets considered. Moreover, we intend to map the data with Linked Data
entities to combine the statistical analyses with semantics.
      </p>
      <p>Acknowledgements This work is supported in part under SOCIAM: The
Theory and Practice of Social Machines; the SOCIAM Project is funded by the UK
Engineering and Physical Sciences Research Council (EPSRC) under grant
number EP/J017728/1.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Bivand</surname>
          </string-name>
          .
          <article-title>Tools for reading and handling spatial objects</article-title>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ceolin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Groth</surname>
          </string-name>
          , and
          <string-name>
            <surname>W. R. van Hage. Calculating</surname>
          </string-name>
          <article-title>the Trust of Event Descriptions using Provenance</article-title>
          .
          <source>In SWPM</source>
          , pages
          <volume>11</volume>
          {
          <fpage>16</fpage>
          . CEUR-WS.org, Nov.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ceolin</surname>
          </string-name>
          , W. R. van Hage,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fokkink</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          .
          <article-title>Estimating Uncertainty of Categorical Web Data</article-title>
          .
          <source>In URSW</source>
          , pages
          <volume>15</volume>
          {
          <fpage>26</fpage>
          . CEUR-WS.org,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. CrimeReports. Crimereports. https://www.crimereports.co.uk/,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>N.</given-names>
            <surname>Cristianini</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Shawe-Taylor</surname>
          </string-name>
          .
          <article-title>An Introduction to Support Vector Machines and other kernel-based learning methods</article-title>
          . Cambridge University Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ebden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Huynh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Moreau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramchurn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Roberts</surname>
          </string-name>
          .
          <article-title>Network analysis on provenance graphs from a crowdsourcing application</article-title>
          .
          <source>In IPAW</source>
          , pages
          <volume>168</volume>
          {
          <fpage>182</fpage>
          . Springer-Verlag,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>F.</given-names>
            <surname>Galton</surname>
          </string-name>
          .
          <article-title>Regression Towards Mediocrity in Hereditary Stature</article-title>
          .
          <source>Journal of the Anthropological Institute</source>
          ,
          <volume>15</volume>
          :
          <fpage>246</fpage>
          {
          <fpage>263</fpage>
          ,
          <year>1886</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Human</given-names>
            <surname>Inference</surname>
          </string-name>
          .
          <source>DataCleaner</source>
          . http://datacleaner.org,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>A. J sang.</surname>
          </string-name>
          <article-title>A logic for uncertain probabilities</article-title>
          .
          <source>Int. Journal Uncertainty Fuzziness Knowledge-Based Systems</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <volume>279</volume>
          {
          <fpage>311</fpage>
          ,
          <year>June 2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>I. N.</given-names>
            <surname>Koch-Weser</surname>
          </string-name>
          .
          <article-title>The Reliability of China's Economic Data: An Analysis of National Output</article-title>
          . http://www.uscc.gov/sites/default/files/Research/ TheReliabilityofChina'sEconomicData.pdf, Jan.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. mySociety. MapIt. http://mapit.mysociety.orgs,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>K.</given-names>
            <surname>Pearson</surname>
          </string-name>
          .
          <article-title>On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling</article-title>
          .
          <source>Philosophical Magazine</source>
          ,
          <volume>50</volume>
          :
          <fpage>157</fpage>
          {
          <fpage>175</fpage>
          ,
          <year>1900</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>R Core</given-names>
            <surname>Team. R:</surname>
          </string-name>
          <article-title>A Language and Environment for Statistical Computing</article-title>
          . R Foundation for Statistical Computing, Sept.
          <year>2012</year>
          . ISBN 3-900051-07-0.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Talend</surname>
          </string-name>
          .
          <article-title>Talend Open Studio for Data Quality</article-title>
          . http://www.talend.com/ products/data-quality,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>United</surname>
          </string-name>
          <article-title>Kingdom Police Home O ce. data.police.uk. data.police</article-title>
          .uk,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>J. VanDerWal</surname>
            , L. Falconi,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Januchowski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Shoo</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Storlie. SDMTools: Species Distribution Modelling</surname>
          </string-name>
          <article-title>Tools: Tools for processing data associated with species distribution modelling exercises</article-title>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. W3C.
          <article-title>PROV-O: The PROV Ontology</article-title>
          . http://www.w3.org/TR/prov-o/,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>F.</given-names>
            <surname>Wilcoxon</surname>
          </string-name>
          .
          <article-title>Individual comparisons by ranking methods</article-title>
          .
          <source>Biometrics Bulletin</source>
          ,
          <volume>1</volume>
          (
          <issue>6</issue>
          ):
          <volume>80</volume>
          {
          <fpage>83</fpage>
          ,
          <year>1945</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>