<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Retrieval Interface for Statistical Research Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Bahls</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Klaus Tochtermann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leibniz Information Centre for Economics (ZBW)</institution>
          ,
          <addr-line>Kiel</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>93</fpage>
      <lpage>103</lpage>
      <abstract>
        <p>Statistical research data is the foundation for empirical studies. Researchers in economics or social sciences often obtain such data from external sources through specially designed retrieval interfaces from statistical o ces, commercial data providers as well as from data agencies and other archives. With the advancements in data cataloguing and acquisition of long tail research data sets from individual scientists and institutes, the opportunity is there to install central services for a more holistic data search. In view of a rapid increase in amount of data available and by association an emerging retrieval problem, retrieval interfaces must make e ective use of provided metadata in order to help nd relevant data sets e ciently. This paper presents a multi-step retrieval interface that aims to support the researchers' natural approach to data search and composition. Starting with an idea of the concepts that are to be compared, users kick o their search with thesauri terms and successively specify requirements according to their priorities until suitable data can be selected easily from a manageable number of matching data sets. The prototype presented in this paper also provides means for convenient data harmonization, which is an essential aspect especially when combining statistical data from di erent sources.</p>
      </abstract>
      <kwd-group>
        <kwd>Research Data Management</kwd>
        <kwd>Semantic Digital Data Library</kwd>
        <kwd>Linked Data</kwd>
        <kwd>Statistics</kwd>
        <kwd>Data Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A signi cant number of scienti c results are based on research data, since
research has become increasingly data-driven over the years [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, to
understand such scienti c publications in depth, documentation on underlying data
is a necessary means. To further provide transparency and enable replicability
in the end, respective data sets must be available as such, for which a reliable
infrastructure is required. Scienti c data needs to be maintained and organized
in archives.
      </p>
      <p>With the advancement of computer technology, scienti c analyses are more
and more carried out with the aid of machines, as it allows for large amounts
of data being processed in short amount of time which has never been possible
before. While this certainly is one reason why science has become signi cantly
data-driven, it also leads to the fact that most scienti c data is maintained
in digital form already. This circumstance and the rise of the Web opens up
possibilities for a powerful information infrastructure for supporting these
aforementioned goals. Information resources nowadays can be delivered to any place in
the world within seconds, laying the ground for delivering the right information
to the right place at the right time, the precept of knowledge management.</p>
      <p>The Web together with its well-established Web 2.0 technologies has already
been recognized as a powerful media for promoting e cient exchange and
advancement in the scienti c domain. In this regard, the Leibniz Association has
recently started the research alliance Science 2.01 with a growing number of 30
associated institutes to jointly venture into a well-organized and integrated
environment of Web-based tools and services for the scienti c community to support
rapid exchange and good scienti c practice.</p>
      <p>
        The vision of a thought-out research data infrastructure ts well into this
theme, and many initiatives have formed in the last years, a whole movement to
e ectively enable exchange, citation and preservation of research data. However,
this task has proven non-trivial, as it opened up exhaustive discussions on
metadata schemes2, organized preservation and curation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], responsibilities [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], data
publication policies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as well as solutions to overcome issues of data protection
and usage rights, only to mention a few. Yet, these e orts have already lead to
signi cant advancements (TheDataHub3, DataCite4, and other).
      </p>
      <p>
        At present, e orts are being made to pick up research data as bibliographic
artifacts for re-use, transparency and citation[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In view of a rapid increase
in amount of data available and by association an emerging retrieval problem,
retrieval interfaces must make e ective use of provided metadata in order to help
nd relevant data sets e ciently.
      </p>
      <p>In this paper, we investigate how to make use of Semantic Web technologies
for providing an e cient and novel approach for the retrieval of statistical data
sets that follows a natural approach for data retrieval in the domain of statistics,
particularly in the context of economics or the social sciences. Section 2
elaborates on the practice of data acquisition in empirical research to gain a clear
picture on the purpose of our system. Related work is discussed in the
subsequent section, and Section 4 explains fundamental design decisions and outlines
a system architecture. Section 5 describes the user interface itself and how the
declared goals have been implemented into features. The paper eventually closes
with conclusions and outlook.
1 http://www.leibniz-science20.de
2 particularly important, as in contrast to textual publications, data cannot be
understood without documentation
3 http://datahub.io/
4 http://www.datacite.org/
2</p>
    </sec>
    <sec id="sec-2">
      <title>Retrieving Statistical Data</title>
      <p>In many cases, empirical researchers in economics and the social sciences are
to put together statistical indicators in large data tables. Typically, each
column represents one indicator while the rows represent respective data per year,
country or other so-called dimension. The data itself may be self-produced in
terms of studies and surveys or acquired from external sources such as
statistical o ces, a liated institutes or purchased from commercial data providers.
However, common practice is to combine several sources, since some indicators
may be obtained from one source while the data for other indicators may be
obtained from another one. In this regard, researchers have to be extra careful to
make sure respective data represents the same or su ciently similar statistical
population.</p>
      <p>
        To gain a clear picture of the goals of this research, we need to clearly
understand the purpose of the system. We have conducted interviews with economic
scientists which helped us gain insights in their work with research data.
Empirical researchers typically start out with an idea of concepts relevant in their
research (e.g. living standards, work conditions, economic growth, etc.). In
addition, they have further details in mind, for instance on reference periods, regions
to be included and distinguished or frequency of data acquisition in case of time
series data. As a result, the data set should be as consistent as possible with
respect to acquisition method, statistical universe and adjustments. To achieve
user acceptance, the system has to be practical in research settings [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and
therefore we aim to support this data harmonization procedure in a light-weight
manner.
      </p>
      <p>As a result, user communication should follow the below steps:</p>
      <sec id="sec-2-1">
        <title>1. Prompt for a list of concepts that are to be compared</title>
      </sec>
      <sec id="sec-2-2">
        <title>2. Let user specify additional requirements on the data</title>
        <p>3. Let explore and select matching data sets, allow for revisiting Step 2</p>
      </sec>
      <sec id="sec-2-3">
        <title>4. O er selected data for download</title>
        <p>After nishing Step 1, data sets associated with the concepts named should
be presented to the user. Speci cation of additional requirements should be
based on the metadata available for the data sets found. As soon as all relevant
requirements are given, the user may inspect and decide on these satisfying data
sets and proceed to download at last.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>There are many repositories on the Web that provide statistical data. Some of
them are provided by statistical o ces and data agencies (e.g. Federal Statistical
O ce of Germany5, EuroStat6, World Bank7), some are associated with
commercial providers (e.g. Thomson Reuters Datastream8, Statista9) and yet others
are maintained by journals, archives, libraries or independent organizations (e.g.
GESIS10, The Data Hub11, Dataverse repository of Economists Online12). All
of these portals are as heterogeneous as the kind and spectrum of data they
provide. Some of them provide interfaces for composition of customized data
tables where users pick and choose indicators and data records according to
their needs. Such features are also provided by the Nesstar system13, one of the
most prominent systems for data publishing and online analysis that is being
used by a large number of institutes. The Social Science Variables Database at
ICPSR14 allows for direct comparison of indicators with respect to a variety of
metadata, giving intuitive means to understand di erences in universe,
acquisition method and other between data sets. However, users of these systems are
to run keyword-based queries and browse through category trees in order to nd
relevant data sets individually, and therefore our approach follows a di erent
paradigm as presented in Section 2.</p>
      <p>
        Technical challenges in dealing with distributed sources and applying the
OLAP paradigm for retrieval of statistical data from the Linked Data cloud
have been addressed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We view this work as a major contribution for
building a scalable backend, whereas our work aims to provide a user interface and
communication design for data search and retrieval within the speci c setting
research data sharing.
      </p>
      <p>
        Other approaches are based on semantic links between data sets and research
articles [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which give textual context for otherwise sparsely described data
content and therefore improve data search by established Information Retrieval
techniques. These data links, typically given by persistent identi ers, however,
point to entire data bundles as a whole, whereas our approach aims to make
single indicators and values available for retrieval.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>System Architecture</title>
      <p>Following the steps presented in Section 2, we elaborate on the system
architecture of our data retrieval system. To support Step 1, a thesaurus should be
used, so that data sets associated with a particular concept can be found easily.
To enable the speci cation of requirements, metadata must be given in detail
5 https://www.destatis.de
6 http://epp.eurostat.ec.europa.eu
7 http://data.worldbank.org
8 http://online.thomsonreuters.com/datastream/
9 http://de.statista.com
10 http://www.gesis.org/en/
11 http://thedatahub.org
12 http://dvn.iq.harvard.edu/dvn/dv/NEEO
13 http://www.nesstar.com
14 http://www.icpsr.umich.edu
and in association with individual indicators and records rather than a separate
metadata block for a zipped data bundle. This enables the system to make sense
of the data in depth and allow for requirement speci cation as explained later
in Section 5.</p>
      <p>
        The research on a data retrieval interface is part of our overall research
activities on an infrastructure for scienti c data for the eld of economics. For
several reasons we regard Semantic Web technologies most suitable for this
purpose, among which is strength in dealing with distributed data and
extensibility, which is required whenever highly speci c long tail data from individual
researchers needs additional vocabulary for description [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. However, the data
format should provide for typical data types, such as oats, strings, dates and
other. It must provide metadata on ne-grained level as to open up possibilities
for retrieval and composition. As a consequence, the retrieval system operates
on statistical data in the format of the RDF Data Cube Vocabulary15 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        The prototype was implemented in Java and JavaScript under the use of the
Play Framework16. The live system was tested on an Apache Tomcat17 and a
Sesame Triple Store18, as the system operates on statistical data provided as
RDF using the RDF Data Cube Vocabulary19 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>User Interface Design</title>
      <p>The system implements a multi-step retrieval interface as described in Section
2. In the following, we are going to refer to the screenshots given in Figure 1
to 8 in parantheses. Since the expected result is a data table after all, the main
screen starts with an empty spreadsheet (1). For Step 1, the user successively
enters the names of the concepts that are to be compared in the empty column
headers as shown in (2). This task is supported by autocompletion on the basis
of concept terms contained in a thesaurus, STW20 in our case. With the selection
of a concept, the system displays the number of associated data sets beneath
the concept label entered before. A click on this number lists all of them in
alphanumerical order (3), and another click reveals a detailed description and
further information on the particular data set (7). Yet, at this point, the number
of data sets might be huge, and the user may decide to formulate requirements
for the data rst as per Step 2. With the selection of a single column header, the
panel on the left lists down the union over all properties and property values
available in the metadata of all the data sets associated with the concept of the
column (4). Hovering over a property or property value produces an info box with
documentation on the vocabulary. Selecting a particular property value speci es
a requirement and tells the system that only those data sets are relevant for
15 http://www.w3.org/TR/vocab-data-cube/
16 http://www.playframework.org
17 http://tomcat.apache.org
18 http://www.aduna-software.com/technology/sesame
19 http://www.w3.org/TR/vocab-data-cube/
20 STW Thesaurus for Economics, http://zbw.eu/stw/
this column that provide this respective property and property value, and the
number of relevant data sets drops. With the selection of two or more column
headers, the panel on the left shows the intersection between the properties and
values of the single columns (5). This feature facilitates harmonization of data,
as it reveals which data characteristics can be uni ed among the columns. To
specify the contents of the rows, one must specify the Dimension property. A
click on the respective header highlights all column headers of the entire table as
to indicate that the property of choice must be available in the data sets of all
columns. The user selects (multiple) values from the properties listed on the left
and the Dimension column lls accordingly (6). This again sets requirements for
the data sets, as it lters all data sets that do not provide respective records.
Eventually, when all requirements are set, the user examines and selects from the
remaining list of data sets for each column (7). If all remaining properties with
multiple options are bound to a value, the table lls with actual data content
(8). As a last step, the table is o ered for download.
Following the call for a research data infrastructure, we have addressed the issue
of data retrieval for the domain of economics and social sciences where large
amounts of scienti c results are based on statistical data. With the prospect of a
rapidly growing amount of data from individual researchers and institutes led
in the future, overviewing all relevant data sets e ciently becomes a problem.
For this purpose, we have designed an innovative retrieval interface that aims
to support researchers in nding and composing data sets according to their
natural way of approaching a research question. The prototype presented in
this paper provides simple means for data harmonization to enable consistency
within statistical population in intuitive ways. Under the use of these features,
we expect a signi cant decrease of time needed for data search and composition
in comparison to the current practice, although this is yet to be evaluated.</p>
      <p>Future improvements of the system should include retrieval from distributed
sources, as this version operates on a single triple store endpoint only Moreover,
the advantages of using subproperty relations should be investigated and made
available to the user. Many other valuable ideas for improvements can be found
with regard to user assistance, e.g. warning noti cations when selected time
series data include breaks, errors or changes in acquisition method which can be
derived from well-maintained metadata.</p>
      <p>Finally, this approach needs to be tested on a large archive of various kinds of
statistical data and evaluated with end users from the target group of empirical
researchers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gray</surname>
          </string-name>
          , J.:
          <article-title>Jim Gray on eScience: A Transformed Scienti c Method (January</article-title>
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Treloar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harboe-Ree</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Data management and the curation continuum: how the Monash experience is informing repository relationships</article-title>
          .
          <source>Proceedings of VALA</source>
          <year>2008</year>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ru</surname>
          </string-name>
          <article-title>mpel, S.: Data Librarianship : Anforderungen an Bibliothekare im Forschungsdatenmanagement (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Vlaeminck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegert</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Welche rolle spielen forschungsdaten eigentlich fur fachzeitschriften? eine analyse mit fokus auf die wirtschaftswissenschaften</article-title>
          .
          <source>Technical report, German Council for Social and Economic Data (RatSWD)</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andersson</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bachem</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Best</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Genova</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          , Los, W.,
          <string-name>
            <surname>Marinucci</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romary</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Van de Sompel, H.,
          <string-name>
            <surname>Vigen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giaretta</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Riding the wave: How Europe can gain from the rising tide of scienti c data</article-title>
          .
          <source>European Union</source>
          (
          <year>2010</year>
          )
          <article-title>Final report of the High Level Expert Group on Scienti c Data: A submission to the European Commission</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Feijen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>What researchers want - a literature study of researchers' requirements with respect to storage and access to research data</article-title>
          (
          <year>February 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Kampgen,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Transforming statistical linked data for use in olap systems</article-title>
          .
          <source>In: Proceedings of the 7th international conference on Semantic systems, ACM</source>
          (
          <year>2011</year>
          )
          <volume>33</volume>
          {
          <fpage>40</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Boland</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ritze</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eckert</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mathiak</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Identifying references to datasets in publications</article-title>
          . In Zaphiris,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Buchanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Rasmussen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Loizides</surname>
          </string-name>
          , F., eds.
          <source>: Theory and Practice of Digital Libraries. Volume 7489 of Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg (
          <year>2012</year>
          )
          <volume>150</volume>
          {
          <fpage>161</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bahls</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tochtermann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Addressing the long tail in empirical research data management</article-title>
          .
          <source>In: Proceedings of the 12th International Conference on Knowledge Management and Knowledge Technologies. i-KNOW '12</source>
          , New York, NY, USA, ACM (
          <year>2012</year>
          )
          <volume>19</volume>
          :
          <fpage>1</fpage>
          {
          <issue>19</issue>
          :
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Field</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gregory</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halb</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tennison</surname>
          </string-name>
          , J.: Semantic statistics:
          <article-title>Bringing together sdmx and scovo</article-title>
          . In Bizer,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Heath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Hausenblas</surname>
          </string-name>
          , M., eds.
          <source>: LDOW</source>
          . Volume
          <volume>628</volume>
          of CEUR Workshop Proceedings., CEUR-WS.org (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>