<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using the Institutional Repository to publish research data</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Southampton</institution>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>8</lpage>
      <abstract>
        <p>For open research data to be fully utilised it must be discoverable. Many types of research dataset are impossible to identify by looking at them so metadata is essential. This is the only major issue with using existing Institutional Repositories to preserve and disseminate data. This paper suggests a simple scheme for facilitating discovery and reuse of open scienti c data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Introduction</p>
      <p>\The greatest crisis facing us is not Russia, not the Atom Bomb,
not corruption in government, not encroaching hunger, nor the morals
of the young. It is a crisis in the organization and accessibility of human
knowledge. We own an enormous \encyclopedia" - which isn't even
arranged alphabetically. Our \ le cards" are spilled on the oor, nor were
they ever in order. The answers we want may be buried somewhere in
the heap, but it might take a lifetime to locate two already known facts,
place them side by side and derive a third fact, the one we urgently
need." - Robert Heinlein, 1950</p>
      <p>It is already possible to use an Institutional Repository (IR) to store and
disseminate datasets, but not yet common practice. This data will be most valuable
when similar datasets from around the world are aggregated and all scientists
can have access to all available data.</p>
      <p>Best practice will require Open formats, Open licenses, provenance and
discoverability. Google, and other search engines, may have solved many of the
problems of nding text, but raw data is a stickier challenge, as it may be
nothing more than a grid of numbers or other arcane formats. This paper suggests a
simple mechanism to make all open datasets discoverable and recommends that
data is stored in Institutional Repositories to provide reliable long term curation
and availability.
Most data created in research is not yet available online. The infrastruture to
enable this already exists in the form of Institutional Repositories. Making raw
data available online allows it to be reused, and also supports the scienti c
process by allowing peers to repeat the analysis of the data and verify conclusions
in a paper. It is impractical for an IR manager to do more than curate the data.
Individual research communities will form their own practices, tagging data in
their local IR to allow it to be discovered by other members of their community.</p>
      <p>There is limited space in a print medium. This makes it impractical to include
many pages of data which will be of interest to very few readers. Much research
is created, reviewed and consumed without ever entering hard-copy, and digital
media requires fewer physical restrictions.</p>
      <p>
        15 years ago, Stevan Harnad made his Subversive Proposal[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that the
scholarly community should be sharing its research online, without barriers, plus a
sketch of how to get there. This is now well under way and at the time of
writing, 63% of publishers now permit some form of a paper to be made available
online.[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
      </p>
      <p>There are many issues with the communication of research data including
collection, provenance, curation, interoperation and dissemination.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Immediate Solution</title>
      <p>
        There are nearly 1000 insitutional-style repositories in the world[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and that
number is increasing. Rather than design complex new systems, the remit of
IRs should be expanded to include research data along with research outputs.
Nothing is required but a change in repository policy, and the addition of a new
option \dataset", plus encouraging researchers to deposit.
      </p>
      <p>Many repositories already support a record type of \other" and using this is
better than nothing, but the data will be di cult to discover.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Making datasets discoverable</title>
      <p>A repository should allow the depositor to identify the content of a dataset. This
needs to be painless, but must not be from a limited set curated by the library
as that will be a barrier to the evolution of scienti c communities. There is no
reason that a repository cannot o er auto-completion of the eld from a short
list, so long as it does not preclude free text input.</p>
      <p>The ideal solution would be to identify the subject and format of the
contents of the dataset by one or more URIs, but URIs are cumbersome. As a
more practical solution, dataset contents can be identi ed by either a URI or
a short text string which will be treated as part of a data format namespace
http://ds.eprints.org/ns/, giving a URI without requiring the depositing
scientist to remember the whole thing. These identi ers can signify either the
content of the dataset (e.g. a analysis of a crystal structure), the format (e.g.
chemical .cif or .cml le), other properties such as strictness of protocols used, or
any combination of these. Generally, an identi er will indicate both format and
content, but this will need to be established by various communities as their
requirements will vary widely. Any dataset can be identi ed by multiple identi ers
as appropriate.</p>
      <p>These identi ers should be included in any electronic dissemination of the
metadata of the repository. Including via the OAI-PMH protocol using dc:subject,
in any RDF or RDFa relating to the record.
http://ds.eprints.org/ns/chem some kind of chemical data
http://ds.eprints.org/ns/chem-cml a chemical in CML format
http://ds.eprints.org/ns/chem-cml-cry a crystal in CML format
http://ds.eprints.org/ns/chem-cml-cry-org an organic crystal in CML format</p>
      <p>This system is deliberately a very loose semantic relationship. A single
identi er can indicate any or all of type, subject and format. To keep the system
managable for each community, as few identi ers as possible should be used. In
the above example the `chem-cml-cry-org' is almost certainly an unhelpful level
of detail, and just `chem' is no more useful than a Library of Congress subject.
A good balance should be to use identi ers which indicate scienti c area and the
format of the data at a level of detail useful to aggregation tools. If the dataset is
in a machine processable format, such as CML, then the identi er should re ect
this to allow aggregation tools to discover and process these datasets. Later, if
needed, semantic information can be returned from ds.eprints.org, naming
established identi ers and indicating that they imply that a dataset has certain
subjects, types and formats.</p>
      <p>Ideally all data should be made available in Open formats, with Open licenses.
Formats and licensing are essential for providing aggregation services and using
the data in future work.</p>
      <p>These changes can and should be made to the standard release of repository
tools such as EPrints, D-Space and Fedora.
4.1</p>
      <p>
        Very Large Datasets
Some datasets may be much too large to make available via HTTP, due to
both expense of bandwidth, and it being impractical to download. Anything
larger than a few terabytes cannot yet be usefully made available via the web
in raw form, but that threshold will increase in time. The existence of large
datasets can still be described in the IR, even if the URI identifying the dataset
is not resolvable. In this case the record should also contain human-readable
information on how to gain access to the dataset.
If the raw dataset just does not contain enough information to be useful, then
the communities need to establish better formats. An interim solution is to make
the URL of the dataset return a manifest le in XML, RDF or similar which
contains the additional data to make the dataset useful.
Some datasets may contain multiple les. In this case, as with metadata being
required, a manifest le can contain the URLs of the other les in the dataset.
If the les have xed names, then a less robust solution would be to load the
other les based on relative URL paths. OAI-ORE[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is suitable for the purpose
and already supported by some repositories.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Policy</title>
      <p>While some researchers are keen to publish their research papers online, many
require a mandate before they will do so as it is one extra task for busy people.
To ensure the majority of scienti c data is made available will require mandates
from funding bodies, accreditation exercises or institutions. When they emerge
they should require that the data must not only be Open, but discoverable.
Without clues to the content, search tools and harvesters may not be able to tell
what many datasets are about.</p>
      <p>Initially much data will not be in Open formats, but communities will
discover bene ts in standardising formats, but only if there is a practical way to
aggregate the data.</p>
      <p>For the rst few years, it is very likely that many researchers will resist
putting data online as they do not feel it is of a standard to publish, although
they based published papers on it. Once it becomes part of the expected process,
then this issue will diminish. To ease the process, researchers should be given
credit for the quality and impact of their raw data, as well as their papers.
Research papers should reference datasets used, both in the text and in electronic
metadata.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Aggregators</title>
      <p>With appropriately discoverable datasets available, the next step will to build
aggregators which can add value to speci c types of dataset. It is impractical for an
IR to provide subject-aware visualisations and search tools, as there will be many
and they will evolve. A more practical approach is to make the datasets
available and discoverable with licences which make it possible for subject-speci c
web sites to provide these services over all datasets of a given type, from all the
IRs in the world.</p>
    </sec>
    <sec id="sec-6">
      <title>Using the IR as storage for subject-speci c or experimental services</title>
      <p>While the Institutional Repository can provide a high degree of security in the
continuity of URLs and preservation, there may be reasons for researchers to
want to deposit their data in other systems. If these are well supported and stable
then this is not a problem, however if these are more experimental then this is
a concern. It is likely that many research projects will set up data repositories
with no clear plan for continuity past the end of these projects.</p>
      <p>A better solution, in many cases, is to use an institutional repository as a
back-end to store the data. The experimental tool can both deposit items in the
IR and then retrieve the data to provide subject-speci c features or analysis.
Where possible, it should disseminate the URI/URL of the raw data in the IR
to proof against the risk of the experimental service going o ine once funding
ends.</p>
      <p>
        A subject speci c tool may act as both a tool for ingest (see g. 3), and
as a way to add value to speci c types of dataset, in the same manner as an
aggregation tool. The SWORD protocol[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is ideal for this purpose. Such subject
speci c tools may well be commercial, or even supplied as an integrated solution
with the next generation of laboratory devices.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Alternatives to URI tagging raw datasets</title>
      <p>The advantage of the author tagging the dataset is that it gives a degree of
trust, as the information is collected and disseminated by an IR and that means
the institution has a vested interest in ensuring that its data is correct and
asdescribed. A downside is that if the community evolves new ways to identify its
data, it is unlikely that anyone will update these tags.</p>
      <p>
        An alternative or complimentary solution would be to use a social tagging
system, such as Delicious[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to tag the URLs or even URIs of datasets, either by
a community e ort (crowd-sourcing) or by a small expert group. However
crowdsourced information may be less reliable and in small communities is liable to
noise. On the other hand, a single authority creating a global list of datasets for
a given format is a single point of failure.
      </p>
      <p>Social Bookmarking or authorities maintaining lists of known datasets are
complementary to using subject tags or URIs to identify datasets. It is certain
that di erent solutions will work for di erent communities.
9</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>Research funders should mandate that raw data produced as a result of their
funding should be made available in Open formats, with Open licenses and made
suitably discoverable at URLs which will be stable for many years. For this to
be possible the researchers must have a sutiable repository for their data, and
research communities will need to decide what level of detail is useful to facilitate
discovery.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Harnad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <string-name>
            <given-names>A Subversive</given-names>
            <surname>Proposal</surname>
          </string-name>
          . Association of Research Libraries. (
          <year>1995</year>
          ) http://www.arl.org/sc/subversive/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <article-title>Registry of Open Access Repositories (ROAR)</article-title>
          . http://roar.eprints.org/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. SHERPA/RoMEO - Publisher copyright policies &amp;
          <article-title>self-archiving</article-title>
          . http://www.sherpa.ac.uk/romeo/statistics.php
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. SWORD protocol. http://www.swordapp.org/</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Delicious</surname>
          </string-name>
          <article-title>- social bookmarking tool</article-title>
          . http://delicious.com/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Open</given-names>
            <surname>Archives Initiative</surname>
          </string-name>
          <article-title>Object Reuse and Exchange (OAI-ORE)</article-title>
          . http://www.openarchives.org/ore/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>