<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Repolytics: Identifying Measurable Insights for Digital Repositories</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vangelis Nomikos</string-name>
          <email>nomikos@intelibility.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelibility Innovative Data Engineering LLC</institution>
          ,
          <addr-line>Atlanta, GA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents specific aspects of the Repolytics platform: a data analytics platform for digital repositories. An overview of the platform is presented alongside an example of how one can employ it's services to analyze a digital repository's data and identify quality related issues. Data is captured, analyzed and used to drive all aspects of our lives in a data driven world. Recently, we have seen a rapid growth of digital repositories and open data catalogues being made available to the public. In the case of digital repositories that target the libraries-archives, scientific-research data domains and open data portals the market is dominated by open source solutions such as DSpace, Omeka, DKAN, CKAN due to their simplicity and low-no cost. Most solutions enable ingest / cataloguing of information either through automated means (SWORD, REST APIs) or through simple and untuitive forms. Quality assurance in most cases comprises of a number of mandatory metadata the user has to enter. The reality however is far more complex and this has a profound effect in the quality of the the ingested data. Repolytics [1] is a platform that aims at filling this gap through intelligent data analytics. The data loaded into the platform are analyzed and specific quality metrics are presented alongside a more thorough analysis per metadata element. The metrics include metadata completeness, accuracy and consistency. Furthermore, similar metrics are calculated for the data as well.</p>
      </abstract>
      <kwd-group>
        <kwd>metadata quality</kwd>
        <kwd>data quality</kwd>
        <kwd>data analytics</kwd>
        <kwd>quality metrics</kwd>
        <kwd>digital repositories</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Metadata quality is an important issue for the Digital Library domain and has
attracted several researcher groups to deal with it. One of the main demands is the
establishment of a conceptual framework consisted of a set of well-defined quality
assessment criteria such as completeness, validity, consistency, timeliness, appropriateness
and accuracy constituents [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. In such a framework any assessment effort would be
based on reliable indications about metadata quality. The first attempt to define a
framework established a narrow set of criteria such as accuracy, completeness and
serviceability [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Some researches expanded the criteria set from which metadata
quality is approached [
        <xref ref-type="bibr" rid="ref6 ref8">6, 8</xref>
        ], while some other narrowed the perspective focusing only
on the completeness criterion [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        An important evolution of this scientific field was the introduction of the
contextdependent metadata quality approach [
        <xref ref-type="bibr" rid="ref3 ref7">3, 7</xref>
        ]. According to this approach metadata
quality issues follow four major concepts: mappings, changes to the information
entity, changes to the underlying entity and context changes. For these concepts a
taxonomy of 22 information quality criteria was developed. The criteria were clustered to
three categories: intrinsic, relation and reputational and are measured via 41 metrics
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Recent research suggests that a metadata quality framework doesn’t have to
“invent new dimensions in order to accommodate the needs of diverse communities of
practice” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] but to give the flexibility to each evaluator to assess the results within a
specific context. This paper follows the context-dependent approach and assumes that
metadata quality strongly depends on the viewpoint of the evaluator and should be
aligned with the application domain for which the metadata were produced and used.
Therefore weighting functions for these factors (evaluator viewpoint, application
domain, metadata usage) should be defined and used to weight the values of the
metadata quality metrics.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Repolytics approach</title>
      <p>The core principle of the Repolytics platform is the provisioning of an expandable set
of middleware services that operates both on metadata and data level of a digital
repository. The most fundamental services of the platform involve:
 (meta)-data integration services
 (meta)-data profiling services
 (meta)-data quality services
The Repolytics platform enables the user to load data from different data sources
either directly (e.g. through a file archive) or by using one of the supported data
providers such as DKAN API and OAI-PMH protocol. Each data source provides one or
more metadata (e.g. OAI-DC, MODS) and data (e.g. CSV, excel) representations.
Once the data has been loaded, each digital object is analyzed and the main workflow
seen in Figure 1 below is executed.</p>
      <p>For each digital object, if metadata is provided, the format is identified, every
metadata element is indexed and a set of core analyzers are invoked for each element. For
specific elements depending on their type, a set of custom analyzers is invoked. If the
data is provided (usually throuh some kind of URL) the actual data is fetched and
verified followed by a profile of each bitstream. Again, a set of custom analyzers is
then invoked depending on the file type.</p>
      <p>Some of the metrics that are calculated include:
 metadata completeness
 distinct values for each metadata element
 accuracy for specific element types (e.g. dates, actors) where all values are
classified according to their class
 itemset frequencies when applicable (e.g. in the case of subject terms and
keywords)
One of the primary challenges include the efficeint visualization of the results to the
end user. For that reason a series of bar charts, radar charts, tables, gauge meters etc
are employed per case. For example, as shown in Figure 2 a radar chart is employed
to fingerprint an entire repository according to it's completeness.</p>
      <p>Similary, an accuracy detector identifies and classifies all metadata element values
according to their class and provides an insight on the accuracy level of each class
(low, medium, high).</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions &amp; Future Work</title>
      <p>In this paper, Repolytics, a platform for repository data analytics is presented. The
platform enables the use to easily load data from supported data sources and analyzes
this data focusing primarily on completeness and accuracy whereas in specific cases,
more specific metrics (such as itemset frequencies) are employed to help gain insight
on highly subjective metrics such as consistency. The platform can also access and
profile data as well.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Repolytics</given-names>
            <surname>Platform</surname>
          </string-name>
          , http://www.repolytics.com
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Herzog</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scheuren</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winkler</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Data Quality and Record Linkage Techniques</article-title>
          . Springer-Verlag, NY (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Tani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Candela</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Dealing with metadata quality: The legacy of digital library efforts</article-title>
          .
          <source>Inf. Process. Manag</source>
          .
          <volume>49</volume>
          (
          <year>2013</year>
          )
          <fpage>1194</fpage>
          -
          <lpage>1205</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Moen</surname>
            ,
            <given-names>W.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McClure</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          :
          <article-title>The Role of Content Analysis in Evaluating Metadata for the U.S. Government Information Locator Service (GILS): Results from an Exploratory Study</article-title>
          . http://digital.library.unt.edu/ark:/67531/metadc36312/citation/ (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Margaritopoulos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Margaritopoulos</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mavridis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manitsaris</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Quantifying and measuring metadata completeness</article-title>
          .
          <source>J. Am. Soc. Inf. Sci. Tech</source>
          .
          <volume>63</volume>
          (
          <year>2012</year>
          )
          <fpage>724</fpage>
          -
          <lpage>737</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ochoa</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duval</surname>
          </string-name>
          ,E.:
          <article-title>Quality Metrics for Learning Object Metadata</article-title>
          .
          <source>In: WorldConference on Educational Multimedia, Hypermedia and Telecommunications</source>
          . (
          <year>2006</year>
          )
          <fpage>1004</fpage>
          -
          <lpage>1011</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Stvilia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Twidale</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>L.C.</given-names>
          </string-name>
          :
          <article-title>A framework for information quality assessment</article-title>
          .
          <source>J. Am. Soc. Inf. Sci. Tech</source>
          .
          <volume>58</volume>
          (
          <year>2007</year>
          )
          <fpage>1720</fpage>
          -
          <lpage>1733</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.-N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wen</surname>
          </string-name>
          ,C.-Y.,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.-P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sum</surname>
          </string-name>
          ,H.-C.:
          <article-title>Metrics for metadata quality assurance and their implications for digital libraries</article-title>
          . In: Xing,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Crestani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            , and
            <surname>Rauber</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.) 13th
          <source>International Conference on Asia-Pacific Digital Libraries</source>
          , Beijing, China. Springer-Verlag, Berlin (
          <year>2011</year>
          )
          <fpage>138</fpage>
          -
          <lpage>147</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>