<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Quality Assessment in Europeana: Metrics for Multilinguality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valentine Charles</string-name>
          <email>valentine.charles@europeana.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juliane Stiller</string-name>
          <email>juliane.stiller@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Kiraly</string-name>
          <email>peter.kiraly@gwdg.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Werner Bailer</string-name>
          <email>werner.bailer@joanneum.at</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nuno Freire</string-name>
          <email>nuno.freire@tecnico.ulisboa.pt</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Europeana Foundation</institution>
          ,
          <addr-line>The Hague, NL</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gesellschaft fur wissenschaftliche Datenverarbeitung mbH Gottingen</institution>
          ,
          <addr-line>DE</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Humboldt-Universitat zu Berlin</institution>
          ,
          <addr-line>DE</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>INESC-ID</institution>
          ,
          <addr-line>Lisbon, PT</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Joanneum Research</institution>
          ,
          <addr-line>Graz, AT</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Europeana.eu aggregates metadata describing more than 55 million cultural heritage objects from libraries, museums, archives and audiovisual archives across Europe. Quality (and particularly multilingual quality) of metadata is crucial for enabling search and data re-use in digital libraries. Capturing multilingual aspects of the data requires us to take into account the full lifecycle of data aggregation including data enhancement processes such as data enrichment. Multilinguality cannot be captured as one measure, but needs to be considered as an intersection of several measures, bringing challenges when interpreting and visualising the results. The paper presents an approach for capturing multilinguality as part of data quality dimensions, namely completeness, consistency and accessibility. We describe the measures de ned and implemented, and provide initial interpretations of the results.</p>
      </abstract>
      <kwd-group>
        <kwd>metadata quality</kwd>
        <kwd>multilinguality</kwd>
        <kwd>digital cultural heritage</kwd>
        <kwd>Europeana</kwd>
        <kwd>data quality dimensions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Europeana.eu6 is Europe's digital platform for cultural heritage. It aggregates
metadata describing more than 55 million cultural heritage objects from a wide
variety of institutions (libraries, museums, archives and audiovisual archives)
across Europe.</p>
      <p>The need for high-quality metadata is particularly motivated by its impact on
search, the overall Europeana user experience, and data re-use in other contexts
such as the creative industries, education and research. One of the key goals
of Europeana is to enable users to nd the cultural heritage objects that are
relevant to their information needs irrespective of their national or institutional
origin and the material's metadata language.</p>
    </sec>
    <sec id="sec-2">
      <title>6 http://www.europeana.eu/portal/en</title>
      <p>
        As highlighted in the White Paper on Best Practices for Multilingual Access
to Digital Libraries [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], most digital cultural heritage objects do not have a
speci c language and can only be searched through their metadata, which is
text in a particular language. The presence of multilingual metadata description
is therefore essential to improving the retrieval of these objects across language
spaces. Understanding Europeana's cross-lingual reach is essential to enhancing
the quality of metadata in various languages.
      </p>
      <p>
        In the Data Quality Committee (DQC), a body formed at Europeana's
initiative, we specify functional requirements that de ne the purpose of the metadata
and guide data-quality evaluation { following a core principle for metadata
assessment (e.g. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). Europeana's international scope means that multilinguality is
an inherent aspect of these requirements. To assess multilinguality in metadata,
we look at commonly-agreed data quality dimensions: completeness, consistency
and accessibility, and de ne and implement quality measures. We capture data
quality along the full data-aggregation lifecycle, taking also into account the
impact of data enhancement processes such as semantic enrichment. The model
the data is represented in, namely the Europeana Data Model (EDM)7, is also
a key element of our work.
      </p>
      <p>In the next section, we present data quality frameworks, dimensions and
criteria that are commonly referred to in the context of data quality measurement.
Section 3 describes how multilingual data is presented in Europeana's data model
and the functional requirements that are the basis of our assessment. The data
quality dimensions we use will also be presented. In section 4, we describe the
implementation of the di erent measures. Section 5 describes rst results and
measures that were taken to improve data along the di erent quality dimensions.
We conclude this paper with an outline of future work.
2</p>
      <sec id="sec-2-1">
        <title>State of the art</title>
        <p>
          Addressing data quality requires the identi cation of the data features that need
to be improved and this is closely linked to the purpose the metadata is serving.
Libraries have always highlighted that bibliographic metadata enables users to
nd material, to identify an item and to select and obtain an entity [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Based on
this, Park [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] expands functional requirements of bibliographic data to
discovery, use, provenance, currency, authentication and administration, and related
quality dimensions. The approach to metadata assessment for cultural heritage
repositories presented in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] also starts from use for a speci c purpose. While the
work mentions the issue of multilinguality, it does not propose speci c metrics
to measure it.
        </p>
        <p>
          Bruce and Hillmann [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] de ne the following measures for quality:
completeness, accuracy, provenance, conformance, logical consistency and coherence,
timeliness and accessibility, grouped in three tiers. Shreeves et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] focus on three
quality dimensions (completeness, consistency and ambiguity) for their
explo
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>7 http://pro.europeana.eu/edm-documentation</title>
      <p>rative study of four collections, which included metadata from several
institutions. Not surprisingly, they found greater variances for collections aggregated
from multiple sources as well as inconsistencies in the encoding scheme { a result
also experienced by Europeana.</p>
      <p>
        Designing an information quality framework, Stvilia et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] have
proposed a taxonomy of 22 measures, grouped into three categories: intrinsic,
relational/contextual and reputational information quality. In this taxonomy,
completeness/precision exists both in the intrinsic (e.g. the number of elements, the
non-empty elements) and in the relational/contextual categories (completeness
wrt. a recommended set of elements). Although multilinguality is not explicitly
mentioned in this paper, it ts in this model as part of contextual completeness.
      </p>
      <p>
        Zaveri et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] propose dimensions for quality assessment for linked data,
grouped into the accessibility, intrinsic, contextual and representational
categories. Completeness is listed as an intrinsic criterion, while multilinguality
is covered by versatility, which is considered a representational criterion. The
ISO/IEC 25012 standard [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] de nes a data quality model with 15 characteristics,
discriminating between inherent and system-dependent ones, but putting many
of the criteria in the overlap between the two classes. Completeness is de ned as
an inherent criterion, while accessibility and compliance are in the overlapping
area. Multilinguality could be seen as being compliant to providing a certain
number of elements in a certain number of languages, and as enabling access to
users who are able to search and understand results in certain languages.
      </p>
      <p>
        Mader et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] include issues related to multilinguality, such as
incompleteness of language coverage and missing language tags, in their analysis of SKOS
vocabularies. Multilinguality is covered by Droge [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] as part of the eight classes
of criteria de ned for assessing the quality of vocabularies. However, no metric
is developed in detail.
      </p>
      <p>
        It becomes apparent that while multilinguality is considered in some works,
it is usually not treated as a separate quality dimension, but rather as part of
other criteria or dimensions existing at quite di erent levels in di erent quality
models. To the best of the authors' knowledge, the only resource which measured
multilingual features in metadata is [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], cited by [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]; here language
attributions across a collection were counted. A rst iteration of a multilingual
saturation score that counted language tags across metadata elds in the Europeana
collections as well as the existence of links to multilingual vocabulary was
introduced by Stiller and Kiraly [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The further development of this metric as well
as the integration of measures for all aspects of multilingual data is described in
this paper.
3
      </p>
      <sec id="sec-3-1">
        <title>Approach</title>
        <p>Firstly, to determine the multilingual degree of metadata across several quality
dimensions we have to understand the di erent ways multilingual information is
expressed in Europeana's data model. Secondly, multilingual information is at
the core of the functional requirements that constitute the services Europeana
is providing for an international audience. Thirdly, these requirements and the
structure of multilingual data inform the criteria and metrics that enable us to
measure multilinguality across several metadata quality dimensions.
3.1</p>
        <sec id="sec-3-1-1">
          <title>Multilingual information in Europeana's metadata</title>
          <p>Multilinguality in Europeana's metadata has two perspectives: concerning the
language of the object itself, and the language of the metadata that describes this
object. First, the described cultural object, insofar as it is textual, audiovisual
or in any other way a linguistic artefact, has a language. The data providers
are urged to indicate the language of the object in the dc:language eld in the
Europeana Data Model (EDM) in this way: &lt;dc:language&gt;de&lt;/dc:language&gt;.
This information is then used to populate a language facet allowing users to
lter result-sets by language of objects. The language information is essential for
users who want to use objects in their preferred language. Second, the language
of metadata is essential for retrieving items and determining their relevance.
Metadata descriptions are textual and therefore have a language. Each value in
the metadata elds can be provided with a language tag (or language attribute).
Ideally, the language is known and indicated by this tag for every literal in each
eld. If several language tags in di erent languages exist, the multilingual value
can be considered to be higher. For instance, consider as an example this data
provided by an institution:
&lt;#example&gt; a ore:Proxy ; # data from provider
dc:subject \Ballet", # literal
dc:subject \Opera"@en # literal with language tag</p>
          <p>The rst dc:subject statement is without language information, whereas the
second tells us that the literal is in English. Europeana assesses data in particular
elds to enrich it automatically with controlled and multilingual vocabularies as
de ned in the Europeana Semantic Enrichment Framework.8 As shown in the
following example, the deferencing of the link (i.e., retrieving all the multilingual
data attached to concepts de ned in a linked data service) allows Europeana to
add the language variants for this particular keyword to its search index.
&lt;#example&gt; a ore:Proxy ; edm:europeanaProxy true ;
# enrichment by Euroepana with multilingual vocabulary
dc:subject &lt;http://data.europeana.eu/concept/base/264&gt;
&lt;http://data.europeana.eu/concept/base/264&gt; a skos:Concept .</p>
          <p># language variants are added to index
skos:prefLabel "Ballett"@no, "Ballett"@de, "Bale"@pt,</p>
          <p>"Baletas"@lt, "Balet"@hr, "Balets"@lv</p>
          <p>The record now has more multilingual information than at the time of
ingestion into Europeana. The di erent language versions from multilingual
vocabularies are likely to be translation variants.
8 https://docs.google.com/document/d/1JvjrWMTpMIH7WnuieNqcT0zpJAXUPo6x4uMBj1pEx0Y/
edit
3.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Functional requirements for multilingual services</title>
          <p>In order to better understand the impact of multilinguality on the overall quality
of the data, we have identi ed the multilingual aspects of several functional
requirements relevant to Europeana. Note that these requirements are also crucial
to better communicating the results of our measures to our data providers, and
hence to motivating improvement of the data.</p>
          <p>Cross-language recall: When a user wants to search for a document or
documents relevant to her needs, and to do so irrespective of the language of the
documents and/or their metadata, multilingual metadata is a necessity. It is
therefore required that all searchable freetext metadata elements9 be
consistently tagged with their language.10
Language based facets: In a language-based facet requirement a clear
distinction needs to be made regarding whether the language refers to the language of
the metadata or the digital representation of the work. The method of measuring
the multilingual saturation is di erent in both cases. While the language of the
metadata will be measured against the presence of language tags, the language
of the content will be captured in the dc:language eld in EDM. The existence
of multilingual metadata records in Europeana demands that language-tagging
occur at the level of the metadata element rather than the record as a whole.
Entity-based facets (people, places, concepts, periods): while users will be
primarily searching for entities in their own language (the city of \Paris"), they
will need to access results in di erent languages (Parijs, Parigi, ...). To support
this scenario we will rely on the translated labels provided by data providers
as part of their data or on multilingual vocabularies to which the data can be
linked. In both cases measuring multilinguality is crucial to identify language
gaps.</p>
          <p>It is also important to determine the impact of work ows contributing to the
increase of multilingual labels (such as machine learning and natural language
processing techniques for language detection, automatic tagging, or semantic
enrichment) in the data in order to de ne the measures but also to interpret the
results.
3.3</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Multilinguality as a facet of quality dimensions</title>
          <p>For measuring multilinguality, we look at three quality dimensions:
completeness, consistency and accessibility. Each of these dimensions assesses
multilinguality from a di erent perspective. The results re ect how well the functional
requirements work.
9 https://docs.google.com/spreadsheets/d/1f6vbeo4mt0stl0yAbVCyp_
-5bzBBCq4Dy6p7xhMcjSI/
10 for instance, in accordance with the ISO 639-1 or 632-2 (https://www.loc.
gov/standards/iso639-2/php/code_list.php) list, the IANA language tag
registry (http://www.iana.org/assignments/language-subtag-registry/
language-subtag-registry), or the Languages Name Authority List (NAL)
(https://open-data.europa.eu/en/data/dataset/language)
Completeness. Completeness is a basic quality measure, expressing the
number (fraction) of elds present in a dataset, and identifying non-empty values in
a record or (sub-)collection. For a xed set of elds completeness is thus
straightforward to measure, and can be expressed as the absolute number or fraction
of the elds present and not empty. However, the measure becomes non-trivial
when data is represented using a data model with optional elds (that may e.g.,
only be applicable for certain types of objects), or with certain elds for which
the cardinality is unlimited (e.g., allowing zero to many subjects or keywords).
These characteristics apply to EDM. In such cases the measure becomes
unbounded, and a few elds with high cardinality may outweigh or swamp other
elds.</p>
          <p>In the context of measuring multilingual completeness, the metric is two-fold.
First, the concept of completeness can be applied to measuring the presence of
elds with language tags or, second, the presence of the dc:language eld. Any
measure of multilinguality must be seen in relation to the results of measuring
completeness. Only elds both present and non-empty can be said to have or
lack language tags and translations. A record which is 80% complete can still
reach 100% multilingual completeness if all present and non-empty elds have
a language tag. The completeness a ects all three functional requirements as
they all work better the more multilingual information exists in a given set of
elds. Consistency. Consistency describes the formal conformity to set
standards and eld rules as well as the logical coherence of the metadata. With
regard to multilinguality, the dimension assesses the variety of language values
in the dc:language eld. The goal is to de ne a standard for language notations
and normalize the eld in this regard (see section 5.2). The consistency measure
is mainly relevant for the language based facet. The better normalization works,
the better the user can lter objects by their language.</p>
          <p>Accessibility. Accessibility describes the degree to which multilingual
information is present in the data, and allows us to understand how easy or hard it
is for users with di erent language backgrounds to access information. So far,
Europeana has little knowledge about the distribution of linguistic information
in its metadata { especially within single records. To quantify the multilingual
degree of data and measure cross-lingual accessibility, the language tag is
crucial. Resulting metrics can be scaled to the eld, record and collection levels.
In practical terms, the accessibility measure serve to gauge cross-language recall
and entity-based facet performance.</p>
          <p>To summarize: with regard to multilinguality, we identi ed the following
dimensions and quality criteria (Table 1 3.3).
4</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Operationalizing the metrics for multilinguality</title>
        <p>
          Implementation requires a good understanding of EDM in order to not exclude
relevant data. The source data go through several levels of data aggregation
before being aggregated into Europeana. Data processes can take place at each
of theses levels and a ect the nal data output. EDM allows us to distinguish
between values provided by the data provider(s) from the information added
by Europeana (for instance by semantic enrichment), doing so by leveraging the
proxy mechanism from the Object Re-use and Exchange (ORE) model. It enables
the representation of resources in the context of aggregations o ering di erent
views on the same resource [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Any implementation of quality measures, and in
particular of multilingual saturation, needs to take into account this distinction
depending on the results it is intended to yield. For instance, the contribution
to the score might be higher if we only consider the Europeana proxy where a
value was enriched with a multilingual vocabulary (e.g. DBpedia).
        </p>
        <sec id="sec-3-2-1">
          <title>4.1 Implementation of the measures</title>
          <p>The di erent metrics for the assessment of multilinguality in metadata are
implemented in the Data Quality Assurance framework.11 The measurement has
four phases. (1) data collection and preparation: the EDM records were collected
via Europeana's OAI-PMH service12, transformed to JSON where each record
was stored in a separate line, and stored in Hadoop Distributed File System13 (2)
record-level measurement: the Java applications14 measuring di erent features
of the records run as Apache Spark jobs, allowing them to scale readily. The
process generates CSV les which record the results of the measurements such
as the number of eld instances, or complex multilingual metrics. (3) statistical
analysis: the CSV les are analyzed using statistical methods implemented in R
and Scala. The purpose of this phase is to calculate statistical tendencies on the
dataset level and create graphical representations (histograms, boxplots). The
results are stored in JSON and PNG les. (4) user interface: interactive HTML
and SVG representations of the results such as tables, heath maps, and spider
charts. We use PHP, jQuery, d3.js and highchart.js to generate them.
11 http://144.76.218.178/europeana-qa/multilinguality.php?id=all
12 http://labs.europeana.eu/api/oai-pmh-introduction. Our client library:
https://github.com/pkiraly/europeana-oai-pmh-client/.
13 Snapshot created at the end of 2015, consisting of more than 46 million records (392</p>
          <p>GB disk space): http://hdl.handle.net/21.11101/0000-0001-781F-7.
14 Source code and binaries: http://pkiraly.github.io/about/#source-codes.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Interpretating the output of the measures</title>
        <p>The status of these metrics with regard to interpreting their results and thus
deriving strategies for metadata improvement di ers considerably. Whereas the
measures for the quality dimensions of completeness and accessibility need more
re nement, as well as more discussion of appropriate display of the results, the
consistency metric has already served to drive a successful project for normalizing
languages. These various developments and their outcomes are reported in this
section.
5.1</p>
        <sec id="sec-3-3-1">
          <title>Completeness</title>
          <p>The completeness measure indicated that 904 (out of 3548) collections have no
value in the dc:language eld, which shows the eld is missing. On a record level,
58,03% of the records have a dc:language eld.15 Another pattern one can detect
in the data relates to the misuse of elds. For example, the metric \cardinality"
allowed us to identify collections that have metadata elds with more than 3
instances of dc:language. In one example, as many as 153 language values were
found associated with this eld, owing to duplication of the language tag. One
recommendation to avoid this in future would be to check the eld for duplicated
entries and reduce them to a single language tag.
5.2</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Consistency</title>
          <p>Measuring consistency in the dc:language eld shows that values are currently
quite heterogeneous in the Europeana dataset. dc:language values are
predominantly normalized in ISO-639-1 or ISO-639-3, but, in contrast, values
nevertheless sometimes occur in natural language sentences that cannot be processed
automatically. We also nd language ISO codes without their reference to the
ISO standard in use, or references to languages by their name. Additionally,
over 400 di erent language variants are present in the eld. This table presents
some general statistics about the presence of ISO-639 codes in the values of
dc:language in the Europeana dataset:</p>
          <p>Total values in the Europeana dataset 33,070,941
Total values already normalized (ISO-639-1, 2 letter codes) 23,634,661
Total values already normalized (ISO-639-3, three letter codes) 4,831,534
The metric helps us to design further language normalization rules which
in turn can be used to further improve the results of the quality measures.
Any visualisation implementation of the quality measures would also bene t
from such normalisation. We have developed a language-normalization operation
15 http://144.76.218.178/europeana-qa/frequency.php
on the Europeana dataset which itself comprises a range of operations - i.e.,
cleaning, normalization and enrichment of data. The following exemplify some
cases of the di erent types of operations:</p>
          <p>Input value</p>
          <p>Normalization output (ISO 639-1)
\English"
\eng"
\English and Latin"
Greek; Latin
\en"
\en"
\en", \la"
\el", \la"</p>
          <p>The output of the normalization algorithm is a value from any of several
authoritative language vocabulary. The algorithms work based on a core
vocabulary, the Languages Name Authority List (NAL) published in the
European Union Open Data Portal16. All normalization operations are internally
performed using the core vocabulary and the alignements to other vocabularies
it contains 17. The nal output can be given in accordance with any of the aligned
vocabularies and can be con gured to use URIs from the NAL vocabulary or
any of the aligned ISO code-sets.
5.3</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>Accessibility</title>
          <p>As noted earlier, our approach to measuring multilingual saturation in metadata
allows us not only to measure the data's quality as it is provided by
contributing institutions, but also provides us with insight into the e ectiveness of
Europeana's data enhancement processes, such as semantic enrichment. The numbers
provided as part of the implementation of the quality measures are either not
enough or sometimes too complex to interpret the results and draw conclusions.
At the time of writing, we are in the process of re ning the saturation score
analyzing rst results of the calculations. Visualisation is therefore a necessary tool
to supplement the measure and guide the data providers in their data analysis.
6</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Conclusion and future work</title>
        <p>In this paper, we presented our approach to assessing the multilingual
quality of data in the context of Europeana. This approach takes into account the
functional requirements identi ed for digital libraries and translates them into
actionable measures. These measures look at the data dimensions of
completeness, consistency, and accessibility. Finally, we also present the initial outcome
of our quality measures, such as the de nition and implementation of
normalisation rules for languages. The numbers provided as part of the quality measures
16 https://open-data.europa.eu/en/data/dataset/language
17 NAL is aligned with ISO 639-1 (currently in use at Europeana), ISO 639-2b, ISO
639-2t and ISO 639-3, providing human-readable labels in all European languages
still need to be re ned. Future work on re ning visualisation reports will help us
both to interpret the numbers and to adjust our measurements. For instance, in
order to get a comprehensive view of the quality of a data, the di erent metrics
will need to be presented together (e.g. multilinguality on top of completeness)
so that the interrelation between the di erent metrics is made visible.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>Functional requirements for Bibliographic records: nal report / IFLA Study Group on the Functional Requirements for Bibliographic Records</article-title>
          . No. vol.
          <volume>19</volume>
          in UBCIM publications ; new series, K.G. Saur, Munchen (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. ISO/IEC 25012. , Software engineering {
          <article-title>Software product Quality Requirements and Evaluation (SQuaRE) { Data quality model (</article-title>
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>White</surname>
          </string-name>
          <article-title>Paper on Best Practices for Multilingual Access to Digital Libraries</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <string-name>
            <surname>Europeana</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bellini</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nesi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Metadata Quality Assessment Tool for Open Access Cultural Heritage Institutional Repositories</article-title>
          .
          <source>In: Information Technologies for Performing Arts, Media Access, and Entertainment</source>
          . pp.
          <volume>90</volume>
          {
          <issue>103</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bruce</surname>
            ,
            <given-names>T.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hillmann</surname>
            ,
            <given-names>D.I.</given-names>
          </string-name>
          : The Continuum of Metadata Quality: De ning, Expressing, Exploiting. In: Hillmann,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Westbrooks</surname>
          </string-name>
          , E. (eds.) Metadata in Practice, pp.
          <volume>238</volume>
          {
          <fpage>256</fpage>
          . ALA Editions, Chicago (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Droge, E.:
          <article-title>Criteria for vocabulary evaluation and comparison</article-title>
          .
          <source>Tech. rep., Humboldt-Universitat zu Berlin</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Guy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Powell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Day</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Improving the Quality of Metadata in Eprint Archives</article-title>
          .
          <source>Ariadne (38)</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Europeana data model primer</article-title>
          .
          <source>Tech. rep. (</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mader</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haslhofer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Finding quality issues in SKOS vocabularies</article-title>
          .
          <source>In: Proc. Intl. Conf. onTheory and Practice of Digital Libraries</source>
          . pp.
          <volume>222</volume>
          {
          <fpage>233</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Palavitsinis</surname>
          </string-name>
          , N.:
          <article-title>Metadata Quality Issues in Learning Repositories</article-title>
          .
          <source>Ph.D. thesis</source>
          , Alcala de Henares,
          <source>Spain (Feb</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          r.:
          <article-title>Metadata Quality in Digital Repositories: A Survey of the Current State of the Art</article-title>
          .
          <source>Cataloging &amp; Classi cation Quarterly</source>
          <volume>47</volume>
          (
          <issue>3-4</issue>
          ),
          <volume>213</volume>
          {
          <fpage>228</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Shreeves</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knutson</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stvilia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Twidale</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cole</surname>
          </string-name>
          , T.W.:
          <article-title>Is \Quality" Metadata \Shareable" Metadata? The Implications of Local Metadata Practices for Federated Collections</article-title>
          .
          <source>In: Proc. ACRL Twelfth National Conference</source>
          . pp.
          <volume>223</volume>
          {
          <fpage>237</fpage>
          .
          <string-name>
            <surname>Minneapolis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Minnesota</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Stiller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiraly</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Multilinguality of Metadata. Measuring the Multilingual Degree of Europeana's Metadata</article-title>
          .
          <source>In: Proc. 15th International Symposium of Information Science (ISI)</source>
          . pp.
          <volume>164</volume>
          {
          <issue>176</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Stvilia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Twidale</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A framework for information quality assessment</article-title>
          .
          <source>J. American Society for Information Science &amp; Technology</source>
          <volume>58</volume>
          (
          <issue>12</issue>
          ) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Vogias</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hatzakis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manouselis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szegedi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Extraction and Visualization of Metadata Analytics for Multimedia Learning Object Repositories: The case of TERENA TF-media network (</article-title>
          <year>2013</year>
          ), https://www.terena.org/mail-archives/ tf-media/pdf547CE3lKFt.pdf
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zaveri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rula</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maurino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pietrobon</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Quality assessment for Linked Data: A Survey</article-title>
          .
          <source>Semantic Web</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <volume>63</volume>
          {
          <fpage>93</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>