<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using the SQuaRE series as a guarantee for GDPR compliance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Simonetta</string-name>
          <email>alessandro.simonetta@gmail.com</email>
          <email>alessandro.simonetta@gmail.com ORCID: 0000-0003-2002-9815</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Cristina Paoletti</string-name>
          <email>mariacristina.paoletti@gmail.com</email>
          <email>mariacristina.paoletti@gmail.com ORCID: 0000-0001-6850-1184</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Venticinque</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Enterprise Engineering, University of Rome Tor Vergata</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Naples</institution>
          ,
          <addr-line>Italy, ORCID: 0000-0003-3286-3137</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-In a context where the availability of information represents the opportunity for companies to gain a competitive advantage in the market through the use of sophisticated AI algorithms, data quality assumes a strategic role. With this paper we want to show that the adoption of an international quality measurement standard such as the one present in the SQuaRE series can on the one hand improve the ethical aspect of machine learning algorithms and on the other hand meet the requirements imposed by the European Community regarding the protection of personal data of citizens in Member States (GDPR). Indeed, although the attention to the protection of personal data is mainly directed towards the aspects of security and confidentiality, in a holistic view we should also evaluate the risks arising from the absence of quality in the data. In this context, we consider consistent and of reference for the international community the choice of the Italian legislator made for the Public Administrations. Since 2013 the Agency for Digital Italy (AgID) has suggested the adoption of ISO/IEC 25012 for public administrations in charge of managing databases of national interest. In the article, we propose a methodological approach that ensures the governance of data quality and some open questions regarding the homogeneity of the selected measures. Index Terms-ISO 25000, ISO 25012, ISO 25024, SQuaRE series, GDPR, data quality, COVID-19</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        According to The Economist [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the data represent the new
oil for the modern business, not only related to the IT services,
but also for what concern the business decision and the
marketing campaigns. Many companies are investing in data
analysis, machine learning based algorithms and in solutions
chosen through data driven approaches. In this scenario they
are realizing that the success or their investment are based
not only on the amount of data, that however is an important
aspect, but mainly on their quality. This could have an impact
on the results of machine learning algorithm that are subject
to bias on results if the dataset is not properly chosen or have
quality problems, i.e. contains unbalanced data. These issues
are more evident if the techniques used are taken to extreme as
for example in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where the use of approximated computing
for low power neural network could be more subject to
errors. Furthermore, the benefit of using methodologies such as
Reinforcement Learning, to contrast the degradation of results
and to distribute the decision system as reported in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and in
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], could be frustrated due to poor data quality.
      </p>
      <p>
        An example is the algorithm used in Florida [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to score
the risk of reiteration for people who went in jail that was
subject to bias due to the wrong composition of the data used
to trainee it and the features selected. Indeed, the algorithm
for calculating the score was trained on a dataset where
the criminals were unbalance towards black people and the
weight given to the past record of crimes committed and their
importance was not properly set. Therefore, where this risk
assessment tool of dangerousness and re-offense risk was used
African Americans scored higher in criticality compared to
Caucasian ones based on the skin color also if their records
were less critical.
      </p>
      <p>This case study tell us the importance of the training data
set and their quality and the great impact that can have on
business decision or citizen life, especially if we concentrate
our study on the use that could do public administration or
private companies about the people data.</p>
      <p>
        Attention to the use of data, its collection and its quality
is a very important issue also in Europe, where for years
the legislator has been addressing these issues and investing
resources to align regulations with the problems arising from
new technologies and new business models [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. An important
step taken in Europe is the introduction of the General Data
Protection Regulation (GDPR) 2016/679 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], defined to
harmonize the data privacy laws among the European countries
and, in order to remodel the methods and the approaches that
the organizations manage the European citizens’ data.
      </p>
      <p>
        The full compliance to this regulation in the past was
addressed mainly focused for what concern security issues, but
aspect as compliance, integrity and correctness of data are
now becoming central. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] some issues linked to GDPR
are addressed and in particular the compliance of data
management and usage for business process. The paper propose
three solutions to reduce data maintenance and information
loss, avoiding degradation through data minimization during
the course of business process. The work covers only part
of the accountability principle in that it is not concerned
with monitoring and measuring the quality of and maintaining
the correctness of the data, but addresses the problem of
degradation about information over the time to ensure that
Copyright © 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
the data once processed is usable for business purposes, in
accordance with regulations, even if some of it loses its
correctness.
      </p>
      <p>
        These aspects are taken into account also in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that deals
with the problems of GDPR compliance in use of Public
Geographic Information System for research and practice.
      </p>
      <p>The team apply the pseudonymisation to GIS information to
guarantee the privacy of the users that share their data.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] the OpenEHR standard is proposed to address some
requirements of data privacy regulation. It gives a guideline
to guarantee that software for electronic health records are
interoperable and secure. The health domain is very
sensitive to data quality since the effects of a minimum error
can cause irreparable damage with death or serious injuries.
      </p>
      <p>
        However, the work addresses the requirements of integrity and
traceability related to data quality; the proposed versioning
assure the indelibility of the clinical record preventing any
information from being deleted. The creation of a new version
of the electronic clinical record is important against the lost,
destruction or accidental arm of data. This is only a part of the
data quality and a section of the requirements that sensitive
information must meet, it does not define KPI or measurement
process. Furthermore, the paper is concentrated on clinical data
only and miss to consider other information as personal data
accuracy that are an important issue. In [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] more details on
data quality are presented, and a clear picture of the problem
affecting clinical records is reported. However, the study focus
only on medical information and is not easy to standardize and
applicate to different domains.
      </p>
      <p>According to GDPR, each processing of personal information
must be performed in accordance with the quality principles
established by art. 5 (adequacy, correctness, update, security,
protection, integrity) following the requested criteria of
accountability of the data owner. He must guarantee not only
the respect of these principles, but also the evidence that he
applied all the actions to protect the data (art. 5, par. 2 and
art. 24, par. 1). The regulation assigns specific obligation to
the responsible for the processing, different compared to those
identified for the owner, and in particular the implementation
of appropriate technical and organizational measures to ensure
the security of the treatment (art. 32) through the concepts of
confidentiality, integrity, availability, resilience and ability to
restore.</p>
      <p>
        The GDPR is not the only regulation in which quality
characteristics (accuracy, completeness, correctness, up-to-date,
security, protection, integrity) are reported with a clear meaning
but difficult to compare in the absence of a common metric
described by a calculation algorithm. Indeed, even the European
Solvency II regulation [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], establishes the need for insurance
companies to have internal procedures and processes in place
to ensure the appropriateness, completeness and accuracy of
the data used in the calculation of their technical provisions
(art. 82 Data quality and application of approximations,
including case-by-case methods, for technical provisions). When
granting the basic solvency capital requirement of approval,
insurance supervisors must verify the completeness, accuracy,
and adequacy of the data used (art. 104 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). In addition,
insurance companies must provide for a regular cycle of
validation of their internal model that includes assessment of
the accuracy, completeness, and adequacy of the data used in
the internal model (art. 124 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>
        In this scenario, where regulations require to ensure certain
characteristics related to the management and maintenance of
the data over time, many efforts are directed towards ensuring
a high level of security for information management, thanks
to the application of the ISO/IEC 27000 series [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], but
fewer organizations are concerned about managing the risk
associated with their management and quality. A good solution
could be the application of the ISO 31000 series [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] which
provides principles, a framework and a process for managing
risk. In section II we describe the state of the art related to the
SQuaRE approach to data quality assurance and what benefits
are recognized from its adoption in several case studies. A
focus will be made on the implications of COVID-19 on the
application of GDPR in the health field. In section III we
will present the extension of Italian institutions’ approach to
data quality and we will extend it as a solution to be adopted
by private organizations, also as an indispensable support to
demonstrate full complaince to GDPR. In section IV we will
identify the limitations of this work and how we intend to
address them in the future. Finally, in section V we will present
concluding remarks.
      </p>
    </sec>
    <sec id="sec-2">
      <title>II. STATE OF ART</title>
      <sec id="sec-2-1">
        <title>A. GDPR Data Quality Compliance in COVID-19 Pandemic</title>
        <p>The pandemic emergency boosted the digitalization and the
use of online services. The smart working catching on many
organizations highlighted problems related to data privacy and
data quality, especially for what concern issues that clash with
the need for infection tracking. New type of communication
and workflows are been developed and adopted during this
time with the objective to be accepted by worker and to
transmit them trust in data management and security. A central
point in this new situation is the compliance of all the data to
GDPR. Many organizations had problems related to guarantee
its requirements and to manage the information in the right
way.</p>
        <p>
          Scientific papers address different aspects of the consequence
of COVID-19 pandemic emergency on data privacy and
protection in these two years. In [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] a literature review is
presented about publications that explore the effect of the
COVID19 outbreak on GDPR compliance. The work identifies some
critiques of the regulation and in particular, it focus on the
ethical use of health data during the pandemic. Furthermore,
the infrastructure use and the absence of controls let the
possibility of cross border transfer of data outside the Europe.
These aspects are treated also in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], where the authors study
the use of personal information for research activities to defeat
the COVID-19 and some criticalities linked to specification
of GDPR. The normative has foreseen procedures to support
research in pandemic and the processing of sensitive data,
included personal and health, but the derogation of some
aspect to national laws is an obstacle to a coordinated global
research. This work enhance also the lack of a framework that
support the proof of compliance of data management to desired
requirements. A particular case of this problem is studied in
[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The paper describes the problems linked to traceability
application offered by mobile devices, Android and iOS based,
to monitor the contacts between people and infections. The
European legislation is analyzed with respect to sector-specific
international rules, as the US Health Insurance Portability and
Accountability Act (HIPAA), highlighting the pros and cons
of its flexibility in responding to critical health situation.
These scenarios show that the action to protect data and the
compliance to regulation is often unfulfilled due to the lack of
a common guidelines and the heterogeneity of normative.
Although, it is reported that data accuracy is a non-core
aspect of data privacy, individuals have the right to correct
inaccurate or incomplete personal data that is processed. Using
the SQuaRE series as a data quality measurement standard and
in order to support GDPR compliance provides for a single
reference with respect to individual national regulations helping
to harmonize the application of legislative specializations in
different countries.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>B. The SQuaRE Approach</title>
        <p>In this section, in order to support our work, we report an
extensive scientific production that focuses on the application
of data quality standards and in particular the SQuaRE series to
achieve measurable and well-defined goals. Such approaches
form the basis of our proposal that aims to achieve full
GDPR compliance by achieving the right level of accuracy
and satisfying the entire data quality characteristic by applying
the SQuaRE series.</p>
        <p>
          Some studies propose a framework for data quality evaluation
to let organizations to be able to support and maintain data
quality. In [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] the described framework is based on ISO/IEC
25012 [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and ISO/IEC 25024 [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and consist of a software
that recognize several patterns that identify common failures
of organization and help to evaluate the KPI defined for this
patterns obtaining a clear picture of the data quality. Small
organizations could be not prepared in the application of the
standards to their data. The framework presented and the
tools, that are starting to spread, represent a solution, that,
if widely adopted, can help such companies to guarantee the
compliance to the standards with advantage for services results
and business decisions.
        </p>
        <p>
          The problem of data quality and in particular of bias present
in dataset used for machine learning algorithm is studied in
[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. This type of systems go under the name of Automated
Decision Systems (ADM) as described in [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. These are tools
in which the decision is taken by an AI algorithm in autonomy.
The use of algorithms in software involves aspects of daily
life such as the marketing campaigns to suggest to customers
on a web platform. The work aim is to verify the balance of
the dataset before it will be used in prediction algorithm and it
cause discriminatory problems. The paper defines metrics with
the dual purpose of identifying bias and providing solutions
that can be used within the measurement framework of the
serie SQuaRE. The integration of the proposed metrics into
the data preparation pipeline for machine learning with the
analysis of intrinsic properties of dataset could anticipate the
emergence of discriminatory behavior of algorithms that in
particular case may contravene laws or infringe human rights.
Nowadays, the most successful organizations are those who
are able to collect data, select the right set and guarantee the
best quality. Their decisions follow a data driven approach
and if the basis are wrong, the strategies implemented and the
services offered will be affected with negative consequences.
Therefore, organizations to be confident with the results of
their processing must trust their data. To achieve this level
of confidence organizations are implementing the regulation
present into the standards and applying process and
framework. Many of them are applying data quality evaluation
process and data quality management in order to obtain
certification for their repositories and not only for the software
that they use to process them. In [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] are reported three case
studies of data quality evaluation and certification process
about repositories. An independent entity verifies and certifies
that the company’s database complies with the requirements
defined in the regulation. The organizations are different each
other for what concern dimension and business domain. The
two visions are analyzed to evaluate the impact of the adoption
of the ISO/IEC 25012, ISO/IEC 25024 and ISO/IEC 25040
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] and their benefit recognized in the three organization
before and after the process. The case studies consist, for
each of them, of two phases. The first evaluate the data
quality and address the issues found; the second consist in
another evaluation on the improved databases to obtain the
certification. The results show that applying their methodology
helps the organization to get a better sustainability in the long
term, improve the knowledge of the business and drive the
organizations in better data quality initiatives for the future.
An element often found in these articles is related to the use
of open formats for archiving aspects to promote portability
regardless of the technology used. These are all concepts
defined in the SQuaRE series.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>III. THE SOLUTION PROPOSED</title>
      <p>Previously, we have shown that there are ambiguous
interpretations about characteristic of data quality linked to
regulations and laws that lead to heterogeneous approaches
and uncertainty in the level of adherence to the requirements
that must be met. These issues are addressed by scientific
works that, also with the support of some case studies,
manage to define methodologies and processes to achieve a
high level of data quality through the application of standards.
In the following we will show our proposal starting from
what has been done by the Italian Public Administration to
achieve the levels of data quality it needs.</p>
      <sec id="sec-3-1">
        <title>A. Italian Digitization Process</title>
        <p>In Italy, the process of modernization and digitization of the
Public Administration began with the Digital Administration
Code (Legislative Decree n.82 of March 7, 2005), which states
that public administration data must be made available and
accessible with information and communication technologies
that allow their use and reuse.</p>
        <p>In 2013, AgID (Digital Italy Agency) with the resolution n. 68
defines, for databases of national interest, the compliance with
the quality characteristics defined in the international standard
ISO/IEC 25012 ”Data quality model”.</p>
        <p>The adoption of the standard for databases of national interest
is thus the reference for both the Public Administration and
private companies, also in order to enhance the value of
information assets and improve the quality of services offered and
their efficiency. The adoption of this data quality governance
system responds to their growing need for dissemination and
transparency of the information processed. At the beginning,
the AgID, in order to simplify the adoption of ISO/IEC 25012,
had identified a minimum set of characteristics (accuracy,
completeness, consistency and currentness) from which to start
and then extend to the entire set. Moreover, in the following
years through the Three-Year Plan for Public Administration,
an essential tool to promote the digital transformation of the
Italian Public Administration, the importance of the use of
ISO/IEC 25012 is reaffirmed.</p>
        <p>
          In particular, in the Plan for the three-year period 2020-2022
[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] AgID promotes the increase of data and metadata quality
(OB.2.2) of which the increase in the number of open datasets
conforming to a subset of quality characteristics derived from
the ISO/IEC standard (R.A.2.2b).
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Method</title>
        <p>The proposed approach is based on four main phases
(Fig. 1): the initial design, the exercise of the measurement
process, the evaluation of the obtained results, and the last
phase about the identification of improvement actions. In the
initial design phase we need to choose the reference context
(conceptual, logical and physical level), we define the quality
requirements and the quality model of data to adopt. During
this phase it is evaluated the opportunity to reduce/extend the
model (override/overload of characteristics and/or measures)
and identify the target entities in the appropriate phase of
the data life-cycle. Fig. 2 shows the relationship between
the data quality characteristics provided in ISO/IEC 25012
and the GDPR principles, noting that these relationships have
been highlighted as examples and depend on the context
of application. Since each quality characteristic is calculated
through the sum of the contribution of several sub-measures,
each organization can choose to give more or less importance
to individual contributions by assigning a weight to them.
In addition, if we consider comparing different quality
characteristics, we may find a different granularity of the constituent
contributions to the measures. In this case, the organization
may choose to adopt new sub-characteristics or select a subset
of them in order to make the measurement system balanced.</p>
        <p>The measurement process can take place through commercial
software or, more simply, by building a set of modular queries,
reusable and moldable on different realities depending on the
model identified. The evaluation of the results must be carried
out by an expert who assesses the level of quality achieved
at the end of each iteration and the progress of the same,
at various levels of granularity, over time. The organization
has to implement, if needed, improvement actions to mitigate
risk and improve services through greater IT efficiency and
enhancement of data as an asset.</p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Expected Results</title>
        <p>
          The application of the proposed approach allows the
organization to measure the quality of their data and to identify issues
related to their acquisition and management. The organization
can have the possibility of quality controlled and certified
information. The automatization of data management and control
eliminates manual queries and cleaning operation. High quality
data support the collaboration and the exchange of information
between internal and external company structures.
On the other hand, if organizations do not control data quality,
they may not only be exposed to bias in sensitive data due to
incomplete dataset, but are also subject to different types of
risk [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ].
        </p>
        <p>The methodology proposed in this paper allows the
organization to demonstrate that all activities have been put in place
to control and properly manage the data.</p>
        <p>The application of a standard capable of supporting the
achieving the quality objectives common to the public and private
sectors, allows to trace a collaborative direction between these
two actors, establishing a virtuous cycle where the application
of the same principles makes it easier for both the verify and
the prove of compliance with certain national or European
rules.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>IV. LIMITATION AND FUTURE WORKS</title>
      <p>In this paper, are not considered problems linked to the
traceability and the validity of the source of information
that can play a key role in automate the process of data
controls over the time. In a scenario where the organization
has traceability on the source of its data or part of it, and
can easily know its validity status, it can automate the
measurement of its quality and take action to meet its targets.
As we presented, the data integrity has great importance
in those applications where the quality of the services is
linked to the composition of dataset. The correctness and
the absence of malicious acts is paramount to avoid that
organizations can offer a wrong results to their customers
or make wrong business decision. The spread of services
based on corrupted data could distribute misinformation and
compromise compliance with the directives.</p>
      <p>
        In order to get this target we will evaluate in our future
work the use of blockchain technology. Its use is starting to
be analyzed especially for domain as healthcare and food
and beverage as reported in [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. In this approach, the data
are distributed over a network of nodes that approve the
correct transaction, with a consensus algorithm, and reject the
malicious one. This architecture guarantee the absence of a
single point of failure and of central control that could be a
point of attack. The data are stored in chain of blocks, where
each of this contains the data, a timestamp and the hash of
previous block. If the data inside a block is changed, its hash
will be different from the hash stored into the chain, so the
blocks will be invalidate.
      </p>
      <p>These features are what we need to guarantee the integrity and
the validation of data and between the different methodologies
that implements the blockchain we will consider the use of
the Ethereum and his smart contracts functionality. These
smart contracts enforce a contract or an agreement between
parties through code without the use of an authority.
The Ethereum Virtual Machine (EVM) provide the services
to publish the smart contracts on the Ethereum blockchain.
These can be used to store variables within them, which in
our case can be the data or information related to them. If
on the one hand this allows to validate the contents in the
blocks through automatic operations, public and shared, there
is a limitation in the use of the blockchain due to the size of
the transactions because of the content that you want to put
in the blocks. In particular, it is difficult to ensure that every
participant agrees to and complies with the relevant rules
on personal data protection in public blockchains. Aspects
such as a data principal, a data fiduciary, or a data processor
on a blockchain network have no clear demarcation.The
compliance of these aspects with the GDPR and how the
SQuaRE series can help to solve them will be discussed
in more detail in future works due to their complexity of
treatment.</p>
    </sec>
    <sec id="sec-5">
      <title>V. CONCLUSION</title>
      <p>
        GDPR compliance is well addressed by following
compliance with the ISO/IEC 27000 and ISO/IEC 31000 series,
however the lack of quality in some features mentioned in art. 5
of the GDPR can lead to errors that impact European citizens.
For example, a health recall campaign for a cancer prevention
screening sent to an outdated residential address causes harm
to the citizen who is not reached by the communication. The
use of blockchain technology can help manage this type of
situation: when the organization that owns the data (i.e., the
address book) needs to change information in the blockchain,
it must run a consensus algorithm with the other parties, so all
actors are informed of the change. Also, if the data is changed
without applying the rules of the blockchain, everyone else
knows, automatically, that that data is invalid. The use of
such technology in applications of this type is not yet mature,
and in-depth studies on the choice of appropriate algorithms
to ensure compliance with regulations must be pursued. The
ISO/IEC 25000 series makes it possible to avoid this type of
problem by continuously measuring the quality of the data
held by organizations. Moreover, the scientific literature is
full of examples where the absence of quality in the learning
data of an automated decision making system leads to biased
analyses especially when the sensitive attributes describing
the individuals in the knowledge base are incomplete. Some
studies [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] are trying to relate the unbalance of the learning
dataset with respect to the fairness of automated classifications.
Thus, the importance of data quality is becoming a strategic
goal for many companies that often find themselves using
replicas of out-of-date data.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>The</given-names>
            <surname>Economist</surname>
          </string-name>
          , “
          <article-title>The world's most valuable resource is no longer oil, but data,” The Economist</article-title>
          , USA, 6th May
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Cardarilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fazzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giardino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. S. S.</given-names>
            <surname>Patetta</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          and Re, “
          <article-title>Approximated computing for low power neural networks</article-title>
          ,
          <source>” TELKOMNIKA</source>
          , vol.
          <volume>7</volume>
          , pp.
          <fpage>1236</fpage>
          -
          <lpage>1241</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Matta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Cardarilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fazzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giardino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nannarelli</surname>
          </string-name>
          , M. Re, and S. Spano`, “
          <article-title>A reinforcement learning-based qam/psk symbol synchronizer</article-title>
          ,
          <source>” IEEE Access</source>
          , vol.
          <volume>7</volume>
          , pp.
          <volume>124</volume>
          <fpage>147</fpage>
          -
          <lpage>124</lpage>
          157,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Canese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Cardarilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fazzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giardino</surname>
          </string-name>
          , M. Re, and
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Spano`, “Multi-agent reinforcement learning: A review of challenges and applications</article-title>
          ,” Applied Sciences, vol.
          <volume>11</volume>
          , no.
          <issue>11</issue>
          ,
          <year>2021</year>
          . [Online]. Available: https://www.mdpi.com/2076-3417/11/11/4948
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Angwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mattu</surname>
          </string-name>
          , and L. Kirchner, “
          <article-title>Machine bias : There's software used across the country to predict future criminals. and it's biased against blacks</article-title>
          .” https://www.propublica.org/,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] Council of Europe, “
          <string-name>
            <surname>Recommendation</surname>
            <given-names>CM</given-names>
          </string-name>
          /Rec(
          <year>2020</year>
          )
          <article-title>1 of the Committee of Ministers to member States on the human rights impacts of algorithmic systems</article-title>
          ,”
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>European</given-names>
            <surname>Parliament</surname>
          </string-name>
          and
          <article-title>Council of Europe, “Regulation (eu) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/ec (general data protection regulation</article-title>
          ),”
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zaman</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hassani</surname>
          </string-name>
          , “
          <article-title>On enabling gdpr compliance in business processes through data-driven solutions,” SN Computer Science</article-title>
          , vol.
          <volume>1</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          , Jun.
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hasanzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kajosaari</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Ha¨ggman, and M. Kytta¨, “A context sensitive approach to anonymizing public participation gis data: From development to the assessment of anonymization effects on data quality,” Computers, Environment and Urban Systems</article-title>
          , vol.
          <volume>83</volume>
          , p.
          <fpage>101513</fpage>
          ,
          <year>2020</year>
          . [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0198971520302465
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sousa</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Gonc¸alves-</article-title>
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
            , G. Bacelar,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Frade</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Pestana</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Cruz-Correia</surname>
          </string-name>
          ,
          <article-title>“openehr based systems and the general data protection regulation (gdpr),” Studies in health technology and informatics</article-title>
          , vol.
          <volume>247</volume>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>95</lpage>
          ,
          <year>01 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V. C.</given-names>
            <surname>Pezoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Kourou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Kalatzis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. P.</given-names>
            <surname>Exarchos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Venetsanopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Zampeli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gandolfo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Skopouli</surname>
          </string-name>
          , S. De Vita,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Tzioufas</surname>
          </string-name>
          , and
          <string-name>
            <surname>D. I. Fotiadis</surname>
          </string-name>
          , “
          <article-title>Medical data quality assessment: On the development of an automated framework for medical data curation,” Computers in Biology and Medicine</article-title>
          , vol.
          <volume>107</volume>
          , pp.
          <fpage>270</fpage>
          -
          <lpage>283</lpage>
          ,
          <year>2019</year>
          . [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0010482519300733
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>European</given-names>
            <surname>Parliament</surname>
          </string-name>
          and Council of Europe, “
          <article-title>Directive 2009/138/EC of the European Parliament and of the Council of 25 November 2009 on the taking-up and pursuit of the business of Insurance and Reinsurance (Solvency II)</article-title>
          ,”
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13] International Organization for Standardization, ”ISO/IEC 27000:
          <year>2018</year>
          , “Information Technology,
          <source>Security Techniques, Information Security Management Systems</source>
          ,Overview and Vocabulary”, International Organization for Standardization Std.,
          <year>2018</year>
          . [Online]. Available: https://www.iso.org/standard/73906.html (accessed Nov,
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14] International Organization for Standardization, ”ISO 31000:
          <year>2018</year>
          <article-title>(en) Risk management - Guidelines”, International Organization for Standardization Std</article-title>
          .,
          <year>2018</year>
          . [Online]. Available: https://www.iso.org/iso31000-risk-management.
          <source>html (accessed Nov</source>
          ,
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>McLennan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Celi</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Buyx</surname>
          </string-name>
          , “
          <article-title>Covid-19:putting the general data protection regulation to thetest,” JMIR Public Health and Surveillance</article-title>
          , vol.
          <volume>6</volume>
          ,
          <issue>04</issue>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Thorogood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ordish</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Beauvais</surname>
          </string-name>
          , “
          <article-title>Covid19 research: Navigating the european general data protection regulation</article-title>
          ,
          <source>” J Med Internet Res</source>
          , vol.
          <volume>22</volume>
          ,
          <year>2020</year>
          . [Online]. Available: http://www.jmir.org/
          <year>2020</year>
          /8/e19799/
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bradford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aboy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Liddell</surname>
          </string-name>
          , “
          <article-title>Covid-19 contact tracing apps: A stress test for privacy, the gdpr, and data protection regimes,” Journal of law and the biosciences</article-title>
          , vol.
          <volume>7</volume>
          , p.
          <year>lsaa034</year>
          ,
          <year>05 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Calabrese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Esponda</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Pesado</surname>
          </string-name>
          , “
          <article-title>Framework for Data Quality Evaluation Based on ISO/IEC 25012</article-title>
          and ISO/IEC 25024,” in VIII Conference on Cloud Computing,
          <source>Big Data &amp; Emerging Topics</source>
          ,
          <year>2020</year>
          . [Online]. Available: http://sedici.unlp.edu.ar/handle/10915/104778
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19] International Organization for Standardization, ”ISO/IEC 25012:
          <year>2008</year>
          <article-title>Software engineering - Software product Quality Requirements and Evaluation (SQuaRE) - Data quality model”, International Organization for Standardization Std</article-title>
          .,
          <year>2008</year>
          . [Online]. Available: https://www.iso.org/standard/35736.html (accessed Nov,
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20] International Organization for Standardization, ”ISO/IEC 25024:
          <article-title>2015 Systems and software engineering - Systems and software Quality Requirements and Evaluation (SQuaRE) - Measurement of data quality”, International Organization for Standardization Std</article-title>
          .,
          <year>2015</year>
          . [Online]. Available: https://www.iso.org/standard/35749.html (accessed Nov,
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Simonetta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Trenta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Paoletti</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A</surname>
          </string-name>
          . Vetro`, “
          <article-title>Metrics for identifying bias in datasets,” SYSTEM</article-title>
          , in press.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fahy</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Appelman</surname>
          </string-name>
          , Netherlands/Research,
          <year>2020</year>
          , pp.
          <fpage>164</fpage>
          -
          <lpage>175</lpage>
          , chapter in
          <source>: Report Automating Society</source>
          <year>2020</year>
          ,
          <string-name>
            <surname>Chiusi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kayser-Bril</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Spielkamp</surname>
          </string-name>
          , M. eds., Berlin: AlgorithmWatch,
          <year>October 2020</year>
          . [Online]. Available: https://www.ivir.nl/publicaties/download/AutomatingSociety-Report-2020.pdf/
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gualo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rodriguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Verdugo</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Caballero</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Piattini</surname>
          </string-name>
          , “
          <article-title>Data quality certification using iso/iec 25012: Industrial experiences</article-title>
          ,
          <source>” Journal of Systems and Software</source>
          , vol.
          <volume>176</volume>
          , p.
          <fpage>110938</fpage>
          ,
          <year>2021</year>
          . [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0164121221000352
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24] International Organization for Standardization, ”ISO/IEC 25040:
          <article-title>2011 Systems and software engineering - Systems and software Quality Requirements and Evaluation (SQuaRE) - Evaluation process”, International Organization for Standardization Std</article-title>
          .,
          <year>2011</year>
          . [Online]. Available: https://www.iso.org/standard/35765.html (accessed Nov,
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Agenzia per l'Italia Digitale</surname>
          </string-name>
          ,
          <source>Piano Triennale 2020-2022. Presidenza del Consiglio dei Ministri</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Simonetta</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Vetro`,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Paoletti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Torchiano</surname>
          </string-name>
          , “
          <article-title>Integrating SQuaRE data quality model with ISO 31000 risk management to measure and mitigate software bias</article-title>
          ,
          <source>” IWESQ</source>
          <year>2021</year>
          , in press.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Grecuccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Giusto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Fiori</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Rebaudengo</surname>
          </string-name>
          , “
          <article-title>combining blockchain and iot: Food-chain traceability and beyond</article-title>
          .”
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetro`</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Torchiano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mecati</surname>
          </string-name>
          , “
          <article-title>A data quality approach to the identification of discrimination risk in automated decision making systems,” Government Information Quarterly</article-title>
          , vol.
          <volume>38</volume>
          , no.
          <issue>4</issue>
          , p.
          <fpage>101619</fpage>
          ,
          <year>2021</year>
          . [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0740624X21000551
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>