<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ethics-aware data governance (Vision Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Letizia Tanca</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Atzeni</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Azzalini</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilaria Bartolini</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Cabibbo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Calderoni</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Ciaccia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valter Crescenzi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Carlos De Martin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Selina Fenoglietto</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Donatella Firmani</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergio Greco</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Isgro</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Maio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Martinenghi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maristella Matera</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Merialdo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristian Molinaro</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Patella</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Prevete</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisa Quintarelli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Santangelo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Tagarelli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guglielmo Tamburrini</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riccardo Torlone</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>The number of datasets available to legal practitioners, policy makers, scientists, and many other categories of citizens is growing at an unprecedented rate. Ethics-aware data processing has become a pressing need, considering that data are often used within critical decision processes (e.g., sta evaluation, college admission, criminal sentencing). The goal of this paper is to propose a vision for the injection of ethical principles (fairness, non-discrimination, transparency, data protection, diversity, and human interpretability of results) into the data analysis lifecycle (source selection, data integration, and knowledge extraction) so as to make them rst-class requirements. In our vision, a comprehensive checklist of ethical desiderata for data protection and processing needs to be developed, along with methods and techniques to ensure and verify that these ethically motivated requirements and related legal norms are ful lled throughout the data selection and exploration processes. Ethical requirements can then be enforced at all the steps of knowledge extraction through a uni ed data modeling and analysis methodology relying on appropriate conceptual and technical tools.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Traditional knowledge extraction (search, query, or data analysis) systems hardly
pay any speci c attention to ethically sensitive aspects and to the ethical and
social problems their outcomes could bring about. However, such aspects are
now becoming prominent, especially with regard to the protection of
fundamental human rights and their underpinnings in normative ethics [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. These
demands are broadly re ected into codes of ethics, in the Responsible Research
and Innovation (RRI) approach of the European Commission HORIZON 2020,
in statements of the European Group on Ethics in Science and Technology (EGE,
https://ec.europa.eu/research/ege) - with regard to research and
innovation in the area of automated data selection and exploration processes - and also
in legally binding regulations such as the EU General Data Protection
Regulation (GDPR, https://www.eugdpr.org). More recently, computer scientists
themselves - via professional organizations such as ACM and Informatics Europe
- have been stressing the importance of raising, within their discipline,
awareness with respect to ethical and societal issues regarding the use of data [
        <xref ref-type="bibr" rid="ref18 ref34">18,
34</xref>
        ]. Coherently with this broad ethical and legal framework, we propose a
vision aiming to enhance and verify the protection and advocacy of these rights
and values throughout each step of the knowledge extraction chain, thus
providing all involved stakeholders with a set of novel computational methodologies
and techniques to protect and promote fundamental ethical guarantees in data
governance. Realizing this vision needs to be achieved in a principled way by
analyzing the ethical challenges that must be addressed, by properly amalgamating
and resolving contrasts between various ethical demands (e.g., transparency vs
protection), and by developing an extensive list of ethical desiderata for data
protection and processing. The resulting Ethical CheckList (ECL) will constitute
the high-level speci cations for any project embracing this vision.
      </p>
      <p>Accordingly, a preliminary goal is to analyze and clarify the relevant
meanings of ethically motivated desiderata about data processing. These notably
include transparency, interpretability and understandability, in addition to
nondiscrimination (fairness), diversity protection and their ethical underpinnings.
Moreover, the conceptual relationships between these desiderata and their
mutual tensions needs to be analyzed, with the aim of identifying acceptable
tradeo s for integrating ethical policies in data selection and exploration processes.
This analysis may be used to identify ethically motivated speci cations for the
various models, technologies and tools to be delivered. With this, our vision
will therefore contribute to establish higher social standards for transparency,
privacy protection, fairness and non-discrimination in data governance.</p>
      <p>The observation and enforcement of ethical principles in data management
are achieved by considering the following steps of the analysis lifecycle: i) source
selection, ii) data integration, and iii) knowledge extraction. Ensuring that the
ECL is applied during all such phases allows all stakeholders to enact law
regulations and ethical principles and verify that these are respected. Our vision will
hence give birth to an ethical data analysis methodology and associated
methodological and technical tools that will enforce the ECL speci cations throughout
the three above mentioned steps of the analysis lifecycle.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>
        Our vision's methodology builds on methods and tools for data management
and on recent studies on data source selection, data integration, and knowledge
extraction. The overall approach is to tackle problems that have a practical
signi cance, providing general methods as well as concrete tools that demonstrate
the approach. The relevant activities can be developed on two parallel tracks.
On one track, a new, ethics-aware, knowledge extraction lifecycle is carried on,
where the experts in ethical, legal, and social disciplines work side-by-side with
the computer scientists to i) identify the ethical requirements, ii) produce the
desiderata checklist and iii) identify the most appropriate modeling tool(s) to
combine them with the application requirements into a coherent framework. The
other track is devoted to the study of novel methods for data source selection,
integration, and knowledge extraction that comply with the identi ed ethical
requirements. The two tracks are bridged by a conceptual model, based on the
Context Dimension Model (CDM) of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] described below, able to express the
ethical requirements of the ECL by associating the various ethical dimensions (i.e.,
fairness, transparency, diversity, and data protection), along with their di erent
in ections and levels of enforcement, with the knowledge extraction lifecycle
activities. Providing a formal model of ethical requirements, to be adopted in each
speci c situation of use, will contribute to mitigate or remove possible biases
from the considered data management methods. Since at present no such
integrated solutions for dealing with ethical issues in data management
exist, our vision will provide a methodological and technological breakthrough
and a concrete answer to issues that the scienti c community has started
investigating in recent years [
        <xref ref-type="bibr" rid="ref10 ref25 ref30">30, 25, 10</xref>
        ] (http://wp.sigmod.org/?p=1900). Existing
notable approaches, such as [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], consider scenarios in which there is no control
on the knowledge extraction lifecycle, thus dealing with the speci c, orthogonal
goal of discovering bias in online information derived by third parties through
the application of knowledge extraction techniques.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Ethical issues in the knowledge extraction lifecycle</title>
      <p>We now elaborate on how the various issues arising in the knowledge extraction
lifecycle are to be dealt with and how they advance the state of the art.
3.1</p>
      <sec id="sec-3-1">
        <title>Source selection</title>
        <p>
          Within this phase we focus on choosing the data source(s) appropriate for the
target objectives in terms of quality and satisfaction of ethical requirements (such
as trustability, personal rights protection, and fairness) taking into account the
di erences between categories and favoring the maximization of source
diversity. Under this view, source selection is a novel data management problem,
motivated by the recent proliferation of data. The research project more related
to our goals is SourceSight [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ], which proposes a system to interactively
discover valuable sets of sources taking into account reliability and data quality,
but without considering ethical issues. Along with institutional datasets, data
from location-based services, and other kinds of open data, we also consider
online social networks, which are nowadays the preferred communication means
for information spread and opinion sharing. We address this problem by
assessing the ethical requirements provided by the ECL in this phase, e.g., the
bias/fairness/authority of the source population. To this end, we plan to design
an iterative process in which the ECL is rst assessed for each candidate source
so as to select the most informative sources for the domain of interest. Innovative
ways of labeling datasets with ethically- and socially-aware metadata are also
needed, with the aim of detecting potential limits or aws early on (e.g., gender
imbalances, ethnic or class misrepresentation or underrepresentation).
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Data integration</title>
        <p>
          Within our vision, suitable tools assist the combination of data from di erent
sources and the extraction of real-world entities. This issue are tackled, for data
reconciliation, by leveraging existing record linkage methods and, for schema
mapping, by tailoring ethically-aware integration views on the basis of the
ethical context. Record linkage seeks to identify which objects refer to the same
real-world entity, and is fundamental for data integration. Leveraging humans
to compare records based on domain knowledge enables high-accuracy linkage
in various domains [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]; however, unrestricted human access to data may not be
suitable in the presence of sensitive information. For this reason, our goals
include the prevention of sensitive informations leaks, and the collection of all the
intermediate data transformation that yielded a given integration result, with
the goal of enforcing the protection of user data and providing a transparent
access to result generation. We also consider review methods for including the
user in the data integration loop, for both data and schema reconciliation, so
that high-quality integrated views and ne-grained feedbacks on the result can
be provided. Merging con icting information is critical for recent data
integration research [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], especially when dealing with web sources and ethical aspects.
Although speci c source properties computed in the previous source selection
step can help in the detection of fake values, the integration step can further
contribute to this problem, empowering human experts with ubiquitous
collaborative review tools (thus promoting fact-checking culture).
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Knowledge extraction</title>
        <p>
          We focus on various kinds of knowledge extraction methods: 1. Result
personalization, 2. Information di usion and in uence propagation in social networks,
3. Explanation models enabling transparency and interpretability of results,
4. Privacy and security in knowledge extraction systems. Note that the research
on topics 1. and 2. is related to the modi cation of existing techniques for
knowledge extraction to guarantee that the process and results satisfy the appropriate
ethical requirements, while topics 3. and 4. study techniques to enforce and verify
that the ethical requirements be satis ed by the analysis process.
Result personalization. Personalization can be broadly de ned as providing an
overall customized, individualized user experience by taking into account the
needs, preferences and characteristics of a user or group of users [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. A data
personalization method may re-rank the items in a collection to be shown to a
user, focus only on items of interest, or recommend additional options. While
personalization delivers relevant content, it also polarizes the perspectives and
diminishes serendipity, whereas searching for relevant information on large datasets
should provide results that are diverse enough [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to ensure a fair coverage of the
available alternatives. Methods that aim to qualify and quantify personalized
experiences and their biasing e ects have been overlooked in the literature.
        </p>
        <p>
          Queries aiming to return only the more interesting results can either provide a
ranking of objects (top-k queries) or consider some form of dominance to exclude
sub-optimal choices (skyline queries [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]). Thanks to a recent contribution [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], free
parameters (e.g., weights), are available to ne-tune the behavior of both kinds
of query. For any speci c dataset, one can characterize the parameter values that
guarantee that the output satis es the required ethical properties (e.g.,
preserving the distribution of a protected attribute). Research on such issues for ranking
queries has provided preliminary results [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ], but no study exists yet for skyline
queries. E cient methods are needed to determine such a set of parameter
values, to study its stability wrt. a change in input data [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], and to provide metrics
to choose among di erent parameter con gurations. This requires analyzing the
overhead incurred by methods providing such ethical guarantees, and studying
the trade-o between the e ciency and the ethical level of the query process.
Such techniques are also useful to ensure that the items of a query result are
diverse enough, providing a balanced view of the result space.
        </p>
        <p>
          The current user context has been adopted as a criterion for personalization
by knowledge ltering. Design methods that support dynamic, context-based
ltering of pertinent resources [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] facilitate the development of software that
takes context into account. In the traditional software lifecycle, the CDM [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
represents the relevant dimensions of context (e.g., current location, situation
of use, role of the user), along with a hierarchy of their possible values; any
set of dimension values represents a possible context. Depending on the current
context, only relevant data are provided to the user. This kind of personalization
is naturally coupled with the ethical dimensions described above, also supporting
user preferences and recommendations aware of both context and ethics [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ].
Information di usion and in uence propagation in social networks. Online social
networks (OSNs) are nowadays the preferred communication means for
spreading information and sharing knowledge, such as advertising products/services,
promoting ideas, sharing opinions. In this regard, in uence maximization is
central, i.e., to identify k initial in uencers that maximize the spread of in uence [
          <xref ref-type="bibr" rid="ref15 ref33">33,
15</xref>
          ]. An important but often overlooked aspect is that success of an information
di usion process might depend not only on the investment-budget (k), but also
on the diversity of the initial in uencers, as well as of the targets to be in
uenced. Members of an OSN present two kinds of diversity: static, which includes
diversity of kind, socio-cultural aspects and other characteristics exogenous to
the OSN; dynamic, which includes the knowledge, community experience, and
shared information acquired over time. The various types of user diversity should
be leveraged to push forward research on information di usion and in uence
propagation along two main directions: i) diversity concerning the targets to
be in uenced and ii) diversity concerning the initial in uencers. The former
allows us to capture non-discrimination or fairness aspects in the outcome of the
di usion process; the latter enables modeling di erent triggering stimuli, which
intuitively capture utilitarian aspects [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] in terms of marketing principles (e.g.,
diversi cation of users skills implies higher productivity). Addressing both
fairness and utility opens to opportunities of ethics-preserving information di usion,
and can support the development of advanced methods around novel perspectives
having ethical implications in OSN data analysis. One such perspective is related
to the ever increasing phenomenon of fake-news/misinformation spread on the
Web: bringing fairness and utility-oriented diversity aspects into fact-checking
and misinformation debunking prompts us to develop sophisticated models to
handle competitive in uence propagation scenarios. There has been little work in
diversity in information di usion and in uence propagation [
          <xref ref-type="bibr" rid="ref1 ref13 ref32 ref5">1, 32, 13, 5</xref>
          ]. Most of
the existing notions of diversity have been developed around structural features
of the network, or are based on user pro le attributes, but no existing approach
proposes diversity-aware solutions in in uence propagation.
        </p>
        <p>
          Explanation models enabling transparency and interpretability of results.
Explanation models for results provided by big data transformations and machine
learning (ML) systems are needed. The transparency requirement must be
balanced with privacy preservation by identifying context-sensitive trade-o s. With
the aim of supporting transparency in big data transformations, we consider
techniques for data lineage, so as to trace the relationships between input and
output data and to identify underlying processes. Data lineage (aka provenance)
concerns data origin and transformation history, and enables transparency by
supporting explanations of results and processes. However, it may also disclose
private or con dential data, and the use of proprietary transformations [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
Extending provenance techniques to Big Data poses new challenges and
opportunities [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]. Providing explanations for ML systems that are often opaque to
human beings is a challenge addressed in the emerging XAI (eXplainable AI)
research area [
          <xref ref-type="bibr" rid="ref20 ref23">20, 23</xref>
          ]. In ML classi cations, explanation requests are expressed
as why-questions: Why were input data associated with class X? [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Answers
may come in the form of I/O explanations (exhibiting prototypes of the output
class) and inner explanations (additionally exhibiting salient components of both
input and intermediate processing data). In I/O explanations one often exhibits
a single prototype. In real-life, however, a single prototype may not be
representative of the entire class, and discarded classi cation possibilities are important
to interpret outcomes. Accordingly, state-of-art tools need to be extended by
extracting multiple prototypes for both classi cation results and discarded classes
for Deep Learning networks and other ML systems, e.g., via statistical methods
like activation maximization [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] and sparse coding approaches [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Similarly,
inner explanation tools need to be extended by identifying components of input
and intermediate processing data contributing most to classi cation outcomes.
Privacy and security in knowledge extraction systems. Focusing on methods for
data protection oriented to the preservation of privacy when releasing analysis
results, we plan to study techniques for trading-o privacy budget for result
accuracy. Aspects concerning privacy and security of personal data are increasingly
relevant, as also testi ed by the GDPR, which uni es data protection laws across
all European Union members. Several techniques have been developed so far for
privacy protection. Di erential privacy promises to enable general data
analytics while protecting individual privacy, yet existing mechanisms do not support
the wide variety of features and sources used in big-data analytics systems [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
Data anonymization attempts to provide privacy while allowing general-purpose
analysis, but recent de-anonymization results [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] proved that it cannot be relied
upon. A further technique is homomorphic encryption, which makes it possible
to perform certain operations on a ciphertext without decrypting it [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ].
Homomorphic encryption techniques were recently coupled with a novel probabilistic
data structure, the Spatial Bloom Filter (SBF) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], designed to secure out
location data. This data structure is suitable for any kind of set-based problem [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ],
and could be conveniently used in a number of applications. We focus on possible
applications of the SBF for data protection. For instance, since an individual's
interest may be seen as his own membership to a speci c set, a promising
application of the SBF is related to those services that rely on outsourced data re ecting
people's interests for marketing actions and other commercial purposes. We
focus on the study of methods for data protection oriented to the preservation of
the privacy of individuals when releasing statistics. Indeed, aggregated data can
be combined for inferring personal information, even if randomized.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Outlook</title>
      <p>Embracing our vision will contribute to establish higher standards in democratic
societies for transparency, privacy protection, fairness and non-discrimination
in data governance. Trust building and public con dence will be fostered by
supporting scrutiny without violating privacy, and by reducing the opaqueness
of current data technologies. The resulting competencies and techniques will
allow addressing the ethical issues in ethics-critical decision-making processes.
On the whole, this vision will contribute to a more cognizant and responsible
use of data, by promoting the ethics-aware development and use of technologies
and systems for collecting, storing and processing data.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Q.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. K.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Zhang.</surname>
          </string-name>
          <article-title>Incorporating structural diversity of neighbors in a di usion model for social networks</article-title>
          .
          <source>Web Intelligence</source>
          <year>2013</year>
          :
          <volume>431</volume>
          {
          <fpage>438</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bolchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Quintarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tanca</surname>
          </string-name>
          . CARVE:
          <article-title>Context-aware automatic view definition over relational databases</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>38</volume>
          (
          <issue>1</issue>
          ):
          <fpage>45</fpage>
          -
          <lpage>67</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>I. Catallo</surname>
          </string-name>
          et al.
          <article-title>Top-k diversity queries over bounded regions</article-title>
          .
          <source>TODS</source>
          <volume>38</volume>
          (
          <issue>2</issue>
          ):
          <volume>10</volume>
          :1{
          <fpage>10</fpage>
          :
          <fpage>44</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J.</given-names>
            <surname>Chomicki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ciaccia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Meneghetti</surname>
          </string-name>
          .
          <article-title>Skyline queries, front and back</article-title>
          .
          <source>SIGMOD Record</source>
          <volume>42</volume>
          (
          <issue>3</issue>
          ):
          <fpage>6</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Calio</surname>
          </string-name>
          et al.
          <article-title>Topology-driven Diversity for Targeted In uence Maximization with Application to User Engagement in Social Networks</article-title>
          .
          <source>IEEE TKDE</source>
          , To appear
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Ciaccia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Martinenghi</surname>
          </string-name>
          .
          <source>Reconciling Skyline and Ranking Queries. PVLDB</source>
          <volume>10</volume>
          (
          <issue>11</issue>
          ):
          <fpage>1454</fpage>
          -
          <lpage>1465</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>L.</given-names>
            <surname>Calderoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Palmieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maio</surname>
          </string-name>
          .
          <article-title>Location privacy without mutual trust: The Spatial Bloom Filter Comput</article-title>
          . Commun., vol.
          <volume>68</volume>
          :. 4{
          <issue>16</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>L.</given-names>
            <surname>Calderoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Palmieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maio</surname>
          </string-name>
          .
          <article-title>Probabilistic Properties of the Spatial Bloom Filters and Their Relevance to Cryptographic Protocols</article-title>
          .
          <source>IEEE Transactions on Information Forensics and Security</source>
          <volume>13</volume>
          (
          <issue>7</issue>
          ):
          <volume>1710</volume>
          {
          <fpage>1721</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>X.L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Berti-Equille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          .
          <article-title>Data Fusion: Resolving Con icts from Multiple Sources</article-title>
          .
          <source>Handbook of Data Quality</source>
          . Springer, Berlin, Heidelberg (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>M. Drosou</surname>
          </string-name>
          et al..
          <article-title>Diversity in big data: a review</article-title>
          .
          <source>Big Data</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ):
          <volume>73</volume>
          {
          <fpage>84</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. S. B.
          <string-name>
            <surname>Davidson</surname>
          </string-name>
          et al.
          <article-title>On provenance and privacy</article-title>
          .
          <source>ICDT</source>
          <year>2011</year>
          :
          <fpage>3</fpage>
          -
          <lpage>10</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>G. Koutrika. Data</given-names>
            <surname>Personalization</surname>
          </string-name>
          .
          <source>In Data Management in Pervasive Systems</source>
          , Springer Verlag, ISBN 978-3-
          <fpage>319</fpage>
          -20061-3:
          <fpage>213</fpage>
          -
          <lpage>234</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Y.
          <string-name>
            <surname>-H. Fu</surname>
          </string-name>
          , C.-Y. Huang, C.-T. Sun.
          <article-title>Using global diversity and local topology features to identify in uential network spreaders</article-title>
          .
          <source>Physica A 433</source>
          (C):
          <volume>344</volume>
          {
          <fpage>355</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. L.
          <string-name>
            <surname>Floridi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Taddeo</surname>
          </string-name>
          .
          <article-title>What is data ethics? Phil</article-title>
          .
          <source>Trans. R. Soc. A374:20160360</source>
          . http://dx.doi.org/10.1098/rsta.
          <year>2016</year>
          .
          <volume>0360</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. S. Galhotra et al.
          <article-title>ASIM: A Scalable Algorithm for In uence Maximization under the Independent Cascade Model</article-title>
          .
          <source>WWW</source>
          <year>2015</year>
          :
          <volume>35</volume>
          {
          <fpage>36</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. S. Galhotra et al.
          <article-title>Robust Entity Resolution using Random Graphs</article-title>
          .
          <source>SIGMOD</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>D.</given-names>
            <surname>Gunning</surname>
          </string-name>
          .
          <article-title>Explainable Arti cial Intelligence (XAI)</article-title>
          .
          <source>Defense Advanced Research Projects Agency (DARPA)</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Informatics</surname>
            <given-names>Europe</given-names>
          </string-name>
          &amp;
          <source>EUACM. When Computers Decide: European Recommendations on Machine-Learned Automated Decision Making</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>N.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Near</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <article-title>Towards Practical Di erential Privacy for SQL Queries</article-title>
          . PVLDB,
          <volume>11</volume>
          (
          <issue>5</issue>
          ):
          <volume>526</volume>
          {
          <fpage>39</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>Z.C.</given-names>
            <surname>Lipton</surname>
          </string-name>
          .
          <article-title>The mythos of model interpretability</article-title>
          .
          <source>arXiv:1606.03490</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>J. Mairal</surname>
          </string-name>
          et al.
          <article-title>Online learning for matrix factorization and sparse coding</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>11</volume>
          :
          <fpage>19</fpage>
          -
          <lpage>60</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>K. Mens</surname>
          </string-name>
          et al.
          <article-title>Modeling and managing context-aware systems variability</article-title>
          .
          <source>IEEE Software</source>
          ,
          <volume>34</volume>
          (
          <issue>6</issue>
          ):
          <volume>58</volume>
          {
          <issue>6</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. G. Montavon et al.
          <article-title>Methods for interpreting and understanding deep neural networks</article-title>
          .
          <source>Digital Signal Processing</source>
          <volume>73</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vitaly. Robust</surname>
          </string-name>
          de
          <article-title>-anonymization of large sparse datasets</article-title>
          .
          <source>IEEE Symposium on Security and Privacy</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25. U. Pagallo.
          <article-title>On the Principle of Privacy by Design and its Limits: Technology, Ethics and the Rule of Law</article-title>
          .
          <source>European Data Protection</source>
          <year>2012</year>
          :
          <fpage>331</fpage>
          -
          <lpage>346</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>P.</given-names>
            <surname>Paillier</surname>
          </string-name>
          .
          <article-title>Public-key cryptosystems based on composite degree residuosity classes</article-title>
          .
          <source>Advances in CryptologyEUROCRYPT (LNCS)</source>
          , vol.
          <volume>1592</volume>
          :
          <issue>223</issue>
          {
          <fpage>238</fpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27. E. Pitoura et al.
          <source>On Measuring Bias in Online Information. SIGMOD Record</source>
          <volume>46</volume>
          (
          <issue>4</issue>
          ):
          <fpage>16</fpage>
          -
          <lpage>21</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28. E.
          <string-name>
            <surname>Quintarelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Rabosio</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Tanca</surname>
          </string-name>
          .
          <article-title>Recommending New Items to Ephemeral Groups Using Contextual User In uence</article-title>
          .
          <source>RecSys</source>
          <year>2016</year>
          :
          <fpage>285</fpage>
          -
          <lpage>292</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>T.</surname>
          </string-name>
          Rekatsinas et al. SourceSight
          <string-name>
            <surname>: Enabling E ective Source</surname>
          </string-name>
          <article-title>Selection</article-title>
          .
          <source>SIGMOD</source>
          <year>2016</year>
          :
          <fpage>2157</fpage>
          -
          <lpage>2160</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>J. Stoyanovich</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Abiteboul</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Miklau. Data</surname>
          </string-name>
          <article-title>Responsibly: Fairness, Neutrality and Transparency in Data Analysis</article-title>
          .
          <source>EDBT</source>
          <year>2016</year>
          :
          <fpage>718</fpage>
          -
          <lpage>719</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>M. Soliman</surname>
          </string-name>
          et al.
          <article-title>Ranking with uncertain scoring functions: semantics and sensitivity measures</article-title>
          .
          <source>SIGMOD</source>
          <year>2011</year>
          :
          <fpage>805</fpage>
          -
          <lpage>816</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32. F.Tanget al.
          <article-title>Diversi edsocialin uencemaximization</article-title>
          .
          <source>ASONAM2014:455{459</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Shi.</surname>
          </string-name>
          <article-title>In uence maximization: near-optimal time complexity meets practical e ciency</article-title>
          .
          <source>SIGMOD</source>
          <year>2014</year>
          :
          <volume>75</volume>
          {
          <fpage>86</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34. USA ACM.
          <article-title>Statement on the Importance of Preserving Personal Privacy - Foundational Privacy Principles and Practices</article-title>
          . (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35. J.
          <string-name>
            <surname>Wang</surname>
          </string-name>
          et al.
          <article-title>Big data provenance: Challenges, state of the art and opportunities</article-title>
          .
          <source>Big Data</source>
          <year>2015</year>
          :
          <fpage>2509</fpage>
          -
          <lpage>2516</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>M. Zehlike</surname>
          </string-name>
          et al.
          <article-title>FA*IR: A Fair Top-k Ranking Algorithm</article-title>
          .
          <source>CIKM</source>
          <year>2017</year>
          :
          <fpage>1569</fpage>
          -
          <lpage>1578</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>