<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Approach to Trade-of Privacy and Classification Accuracy in Machine Learning Processes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Loredana Caruccio</string-name>
          <email>lcaruccio@unisa.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domenico Desiato</string-name>
          <email>domenico.desiato@uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Polese</string-name>
          <email>gpolese@unisa.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Genovefa Tortora</string-name>
          <email>tortora@unisa.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Zannone</string-name>
          <email>n.zannone@tue.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Bari Aldo Moro</institution>
          ,
          <addr-line>via Edoardo Orabona n.4, 70125 Bari (BA)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, University of Salerno</institution>
          ,
          <addr-line>via Giovanni Paolo II n.132, 84084 Fisciano (SA)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Eindhoven University of Technology</institution>
          ,
          <addr-line>Eindhoven</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Machine learning techniques applied to large and distributed data archives might result in the disclosure of sensitive information. Data often contain sensitive identifiable information, and even if these are protected, the excessive processing capabilities of current machine learning techniques might facilitate the identification of individuals. This discussion paper presents a decision-support framework for data anonymization. The latter relies on a novel approach that exploits data correlations, expressed in terms of relaxed functional dependencies (rfds), to identify data anonymization strategies for providing suitable trade-ofs between privacy and data utility. It also permits to generate anonymization strategies leveraging multiple data correlations simultaneously to increase the utility of anonymized datasets. In addition, our framework provides support in the selection of the anonymization strategies by enabling an understanding of the trade-ofs between privacy and data utility ofered by the obtained strategies. Experiments on real-life datasets show that our approach achieves promising results in data utility while guaranteeing the desired privacy level. Additionally, it allows data owners to select anonymization strategies balancing their privacy and data utility requirements.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Privacy preserving machine learning</kwd>
        <kwd>k-anonymity</kwd>
        <kwd>Relaxed functional dependencies</kwd>
        <kwd>Generalization strategies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The increasing amounts of data available together with the advances in information technology
have brought several benefits and opened new opportunities for the industry, individuals, and
society. In particular, Big Data analytics has enabled the development of increasingly sophisticated
applications ranging from personalized medicine and e-commerce to crowd management and
fraud detection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, these applications have also introduced new privacy and ethical
challenges [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Big Data typically holds large amounts of personally identifiable information
(e.g., criminal records, shopping habits, credit and medical history, and driving records), which
can enable mass surveillance and profiling programs and raise several privacy issues [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>To prevent these issues arising, data protection and privacy frameworks usually define strict
requirements on the collection and processing of personally identifiable information. For instance,
the General Data Protection Regulation (GDPR)1 requires organizations to collect, process, and
share personal data only for legitimate and lawful purposes, and to periodically identify privacy
risks that can afect the data subjects.</p>
      <p>
        Employing all the measures and procedures for the protection of personally identifiable
information, as required by data protection regulations and, especially, by the GDPR, can be
expensive for organizations. Thus, many organizations need to ensure that the personal data they
collect for data analytics are suficiently anonymized to reduce the associated compliance burdens
[
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].2 To this end, they often eliminate any unique identifier for each user when collecting
personal data. However, this in itself may not solve the problem, since removing unique identifiers
might not be suficient to guarantee data anonymity [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In fact, anonymized data could be
de-anonymized through cross-referencing with data gathered from other sources [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Moreover, the
application of machine learning techniques to anonymized data might still lead to the disclosure of
sensitive and confidential information about data subjects, thanks to the power of current predictive
models. On the other hand, we might still want to enable machine learning and data analytics
processes to extract useful knowledge and insights from data while avoiding the disclosure of
sensitive information. Thus, the challenge is to devise anonymization techniques that do not allow
re-identification of individuals by using machine learning techniques on anonymized data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Related work. Several techniques relying on cryptography, randomization, and perturbation,
have been proposed to anonymize data in data sharing and analytics contexts [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. In our proposal,
we focus on anonymization techniques based on generalization. The latter consists of replacing
attribute values with more generalized ones to make the records in a dataset indistinguishable
from each other [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. When generalizing data to protect the privacy of individual records,
some information is lost, which can impact the data utility for further analysis [
        <xref ref-type="bibr" rid="ref6 ref7">7, 6</xref>
        ]. Current
solutions usually suggest approaches to meet anonymity requirements while minimizing loss of
information or trade-of between privacy and data utility requirements [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. However,
generalization strategies often fail to consider correlations in the data, resulting in excessive
penalties to data utility.
      </p>
      <p>
        Contribution. In this discussion paper, we describe the data anonymization framework proposed
in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In detail, such framework exploits (multiple) data correlations, represented as relaxed
functional dependencies (rfds) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], in order to define generalization strategies that guarantee the
required level of privacy and supports the entity responsible for the anonymization of the data
(e.g., the data owner) in balancing privacy and data utility requirements.
      </p>
      <p>The remainder of the paper is organized as follows. Section 2 presents the problem statement
and Section 3 describes the proposed approach. Section 4 presents experimental results, whereas
Section 5 concludes the paper and provides directions for future work.
1GDPR - Final version URL: http://data.consilium.1125europa.eu/doc/document/ST-5419-2016-INIT/en/pdf.
2Notice that the principles of the GDPR do not apply to anonymized information, i.e., information from which the data
subject is no longer identifiable.</p>
      <p>age workclass fnlwgt education
81976543120 34543355338929773098 PPSPPPPSSPtrrrrrrreeaiiiiiiillvvvvvvvfftaaaaaaa--eeettttttt-eeeeeeemmgoppv--nnoott--iinncc 412121287255068533717990494355846154711614488421649272916 91BBBBHHMMt1aaaaSShaatcccch--ssgghhhhtteerreeeeaarrllllssooooddrrrrssss
maritial-status occupation relationship sex capital-gain classes
Never-married Adm-clerical Not-in-family Male 2174 &gt;50K
Married-civ-spouse Exec-managerial Husband Male 0 &gt;50K
Divorced Handlers-cleaners Not-in-family Male 0 &lt;=50K
Married-civ-spouse Handlers-cleaners Husband Male 0 &lt;=50K
Married-civ-spouse Prof-specialty Wife Female 0 &gt;50K
Married-civ-spouse Exec-managerial Wife Female 0 &lt;=50K
Married-spouse-absent Other-service Not-in-family Female 0 &gt;50K
Married-civ-spouse Exec-managerial Husband Male 0 &lt;=50K
Never-married Prof-specialty Not-in-family Female 14084 &gt;50K</p>
      <p>Married-civ-spouse Exec-managerial Husband Male 5178 &gt;50K</p>
    </sec>
    <sec id="sec-2">
      <title>2. Problem statement</title>
      <p>
        Classification models capture correlations between the attributes of individuals and a class value,
and are often used to predict the class value for any unseen new observation. Classification models
are built from a training dataset, which might contain sensitive information. This information
could be inferred from the classification model by exploiting the correlations encoded in the model
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. To this end, training data are usually anonymized by removing identifiable information
before the classifier is trained. However, data can still be re-identified using quasi-identifiers [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Example 1. Let us consider the sample dataset in Table 1, which is extracted from the Adult
dataset3. Each tuple describes an individual, where age, workclass, fnlwgt, education,
maritial-status, occupation, relationship, sex, and capital gain are attributes
characterizing her, whereas attribute classes indicates whether her annual income is greater or lower
than 50. From this sample dataset it is possible to narrow down tuple 1 to a specific individual
by looking, for instance, at the age attribute, as this is the only tuple for which age is equal to 39.
      </p>
      <p>This simple example shows that only removing identifiable information from a dataset might
not be suficient to guarantee anonymization. Anonymized data can be re-identified by linking
the data by means of other data sources [18]. An anonymization model largely used for this
purpose is -anonymity [19], which requires that at least  individuals in the dataset share the
same set of attribute values.</p>
      <p>In this work, we propose a novel anonymization technique that uses generalization and
anonymity validation to anonymize a dataset while minimizing the loss of data utility. To this end,
we exploit data correlations in the dataset, expressed in terms of relaxed functional dependencies
(rfds), as a guideline to define suitable generalization strategies.</p>
    </sec>
    <sec id="sec-3">
      <title>3. A decision-support framework for data anonymization</title>
      <p>
        This section presents a decision-support framework for data anonymization we propose in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In
detail, given an input dataset and a taxonomy of its quasi-identifiers, we first extract generalization
rules expressed in terms of rfds and use them to determine which attributes should be generalized
and at which level. To assess the quality of a generalization rule, we first apply it to the input
dataset to replace attribute values with more general ones, and then compute the anonymity level
and the data utility for the resulting generalized dataset. In a second step, we extend the coverage
of the rfds that satisfy a given level of anonymity by joining generalization rules to increase
data utility. The data anonymization and utility provided by the obtained extended rfds are then
assessed as in the previous step. The obtained generalization rules provide data owners with a
view of which generalization rules can be used to anonymize their datasets and their efects in
terms of data utility and anonymization.
      </p>
      <sec id="sec-3-1">
        <title>3.1. Generalization rule extraction</title>
        <p>The first phase of our approach aims to extract generalization rules in terms of rfds and to
determine the level of anonymity and data utility they achieve when applied on a dataset. rfds
are extracted from the input dataset, along with the generalization levels (defined with respect to
the given attribute taxonomies), by integrating the semantics of roll-up dependencies [20].
Definition 1. (Roll-up dependency) . Let  be a genschema of a relation schema  and
,  ⊆ (), a roll-up dependency (rud) Φ1 → Φ2 is valid on an instance  of , if and
only if for each tuple pair (1, 2) of , if Π  (1) and Π  (2) are  -equivalent, then also Π  (1)
and Π  (2) must be  -equivalent, where 1, 2 are said to be  -equivalent if they become equal
after rolling up their attribute values at most as many levels as the ones specified by  .</p>
        <p>During rfd extraction, we only consider rfds having the classification attribute (i.e., attribute
classes in the example dataset of Table 1) on the right-hand side, with generalization level
equal to 0. This is because we are interested in the generation of anonymized datasets that can
be used to train a classification model. Accordingly, our focus is on correlations involving the
classification attribute and preserving its original values.</p>
        <p>Example 2. Suppose that the following rfd is extracted from the dataset of Table 1:
age≤ 3, fnlwgt≤ 2 → classes≤ 0. The right-hand side of the rfd contains the
classification attribute classes, whereas the left-hand side contains the subset of attributes age and
fnlwgt to be generalized. The generalization level is defined by the values after the tag “ ≤ ”.</p>
        <p>
          Moreover, data owners are left with the task to determine which generalization rules should
be used for the anonymization of their datasets. This can be a complex task, as a large number of
rfds can be potentially extracted from the dataset itself [
          <xref ref-type="bibr" rid="ref4">21, 4</xref>
          ], and not all of them might satisfy
the desired level of anonymity. In addition, rfds usually capture basic correlations in the data,
involving a limited number of attributes and, thus, limiting the data utility that can be achieved
from their application. Increasing the number of attributes on the left-hand side of an rfd will
make it possible to involve more attributes in the anonymization of the dataset, and thus, increase
its data utility [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. However, the use of more attributes could reduce the level of anonymity
guaranteed by the generalization rules. Therefore, the data utility can be improved only where,
and to the extent that, the minimum level of anonymity required by the data owner is satisfied.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Generalization rule selection and improvement</title>
        <p>This phase of the approach aims to generate a set of candidate generalization rules from the rfds
derived in the previous phase of the approach (cf. Section 3.1), which satisfy at least a given level
of anonymity and, at the same time, limit the data utility loss due to the anonymization process.</p>
        <p>rfds may not guarantee a level of anonymity that is acceptable for the data owner. In particular,
the data owner might define minimum anonymization requirements for a dataset to be shared with
other parties. According to the -anonymity model, we model these requirements as a user-defined
threshold , indicating the minimum anonymity level that the dataset should satisfy in order to be
considered for sharing. We use the threshold  to determine whether an rfd provides a suficient
level of anonymity. To check if an rfd is suitable for anonymization, the rfd is applied to the
original dataset and the anonymity level  of the obtained anonymized dataset is computed using
the -anonymity model. If the anonymity level  of the obtained anonymized dataset is equal
or greater than the user-defined threshold , then the rfd satisfies the minimum anonymization
requirements, and it is considered in the anonymization process; otherwise, the rfd is discarded.</p>
        <p>The rfds capture only basic correlations in the data, hence limiting the data utility that can
be achieved through their application. To this end, we analyze the attributes involved in the rfds
and define a coverage strategy to increase the number of selected attributes to be used for the
anonymization of the dataset. Our strategy compares the rfds and determines which ones can be
combined to improve data utility. The intuition is that joining rfds allows to account for multiple
data correlations simultaneously, hence increasing the number of attributes that can be used.
Since combined rfds have to be valid on the considered dataset, not all rfds can be combined.</p>
        <p>We introduce the notion of compatible rfds, which specifies when two rfds can be joined.
Intuitively, two rfds are compatible if and only if their left-hand side attributes are disjoint or
occur with the same generalization level, as formalized in Definition 2.</p>
        <p>Definition 2 (rfd Compatibility). Let Φ → ≤ 0 and Φ′′ → ≤ 0 be two rfds such that
 = {1, . . . , }, ′ = {1, . . . , }, and each attribute  ( ) is associated with a
generalization level  (′ ) in Φ (Φ ′). We say that the two rfds are compatible if and only if:
•  ∩ ′ = ∅, or
• ∀ ∈  and  ∈ ′, such that  =  ∈  ∩ ′, then  = ′ .</p>
        <p>Notice that, in our approach we consider possible generalization rules according by guaranteeing
the following constraints: (i) a set of rfds can be combined into a new rfd if and only if each rfd
is compatible with the others; and (ii) if the combination of two rfds does not meet the minimum
anonymity level, any combination of rfds that includes those rfds will not satisfy the minimum
anonymity level, and hence, it will be discarded. Identifying the optimal candidate rules can be
seen as a multi-objective optimization problem and, thus, we use the notion of Pareto-optimality
and Pareto frontier [22] to guide the data owner in the selection of suitable generalization rules.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>We performed a number of experiments to evaluate the approach proposed in Section 3. In
particular, we studied whether joining rfds and, thus, accounting for a larger set of attributes,
result in anonymization strategies that allow to obtain anonymized datasets with higher data
utility. Moreover, we investigated the trade-of between anonymization and data utility that can be
achieved by using generalization rules and how to devise strategies for selecting the generalization
rules to be used for data anonymization. More specifically, our experiments were driven by the
following research questions:
RQ1: What is the impact of combining generalization rules on data utility?
RQ2: Which trade-of between privacy and data utility can be achieved by using generalization rules?
RQ3: How much efort is required by a data owner to identify the generalization rule to apply?</p>
      <p>The first research question ( RQ1) aims to test our hypothesis and provide insights on the impact
that combined generalization rules produce on the data utility. RQ2 aims to assess the trade-of
between anonymization and data utility that can be achieved using generalization rules. RQ3
aims to evaluate the efort required to a data owner to determine the generalization rule to apply
for the anonymization of her dataset, in terms of the number of rules returned by our approach.</p>
      <sec id="sec-4-1">
        <title>4.1. Results</title>
        <p>
          In this section, we present the results of experiments and answer our research questions. In
particular, to determine the anonymity level ofered by a generalization rule, we apply the
generalization rule to the original dataset and compute the minimum number of tuples in the
generalized dataset that are indistinguishable with respect to the quasi-identifiers, representing the
-anonymity level that the generalized dataset can guarantee by applying such a generalization
rule. Moreover, data utility is measured in terms of classification accuracy and information gain.
In the experiment, classification accuracy was computed using both the decision tree and the
Support Vector Machine classifiers, whereas information gain was computed by using the decision
tree classifier only. The remainder of this section discusses only a portion of the obtained results,
whose complete version can be found in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>RQ1: What is the impact of combining generalization rules on data utility? This research
question aims to evaluate the benefits of combining generalization rules, represented through rfds,
to generate strategies for data anonymization, which maximize data utility while guaranteeing
a desired level of privacy. We expected that, on average, the combination of rfds provides
generalization rules with higher data utility compared to those directly extracted from the data.
To measure this, we compare such sets of rules in terms of classification accuracy and information
gain. In the analysis, we consider all generalization rules that achieve an anonymity level of at
least 2 (i.e.,  ≥ 2). Figures 1 show the accuracy that can be achieved using the generalization
rules directly extracted from the data (red boxes) and using the combined rules (blue boxes) at
the varying of sampling percentage for the Electricity dataset4. We can observe that, combining
generalization rules improves the accuracy (obtained using the ID3 decision tree classifier) for
all sampling percentages, except for the 50% sampling percentage for the Electricity dataset. This
is because many generalization rules extracted for this sampling percentage contain the same
attributes with diferent generalization levels and, thus, they are incompatible or their combination
4https://datahub.io/machine-learning/electricity
50%</p>
        <p>20% 10%
101</p>
        <p>Anonymity level (log scale1)03
102</p>
        <p>104
Sampling 5%
Sampling 10%
Sampling 20%
Sampling 50%
violated the privacy requirement over  (i.e.,  &lt; 2). Our experiments also show that combining
generalization rules improves information gain for the Electricity dataset, as illustrated in Figures 2.</p>
        <p>RQ2: Which trade-of between privacy and data utility can be achieved using
generalization rules? We expect that data utility decreases when the anonymity level increases.
This is because achieving a higher level of anonymity requires higher generalization levels, leading
to less specificity of data. To understand which trade-of between privacy and data utility can be
achieved, we quantify these efects by showing how accuracy and information gain vary when the
anonymity level increases. Figures 3 show the trade-of between accuracy and anonymity level for
the Electricity dataset. The -axis reports the anonymity levels (in log scale), whereas the -axis
reports the best accuracy that can be achieved by applying the generalization rules that satisfy
a given anonymity level. The baseline accuracy is obtained over the non-anonymized version
of the datasets. Each vertical dashed line in the plots represents the maximum anonymity level
that can be achieved using a given sampling percentage (5%, 10%, 20%, and 50%). As expected,
we can observe that, for the Electricity dataset, the accuracy decreases when the anonymity
level increases, and that the highest anonymity level is achieved for the 5% sampling. Figure
4 shows the trade-of between information gain and anonymity level for the Electricity dataset.
Similarly to the results obtained for accuracy, information gain decreases when the anonymity
level increases, and the highest anonymity level is achieved for the 5% sampling.</p>
        <p>RQ3: How much efort is required by a data owner to identify the generalization rule
to apply? A large number of generalization rules can be potentially returned by our approach,
leaving the data owner with the burden to identify which generalization rule should be applied. To
assist the data owner in this task, we employed an approach based on Pareto-optimality to identify
those rules providing a suitable trade-of between privacy and data utility. Next, we evaluate such
approach and, in general, the efort required to a data owner to determine the generalization rule to
apply, in terms of the number of rules returned by our approach. Figure 5 reports the total number
of rfds obtained using our approach at the increase of the anonymity level for each sampling
percentage, before (left plot) and after (right plot) the application of Pareto-optimality, for the
Electricity dataset. We can observe that the sampling percentage has a large impact on the number
of rules, the use of lower sampling percentages typically results in a larger number of generalization
rules. An in-depth analysis (not reported here for lack of space) shows that the number of combined
generalization rules is also higher for lower sampling percentages. This is mainly due to the
fact that generalization rules obtained for lower sampling percentages typically involve few
attributes, yielding many possibilities to combine them with each other. The results also show
that the application of Pareto-optimality significantly reduces the number of generalization rules
to be considered by data owners when anonymizing their datasets. For example, the use of
Pareto-optimality yields a reduction of the total number of generalization rules, which achieve
at least an anonymity level equal to 2, between 60% and 63% for the Electricity dataset where
the largest reduction is obtained for the 5% sampling. When deriving rules using low sampling
percentages, Pareto-optimality tends to preserve more combined rules than rules directly extracted
from the data, whereas this consideration is reversed when the sampling percentage increases.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This work presents a decision-support framework for data anonymization with application
to machine learning processes. The approach extracts rfds from the data to define possible
generalization rules and combine them to derive anonymization strategies guaranteeing a higher
data utility. Pareto-optimality is then employed to identify those generalization rules that provide
optimal trade-ofs between privacy and data utility. Results show that the proposed approach
enables a data owner to identify efective anonymization strategies.</p>
      <p>In the future, we plan to investigate the application of other data utility and privacy metrics
and study their impact on the trade-of between anonymization and data utility, as well as the
applicability of our approach to other data sharing contexts [23].</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This Publication was produced with the co-funding of the European union - Next Generation EU:
NRRP Initiative, Mission 4, Component 2, Investment 1.3 – Partnerships extended to universities,
research centers, companies and research D.D. MUR n. 341 del 5.03.2022 – Next Generation EU
(PE0000014 - "Security and Rights In the CyberSpace - SERICS" - CUP: H93C22000620001).
comprehensive survey, IEEE Access 9 (2021) 8512–8545.
[18] H. Goldstein, N. Shlomo, A probabilistic procedure for anonymisation, for assessing the risk
of re-identification and for the analysis of perturbed data sets, Journal of Oficial Statistics
36 (2020) 89–115.
[19] P. Samarati, L. Sweeney, Generalizing data to provide anonymity when disclosing
information, in: Symposium on Principles of Database Systems, ACM, 1998, p. 188.
[20] T. Calders, R. T. Ng, J. Wĳsen, Searching for dependencies at multiple abstraction levels,</p>
      <p>ACM Transactions Database Systems 27 (2002) 229–260.
[21] L. Caruccio, V. Deufemia, G. Polese, Mining relaxed functional dependencies from data,</p>
      <p>Data Mining and Knowledge Discovery 34 (2020) 443–477.
[22] S. Petchrompo, D. W. Coit, A. Brintrup, A. Wannakrairot, A. K. Parlikad, A review of pareto
pruning methods for multi-objective optimization, Computers &amp; Industrial Engineering 167
(2022) 108022.
[23] J. Feng, L. T. Yang, N. J. Gati, X. Xie, B. S. Gavuna, Privacy-preserving computation in
cyber-physical-social systems: A survey of the state-of-the-art and perspectives, Information
Sciences 527 (2020) 341–355.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Meĳaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C. M.</given-names>
            <surname>Cappers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G. M.</given-names>
            <surname>Mengerink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zannone</surname>
          </string-name>
          ,
          <article-title>Predictive analytics to prevent voice over IP international revenue sharing fraud, in: Data and Applications Security</article-title>
          and
          <string-name>
            <surname>Privacy</surname>
            <given-names>XXXIV</given-names>
          </string-name>
          , volume
          <volume>12122</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>241</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rathore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Loia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-S.</given-names>
            <surname>Jeong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <article-title>Social network security: Issues, challenges, threats, and solutions</article-title>
          ,
          <source>Information sciences 421</source>
          (
          <year>2017</year>
          )
          <fpage>43</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Caruccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Piazza</surname>
          </string-name>
          , G. Polese, G. Tortora,
          <article-title>Secure IoT analytics for fast deterioration detection in emergency rooms</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>215343</fpage>
          -
          <lpage>215354</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Caruccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Desiato</surname>
          </string-name>
          , G. Polese,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Tortora, GDPR compliant information confidentiality preservation in big data processing</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>205034</fpage>
          -
          <lpage>205050</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Desiato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Tortora</surname>
          </string-name>
          ,
          <article-title>A methodology for gdpr compliant data processing</article-title>
          ., in: SEBD,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zigomitros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Casino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Solanas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Patsakis</surname>
          </string-name>
          ,
          <article-title>A survey on privacy properties for data publishing of relational data</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>51071</fpage>
          -
          <lpage>51099</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Cang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gope</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Min, Data anonymization evaluation for big data and IoT environment</article-title>
          ,
          <source>Information Sciences 605</source>
          (
          <year>2022</year>
          )
          <fpage>381</fpage>
          -
          <lpage>392</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Kuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Privacy preservation for machine learning training and classification based on homomorphic encryption schemes</article-title>
          ,
          <source>Information Sciences 526</source>
          (
          <year>2020</year>
          )
          <fpage>166</fpage>
          -
          <lpage>179</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>M. I. Pramanik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Y.</given-names>
            <surname>Lau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Hossain</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Rahoman</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          <string-name>
            <surname>Debnath</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Rashed</surname>
            ,
            <given-names>M. Z.</given-names>
          </string-name>
          <string-name>
            <surname>Uddin</surname>
          </string-name>
          ,
          <article-title>Privacy preserving big data analytics: A critical analysis of state-of-the-</article-title>
          <string-name>
            <surname>art</surname>
          </string-name>
          ,
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <article-title>e1387</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sweeney</surname>
          </string-name>
          ,
          <article-title>Achieving k-anonymity privacy protection using generalization and suppression</article-title>
          ,
          <source>International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems</source>
          <volume>10</volume>
          (
          <year>2002</year>
          )
          <fpage>571</fpage>
          -
          <lpage>588</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ashkouti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khamforoosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheikhahmadi</surname>
          </string-name>
          , DI-Mondrian:
          <article-title>Distributed improved mondrian for satisfaction of the l-diversity privacy model using apache spark</article-title>
          ,
          <source>Information Sciences 546</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>T. K. Esmeel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Hasan</surname>
            ,
            <given-names>M. N.</given-names>
          </string-name>
          <string-name>
            <surname>Kabir</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Firdaus</surname>
          </string-name>
          ,
          <article-title>Balancing data utility versus information loss in data-privacy protection using k-anonymity</article-title>
          ,
          <source>in: Conference on Systems, Process and Control</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>158</fpage>
          -
          <lpage>161</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-C. Chang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>Privacy-preserving high-dimensional data publishing for classification</article-title>
          ,
          <source>Computers &amp; Security</source>
          <volume>93</volume>
          (
          <year>2020</year>
          )
          <fpage>101785</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>K. LeFevre</surname>
            ,
            <given-names>D. J. DeWitt</given-names>
          </string-name>
          , R. Ramakrishnan,
          <article-title>Workload-aware anonymization</article-title>
          ,
          <source>in: Proceedings of SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>277</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Caruccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Desiato</surname>
          </string-name>
          , G. Polese, G. Tortora,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zannone</surname>
          </string-name>
          ,
          <article-title>A decision-support framework for data anonymization with application to machine learning processes</article-title>
          ,
          <source>Information Sciences 613</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Caruccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Deufemia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          , G. Polese,
          <article-title>Discovering relaxed functional dependencies based on multi-attribute dominance</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>33</volume>
          (
          <year>2021</year>
          )
          <fpage>3212</fpage>
          -
          <lpage>3228</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Majeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Anonymization techniques for privacy preserving data publishing: A</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>