<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Pizzo Calabro (VV),
Italy
" elena.battaglia@unito.it (E. Battaglia); ruggero.pensa@unito.it (R. G. Pensa)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>DP-DILCA: Learning Diferentially Private Context-based Distances for Categorical Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Battaglia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruggero G. Pensa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Turin</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Distance-based machine learning methods have limited applicability to categorical data, since they do not capture the complexity of the relationships among diferent values of a categorical attribute. Nonetheless, categorical attributes are common in many application scenarios, including clinical and health records, census and survey data. Although distance learning algorithms exist for categorical data, they may disclose private information about individual records if applied to a secret dataset. To address this problem, we introduce a diferentially private algorithm for learning distances between any pair of values of a categorical attribute according to the way they are co-distributed with the values of other categorical attributes forming the so-called context. We show empirically that our approach consumes little privacy budget while providing accurate distances.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;diferential privacy</kwd>
        <kwd>metric learning</kwd>
        <kwd>categorical attributes</kwd>
        <kwd>distance-based methods</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Most machine learning and data analysis methods rely, directly or indirectly, on their ability
to compute distances or similarities between data objects. Although diferent definitions of
distance/similarity exist, they are relatively easy to compute, provided that data are given in
form of numeric vectors. Additionally, for most of the above-mentioned distance-based methods,
diferentially private counterparts of them have been proposed as well. Diferential privacy [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
is a computational paradigm which guarantees that the output of a statistical query applied
to a secret dataset does not allow to understand whether a particular data object is present in
the dataset or not. In recent years, many diferentially private variants have been proposed for
most distance based algorithms, including kNN [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], SVM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and k-means [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        When data are described by categorical features/attributes, instead, distances can only account
for the match or mismatch of the values of an attribute between two data objects, leading to
poorer and less expressive proximity measures (e.g., the Jaccard similarity). And yet, intuitively,
a patient whose disease is “gastritis” should be closer to a patient afected by “ulcer” than to
one having “migraine”. An eficient solution consists in using some distance learning algorithm
to infer the distance between any pair of diferent values of the same categorical attribute from
data. Among all existing methods, DILCA [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is one of the most efective. DILCA’s objective is
to compute the distance between any pair of values of a categorical attribute by taking into
account the way the two values are co-distributed with respect to the values of other categorical
attributes forming the so-called context. According to DILCA, if two values of a categorical
attribute are similarly distributed w.r.t. the values of the context attributes, then their distance
is lower than that computed for two values of the same attribute that are divergently distributed
w.r.t. the values of the same context attributes. DILCA has been successfully used in diferent
scenarios including clustering [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], semi-supervised learning [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and anomaly detection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
However, if applied to a secret dataset, it may disclose a lot of private information.
      </p>
      <p>In this paper, we address the problem of learning meaningful distances for categorical data in
a diferentially private way. To this purpose, we first introduce a diferentially-private extension
of DILCA adopting the exponential mechanisms. We show experimentally that it provides
accurate distances even with relatively small values of privacy budget . Additionally, we
show that our algorithm (which we call DP-DILCA) is efective in two distance-based learning
scenarios, including clustering and k-NN classification.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>In this section, we introduce the necessary background required to understand the theoretical
foundations of our method and, contextually, we introduce its related scientific literature.</p>
      <sec id="sec-2-1">
        <title>2.1. Diferential Privacy</title>
        <p>
          Diferential privacy [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is a privacy definition that guarantees the outcome of a calculation to
be insensitive to any particular record in the data set. More formally, we report the following
definition [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]:
and ∀ ∈ ℛ, ((ℳℳ((′))==)) ≤ .
        </p>
        <p>Definition 1 ( -diferential privacy). Let ℳ : Ω →− ℛ be a randomized mechanism (i.e. a
stochastic function with values in a generic set ℛ) and consider a real number  &gt; 0. We say that
ℳ preserves -diferential privacy if for all pair of datasets , ′ difering for only one record</p>
        <p>
          Diferential privacy satisfies two important properties: composition and post-processing [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
The composition property states that by combining the results of several diferentially private
mechanisms, the outcome will be diferentially private too, and the overall level  of privacy
guaranteed will be the sum of the level of privacy of each mechanism. On the other hand, the
post-processing property says that once a quantity  has been computed in a diferentially
private way, any following transformation of this quantity is still diferentially private, with no
need to spend part of the privacy budget for it.
        </p>
        <p>
          Several mechanisms and techniques preserving diferential privacy have been proposed in
literature. Two of the most famous mechanisms are the Laplace and the Exponential mechanisms
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. They both calibrate the amount of random noise they inject in the computation by looking
at the sensitivity of the function (or utility function) considered:
Definition 2 (Global sensitivity). Let  : Ω →− R be a numeric function. The global
sensitivity () is a measure of the maximal variation of function  when computed over two datasets
difering for only one record and is defined as () = max∼ ′ ||() − (′)||1.
2.2. DILCA
Measuring similarities or distances between two data objects is a crucial step for many machine
learning and data mining tasks. While the notion of similarity for continuous data is relatively
well-understood and extensively studied, for categorical data the similarity computation is not
straightforward. The simplest comparison measure for categorical data is overlap [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]: given
two tuples, it counts the number of attributes whose values in the two tuples are the same.
The overlap measure does not distinguish diferent values of attributes, hence matches and
mismatches are treated equally. Among all the proposed methods for distance computation,
we focus on DILCA [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], a framework to learn context-based distances between each pair of
values of a categorical attribute  . The main idea behind DILCA is that the distribution of
the co-occurrences of the values of  and the values of the other attributes in the dataset
may help define a distance between the values of  (intuitively, two values that are similarly
co-distributed w.r.t. all the other values of all the other attributes are similar and so they should
be close in the new distance). However, not all the other attributes in the dataset should be
taken in consideration, but only those that are more relevant to  . We call this set of relevant
attributes with respect to  the context of  . DILCA distance is defined as follow
Definition 3 (DILCA distance). Let 1, . . .  be the values of attribute  . For each pair , 
with ,  = 1, . . . , , the distance between  and  is computed as
⎯
⎸ ∑︀
⎸
(,  ) = ⎷
∈( ) ∑︀|=|1( (|) −  ( |))2
        </p>
        <p>∑︀∈( ) ||
where ( ) is the set of the attributes belonging to the context of  , || is the number of
values attribute  can assume, and  (|) is the conditional probability that  takes value 
given that  has value .</p>
        <p>The conditional probabilities  (|) are estimated from the data: the contingency table
between attributes  and  is constructed and this contingency table can be interpreted as the
empirical joint distribution of the two variables.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. DP-DILCA</title>
      <p>
        In this section, we introduce a method whose final goal is to inject some form of randomness
in DILCA in order to make the resulting distances among the values of the target attribute 
diferentially private. There are two moments when DILCA algorithm accesses the original
(secret) dataset: the context and the contingency tables computation phases. If the context
selection is made preserving ℎ · -diferential privacy (where ℎ ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]) and the computation of
all the contingency tables is made preserving (1 − ℎ)-diferential privacy, then the composition
Select an object  ∈ ℱ with probability proportional to 
( ) ←
ℱ ← ℱ ∖ {
};
( ) ∪ {};
︁( ℎ· (,) )︁ ;
2· · 
Result: The set ( )
2 (︁
      </p>
      <p>1
(2) + ( ) ;</p>
      <p>︁)
1, . . . , } ∖ { };
3 ( ) ← ∅
4 for  = 1 to  do</p>
      <p>;
1  ←
2 ℱ ← {
5
6
7
8 end
privacy.</p>
      <p>Algorithm 1:  (, , ℎ · , )</p>
      <p>context
Input: The original dataset  with  records and attributes  = {1, ...., }, the
target attribute  ∈  , the privacy budget ℎ, the number  of attributes in the
and the post-processing theorems guarantee that the overall algorithm preserves -diferential</p>
      <p>
        The context selection procedure used by DILCA is an application of a filter method for
supervised feature selection. Indeed, some work has been done on diferentially private feature
selection. For instance, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] present two alternative diferentially private
implementations of a feature selection method that preserves nearest-neighbor classification capability. In
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], instead, the authors study the sensitivity of several association measures used for feature
selection and integrate the noised version of these measures in two diferentially private classifiers.
Here, we propose a diferentially private selection method that measures the connection of two
attributes by looking at the (distorted) Mutual Information between them and then extracts the 
most relevant attributes. Mutual Information is a widely used measure of association in the
superlent to finding the  which maximizes ′(,  ) = () − (,  ).
vised feature selection problem and it can be computed as (,  ) = ()+( )− (,  )
(where (· ) is the entropy function). Thus, finding the  which maximizes (,  ) is
equiva and  , an upper bound of the sensitivity of ′(,  ) is 2 (︁
1
(2) + ( ) .
      </p>
      <p>︁)
Theorem 3.1 (Sensitivity of ′(,  )). Given a dataset  with  records and two attributes</p>
      <p>Algorithm 1 describes the diferentially private version of context selection. It requires the
specification, as input parameter, of the desired number  of attributes in the context of the target
attribute. When setting the value of parameter , one must consider that lower values of  are
preferable, from a diferentially private point of view. In step 5 of Algorithm 1, the exponential
mechanism is applied  times, in order to extract the top  attributes: each application of the
exponential mechanism requires part of the overall privacy budget; thus, the smaller  is, the
higher the accuracy of the selected context. Algorithm 1 preserves ℎ-diferential privacy.</p>
      <p>Once the context of target attribute  has been selected, DILCA algorithm computes the
contingency table  (,  ) between  and , for each  in context. Instead of the exact
value  (,  ), we compute a distorted contingency table via the Laplace mechanism. This
1.0
0.8
e
r0.6
o
c
-s0.4
F
0.2
mechanism needs the specification of two parameters, the sensitivity of the function  , that
can be proved to be 2, and the privacy budget. Since the total number of needed contingency
tables is  = |( )|, the privacy budget we spend for each contingency table is (1− ℎ) .
In this way, the overall DP-DILCA algorithm preserves -diferential privacy.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>In this section, we describe the experiments conducted to evaluate the performances of
DPDILCA. For this evaluation, we use four real-world datasets: adult1, NLTCS2, IPUMS-BR and
IPUMS-ME3.</p>
      <sec id="sec-4-1">
        <title>4.1. Assessment of context selection</title>
        <p>In the first experiment, we run DP-DILCA on the real-world datasets in order to assess the
quality of the context they select. For each dataset, we consider one attribute at a time as
target attribute and we compute its diferentially private context for increasing levels of privacy
budget . Then we compare the context selected by DP-DILCA with the context obtained with
the corresponding non-private method. In all the experiments we set  = 3. To evaluate the
similarity between the private and non-private context for each target attribute, we use the
F-score, i.e., the harmonic mean of precision and recall. For each , we repeat the experiments 30
times and we compute the mean value of all scores. Figure 1 shows the results of our comparison:
for each  we report the average value of the F-score over all the attributes of each dataset. In
all the datasets, the results achieved by DP-DILCA increase with respect to .</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Assessment of the distance matrices</title>
        <p>In this section we repeat the same experiments on the real-world data presented in Section 4.1,
but we focus on the final output of DP-DILCA: the distances between the values of the target
attribute. As before, for each dataset we consider one attribute at a time as target and we
compute the diferentially private distance matrix associated to its values, for increasing levels
of privacy budget . Then we compare the distances obtained with DP-DILCA with those
1https://archive.ics.uci.edu/
2http://lib.stat.cmu.edu/
3https://international.ipums.org/
obtained with the corresponding non-private method. In this experiment we set  = 3 and
ℎ = 0.3. We quantify the linear correlation between the private distance matrix  ′, with shape
 × , and its non-private counterpart  through the sample Pearson’s  correlation coeficient.
The  coeficient takes values between -1 (perfect negative correlation) and 1 (perfect positive
correlation). If the two matrices are not correlated we will have  ∼= 0.</p>
        <p>For each , we repeat the experiments 30 times and we compute the mean value of the sample
Pearson correlation coeficient. Figure 2 shows the results of our computations: for each  we
report the average value of the measure over all the attributes of each dataset. The results show
that there is positive correlation between private and non-private distances and the Pearson’s
coeficient increases as  grows. Notice that the Pearson coeficient is always 1 when the target
attribute has only two values. For this reason, for NLTCS (Figure 2(b)), which consists of binary
attributes only, the Person’s correlation is always maximum.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Experiments on clustering and classification</title>
        <p>
          In this section, we assess the efectiveness and utility of the distances computed by our
differentially private algorithms. To this purpose, we embed DP-DILCA into two distance-based
learning algorithms: the Ward’s hierarchical clustering algorithm and the kNN classifier. Both
the algorithms take as input the matrix of the pairwise distances between the data objects.
DP-DILCA’s output is the distance between values of a categorical attribute; if it is applied to
all attributes in  , then the distance between any pair of objects ,  , both described by 
can be computed as (,  ) = √︀∑︀∈  [.,  .]2, where  is the distance
matrix returned by DP-DILCA for attribute  and . and  . are the values of attribute 
on objects  and  [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. We will refer to this metric as  . Similarly, we will
call  the metric obtained by the non-private DILCA algorithm.
        </p>
        <p>
          We run the experiment about clustering as follows: for each real-world dataset, we compute
the object distance matrix using the diferent private and non private metrics, then we run
Ward’s hierarchical clustering with these matrices as input. Since the hierarchical algorithm
returns a dendrogram which, at each level, contains a diferent number of clusters, we consider
the level corresponding to the number of clusters equal to the number of classes. We call the
overall clustering models   and , depending on the distance metric
adopted. We evaluate the quality of the results through the adjusted rand index (ARI) computed
w.r.t. the actual classes [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. For this reason we run this experiment on dataset adult only.
Figure 3(a) shows the mean ARI results over 30 experiments. The value of  on the  axis of
the plot is the overall privacy budget used for the learning of the metric, while the privacy
budget spent for computing the distances among values of a single attribute is  . The ARI
values of the clustering model with private distance computation grow with respect to the
privacy budget, and for high values of , they get results close to those of the clustering with
non-private distances.
        </p>
        <p>As last experiment, we run the kNN classification algorithm, with  = 5. We perform a
4-fold cross-validation: one fold is retained as test set, then the metrics  , and
 are learned on the remaining 3 folds and the classification model is trained on
the same set. We call the overall models    and  , depending on the
distance learning algorithm used. For each dataset, we apply the four kNN models 30 times and
compute the mean accuracy of the classification on the the test set. The process is repeated
four times and the results are further averaged on the four test sets. In Figure 3(b) we report
the mean accuracy of all the models for increasing levels of privacy budget . The results of
   are always very close to those of  , even for very low levels of .</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We have introduced a new family of diferentially private algorithms for the data-driven
computation of meaningful and expressive distances between any two values of a categorical attribute.
Our approach is built upon an efective context-based distance learning framework whose
output, however, may reveal private information if applied to a secret dataset. For this reason,
we have proposed a randomized algorithm, based on the Laplace and exponential mechanisms,
that satisfies -diferential privacy and returns accurate distance measures even with relatively
small privacy budget consumption. Additionally, the metric learnt by our approach can be used
profitably in distance-based machine learning algorithms, such as hierarchical clustering and
kNN classification.
This work is supported by Fondazione CRT (grant number 2019-0450).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Dwork</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>The algorithmic foundations of diferential privacy</article-title>
          ,
          <source>Found. Trends Theor. Comput. Sci. 9</source>
          (
          <year>2014</year>
          )
          <fpage>211</fpage>
          -
          <lpage>407</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Gursoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Inan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Nergiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Saygin</surname>
          </string-name>
          ,
          <article-title>Diferentially private nearest neighbor classification, Data Min</article-title>
          .
          <source>Knowl. Discov</source>
          .
          <volume>31</volume>
          (
          <year>2017</year>
          )
          <fpage>1544</fpage>
          -
          <lpage>1575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chaudhuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Monteleoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Sarwate</surname>
          </string-name>
          ,
          <article-title>Diferentially private empirical risk minimization</article-title>
          ,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>1069</fpage>
          -
          <lpage>1109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Bertino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <article-title>Diferentially private k-means clustering and a hybrid approach to private optimization</article-title>
          ,
          <source>ACM Trans. Priv. Secur</source>
          .
          <volume>20</volume>
          (
          <year>2017</year>
          )
          <volume>16</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          :
          <fpage>33</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ienco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Pensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meo</surname>
          </string-name>
          ,
          <article-title>From context to distance: Learning dissimilarity for categorical data clustering</article-title>
          ,
          <source>ACM Trans. Knowl. Discov. Data</source>
          <volume>6</volume>
          (
          <year>2012</year>
          ) 1:
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          :
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ienco</surname>
          </string-name>
          , R. G. Pensa,
          <article-title>Positive and unlabeled learning in categorical data</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>196</volume>
          (
          <year>2016</year>
          )
          <fpage>113</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ienco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Pensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meo</surname>
          </string-name>
          ,
          <article-title>A semisupervised approach to the detection and characterization of outliers in categorical data</article-title>
          ,
          <source>IEEE Trans. Neural Networks Learn. Syst</source>
          .
          <volume>28</volume>
          (
          <year>2017</year>
          )
          <fpage>1017</fpage>
          -
          <lpage>1029</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kasif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Salzberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Waltz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rachlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Aha</surname>
          </string-name>
          ,
          <article-title>A probabilistic framework for memory-based reasoning</article-title>
          ,
          <source>Artif. Intell</source>
          .
          <volume>104</volume>
          (
          <year>1998</year>
          )
          <fpage>287</fpage>
          -
          <lpage>311</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Diferentially private feature selection</article-title>
          ,
          <source>in: Proceedings of IJCNN</source>
          <year>2014</year>
          , IEEE,
          <year>2014</year>
          , pp.
          <fpage>4182</fpage>
          -
          <lpage>4189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <article-title>Local learning-based feature weighting with privacy preservation</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>174</volume>
          (
          <year>2016</year>
          )
          <fpage>1107</fpage>
          -
          <lpage>1115</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Anandan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clifton</surname>
          </string-name>
          ,
          <article-title>Diferentially private feature selection for data mining</article-title>
          ,
          <source>in: Proceedings of ACM IWSPA@CODASPY</source>
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hubert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Arabie</surname>
          </string-name>
          ,
          <article-title>Comparing partitions</article-title>
          ,
          <source>Journal of Classification</source>
          <volume>2</volume>
          (
          <year>1985</year>
          )
          <fpage>193</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>