<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mining and Visualizing Uncertain Data Ob jects and Named Data Networking Tracs by Fuzzy Self-Organizing Map</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amin Karami</string-name>
          <email>amin@ac.upc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manel Guerrero-Zapata</string-name>
          <email>guerrero@ac.upc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Architecture Department (DAC), Universitat Polit`ecnica de Catalunya (UPC)</institution>
          ,
          <addr-line>Campus Nord, C. Jordi Girona 1-3. 08034 Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Uncertainty is widely spread in real-world data. Uncertain data -in computer science- is typically found in the area of sensor networks where the sensors sense the environment with certain error. Mining and visualizing uncertain data is one of the new challenges that face uncertain databases. This paper presents a new intelligent hybrid algorithm that applies fuzzy set theory into the context of the Self-Organizing Map to mine and visualize uncertain objects. The algorithm is tested in some benchmark problems and the uncertain tracs in Named Data Networking (NDN). Experimental results indicate that the proposed algorithm is precise and e↵ective in terms of the applied performance criteria.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Uncertainty is a frequent issue in data analysis. The various factors that lead to
data uncertainty include: approximate measurement, data sampling fault,
transmission error or latency, data integration with noise, data acquisition by device
error, and so on [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These factors produce vague and imprecise data.
Visualizing uncertain data is one of the new challenges in the uncertain databases [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Among the many visualization techniques, the Self-Organizing Map (SOM) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
is widely and successfully applied due to its good result. SOM is a very popular
unsupervised learning algorithm based on the classical set theory. An important
application of SOM is discovering the topological relationship among
multidimensional input vectors and mapping them to a low dimensional output which
is easy for further analysis by experts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The process of SOM training
requires a certain and an unambiguous input data either belongs or not belong to a
weight vector (cluster), where the membership evaluation is boolean. In contrast,
uncertain and vague input vectors are not either entirely belong or not belong
to a weight vector. A data may be considered vague and imprecise where some
things are not either entirely true nor entirely false and where the some things
are somehow ambiguous. For instance, fuzzy location in the right side of Fig. 1 is
a way to represent the item of vague information: the object is approximately at
position (4, 3), in which the grey levels indicate membership values with white
representing 0 and black representing 1. In contrast, the left side of Fig. 1 shows
the exact position of a certain data where the membership evaluation of
centers (weights) is boolean. There has been a lot of research in the application
of Fuzzy sets theory to model vague and uncertain information [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The Fuzzy
set (FS) theory introduced by Zadeh [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is a more flexible approach than
classical set theory, where objects belong to sets (clusters) with certain degree of
membership ranging [0..1]. This makes FS theory suitable for representing and
visualizing uncertain data [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Therefore, a combination of SOM and FS is able
to illustrate dependencies in the uncertain data sets in a very intuitive manner.
SOM is indeed originally intended as a classification method, not a visualization
method so there are a few additions to apply SOM for visualization. Li et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
proposed a mining and visualizing algorithm for uncertain data, called USOM
which combines fuzzy distance function and SOM. In this paper, we employ the
FS theory in the context of SOM algorithm to mine and visualize the
uncertain objects in the uncertain databases. Experimental results over four classic
benchmark problems and a new network architecture as Named Data
Networking (NDN) show that the proposed method outperforms standalone SOM and
USOM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in terms of the applied performance metrics. The remainder of the
paper is organized as follows: Section 2 presents self-organizing map. Section 3
presents our contribution. Section 4 evaluates the new approach experimentally.
Section 5 is the conclusion and future work.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Self-Organizing Map (SOM)</title>
      <p>
        SOM (also known as Kohonen SOM) is a very popular algorithm based on
competitive and unsupervised learning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The SOM projects and represents higher
dimensional data in a lower dimension, typically 2-D, while preserving the
relationships among the input data. The main process of SOM is generally
introduced in three main phases: competition, cooperation and adaptation which are
described in detail in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>The Proposed Method</title>
      <p>
        The procedure of the proposed method, application of fuzzy set theory in the
context of SOM for mining and visualizing uncertainties is as follows. A diagram
of the proposed method is shown in Fig. 2.
1. Fuzzy competition: in hard competition, the input vector is divided into
distinct weights (clusters), where each input element belongs to exactly one
weight. In fuzzy competition, input vector can belong to more than one
weight, and associated with each element by a set of membership levels.
Fuzzy c-means (FCM) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] method allows one piece of input data to belong
to two or more clusters (weights). The standard function is:
      </p>
      <p>Ux =
1</p>
      <p>2</p>
      <p>Pj ( dd((wweeiigghhttkj,,xx)) ) m 1
Where, Ux is the membership value of each input vector x to all weights,
j = 1, 2, ..., w, and m is the level of cluster fuzziness which is commonly set
to 2. By the fuzzy competition all the neurons are wining neurons with the
membership degree ranging [0..1].
2. Fuzzy cooperation: in fuzzy cooperation, all wining neurons cooperate with
their neighboring neurons in terms of the membership degree by Eq. 2. For
the size of the neighborhood, we employed the Gaussian function that shrinks
on each iteration until eventually the neighborhood is just the BMU itself.</p>
      <p>h(j, i) = Uxi ⇥ exp( 2dj,2i2 ) i, j = 1, 2, .., n; i 6= j
Where, i is the number of the wining neurons including all the neurons with
di↵erent membership degrees, j is the number of the cooperating neighbor
neurons. Uxi is the membership value of input vector x from ith wining
neuron. h(j, i) is the topological area centered around the wining neuron
i and the cooperating neuron j. The size of the neighborhood needs to
decrease with time. A popular time dependence is an exponential decay by:
(t) = 0exp(
t
)
(1)
(2)
(3)
Where, (t) is the width of the lattice at time t, 0 is the width of the lattice
at time t0, and is the time constant.
3. Fuzzy adaption: the adaption phase is the weight update by:
wj = wj + Uj ⇥ (⌘h (j, i) ⇥ (x
wj )) i, j = 1, 2, .., n; i 6= j
(4)</p>
      <p>Where, Uj is the membership value of input x from neuron j.</p>
      <p>
        These three phases are repeated, until the maximum number of iterations is
reached or the changes become smaller than a predefined threshold.
To assess the accuracy and performance of the proposed method, four classic
benchmark problems from the UCI machine learning repository [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] are applied.
The selected data sets are Iris (4-D), Glass (9-D), Wine (13-D), and Zoo (17-D).
In practice, uncertainties are usually modeled in the form of Gaussian
distributions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For some attributes in data sets, we add a Gaussian noise with a zero
mean and the standard deviation with the normal distribution [0, f ], where, f is
an integer parameter from the set of {1, 2, 3} to define di↵erent uncertain levels.
The experiments on each method were repeated 10 times independently. We
evaluate the several SOM network structures on applied uncertain data sets
which the optimal ones are Iris with 16x16 nodes, Glass with 16x16 nodes, Wine
with 17x17 nodes, and Zoo with 15x15 nodes.
      </p>
      <p>Table 1 shows that our proposed method outperforms SOM and USOM
methods in terms of the Quantization Error (QE) and Topographic Error (TE). The
proposed method seems to be more time consuming (with Exec.) than the other
Data</p>
      <p>K=1
Iris (16x16) 0.937
Glass (16x16) 0.913
Wine (17x17) 0.917
Zoo (15x15) 0.957
methods due to the application of fuzzy set theories in the context of the SOM,
in which all the neurons are winner with di↵erent membership grading. However,
the proposed method can find a better solution with less times of increment on
computational time (with Inc.) than the other methods due to its fast
convergence speed. The trustworthiness and continuity values for K={1, 10, 20} are
shown in Tables 2 and 3, respectively. The trustworthiness and continuity
measures show that the proposed method obtains the better results as compared to
SOM and USOM. The results show that the proposed method with the
application of fuzzy set theory in the context of the SOM yields high accuracy as
compared to other methods without very much computational cost. Since our
proposed method performs well as compared to SOM and USOM, we visualize
uncertainties in the applied uncertain data sets. To facilitate the interpretation
of results, we use the U-Matrix (unified distance matrix) where visualize the
high-dimensional uncertain data into a 2-D space in Fig. 3. In this figure, the
blue hexagons represent the neurons (weights). The darker colors in the regions
between neurons represent larger distance, while the lighter colors represent
smaller distances. Fig. 3(a) shows that the constructed 4-D uncertain Iris SOM
network has been clearly clustered into three distinct groups. The Glass SOM
(a)
(b)
(c)
(d)</p>
      <p>
        [400..1500] Ienxtisetreenstt flcoonotdeinntg uantitfaocrkmslfyoranndone-xepxoisnteennttiaalnlyd
network (Fig. 3(b)) has been apparently classified 9-D uncertain data objects
into six distinct types of glass. Figs. 3(c) and 3(d) show the three and the seven
distinct groups of 13-D and 17-D uncertain data from Wine and Zoo data sets,
respectively. The results confirm that the proposed method performs well in
mining and visualizing uncertain data into somewhat expected distinct groups.
After evaluating the robustness and the accuracy of our proposed method with
some benchmark problems, we apply the proposed method for visualizing
uncertain tracs in Named Data Networking (NDN). NDN [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] is a promising network
architecture being considered as a possible replacement to overcome the
fundamental limitations of the current IP-based Internet. Trac uncertainty refers
to trac volumes belong to more than one pattern, and associated with each
pattern by a set of membership levels. Fuzzy approach can reduce the false
positive rate with higher reliability in identifying the pattern of trac volumes, due
to any uncertain attack data may be similar to some normal patterns [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We
conduct the same testbed configuration from papers [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The employed
features for trac generation come from paper [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] as well as the ratio of (1) cache
hit, (2) dropped Interest packet, (3) dropped data packets, (4) satisfied
Interest packet, and (5) timed-out Interest packets in each 1 sec time interval. The
structure of the trac generated is shown in Table 4. We modeled uncertainties
for some attributes in NDN trac samples in the form of Gaussian distributions
similar to Section 4.1. Fig. 4 maps the 11-D uncertain trac samples to the 2-D
space through our proposed method. This figure shows that the our proposed
method performs somewhat well in mining and visualizing uncertainties into
predefined distinct groups. Fig. 4 illustrates that there are some small groups of
clustered data points with the lighter regions. These small clusters may contain
some normal or attack data that try to be incorrectly placed in the neighboring
regions, due to their uncertain nature. The results in Table 5 show that our
proposed method o↵ers the best performance and outperforms suciently other
preexisting methods.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this paper, we propose a new hybrid algorithm for mining and visualizing
uncertain data. We investigate the implementation of fuzzy set theory in the
design of SOM neural network in order to improve the accuracy of visualizing
uncertain data bases. The experimental results over the uncertain benchmarking
data sets and the uncertain tracs in Named Data Networking show that the
proposed method is e↵ective and precise in terms of the applied performance
criteria. We plan to improve the proposed method for various uncertain models
and big uncertain network trac data in the future.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>This work was partially supported by projects TIN2013-47272-C2-2 and
SGR2014-881.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Pavle Miloˇsevi´c, Bratislav Petrovi´c, Dragan Radojevi´c, and
          <article-title>Darko Kovaˇcevi´c. A software tool for uncertainty modeling using interpolative boolean algebra</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>62</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Jiaqi</given-names>
            <surname>Ge</surname>
          </string-name>
          , Yuni Xia, and
          <string-name>
            <given-names>Yicheng</given-names>
            <surname>Tu</surname>
          </string-name>
          .
          <article-title>A discretization algorithm for uncertain data</article-title>
          .
          <source>In Proceedings of the 21st International Conference on Database and Expert Systems Applications: Part II</source>
          , pages
          <fpage>485</fpage>
          -
          <lpage>499</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Le</given-names>
            <surname>Li</surname>
          </string-name>
          , Xiaohang Zhang, Zhiwen Yu,
          <string-name>
            <given-names>Zijian</given-names>
            <surname>Feng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ruiping</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <article-title>Usom: Mining and visualizing uncertain data based on self-organizing maps</article-title>
          .
          <source>In International Conference on Machine Learning and Cybernetics (ICMLC)</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>804</fpage>
          -
          <lpage>809</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kohonen.</surname>
          </string-name>
          Self-Organizing Maps. Springer, Berlin, Heidelberg,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>Khalilia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mihail</given-names>
            <surname>Popescu</surname>
          </string-name>
          .
          <article-title>Topology preservation in fuzzy selforganizing maps</article-title>
          .
          <source>Advance Trends in Soft Computing</source>
          ,
          <volume>312</volume>
          :
          <fpage>105</fpage>
          -
          <lpage>114</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Juha</given-names>
            <surname>Vesanto</surname>
          </string-name>
          and
          <string-name>
            <given-names>Esa</given-names>
            <surname>Alhoniemi</surname>
          </string-name>
          .
          <article-title>Clustering of the self-organizing map</article-title>
          .
          <source>IEEE Transactions on Neural Networks</source>
          ,
          <volume>11</volume>
          (
          <issue>3</issue>
          ):
          <fpage>586</fpage>
          -
          <lpage>600</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>D.</surname>
          </string-name>
          Herrero-P´erez, H. Martnez-Barber´
          <article-title>a, K. LeBlanc, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Saotti</surname>
          </string-name>
          .
          <article-title>Fuzzy uncertainty modeling for grid based localization of mobile robots</article-title>
          .
          <source>International Journal of Approximate Reasoning</source>
          ,
          <volume>51</volume>
          (
          <issue>8</issue>
          ):
          <fpage>912</fpage>
          -
          <lpage>932</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lotfi</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zadeh</surname>
          </string-name>
          .
          <article-title>Fuzzy sets</article-title>
          .
          <source>Information Control</source>
          ,
          <volume>8</volume>
          :
          <fpage>338</fpage>
          -
          <lpage>353</lpage>
          ,
          <year>1965</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Feng</given-names>
            <surname>Qi</surname>
          </string-name>
          and
          <string-name>
            <surname>A-Xing Zhu</surname>
          </string-name>
          .
          <article-title>Comparing three methods for modeling the uncertainty in knowledge discovery from area-class soil maps</article-title>
          .
          <source>Computers &amp; Geosciences</source>
          ,
          <volume>37</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1425</fpage>
          -
          <lpage>1436</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>J. C.</surname>
          </string-name>
          <article-title>Bezdek. Pattern Recognition with Fuzzy Objective Function Algoritms</article-title>
          . Plenum Press, New York,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A.</given-names>
            <surname>Asuncion</surname>
          </string-name>
          and
          <string-name>
            <surname>D.J. Newman.</surname>
          </string-name>
          <article-title>UCI machine learning repository</article-title>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Gonzalo</surname>
          </string-name>
          <article-title>A. Ruz and Duc Truong Pham</article-title>
          .
          <article-title>Nbsom: The naive bayes self-organizing map</article-title>
          .
          <source>Neural Comput. Appl.</source>
          ,
          <volume>21</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1319</fpage>
          -
          <lpage>1330</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Diana K. Smetters James D. Thornton Michael F. Plass Nicholas H. Briggs Jacobson</surname>
            , Van and
            <given-names>Rebecca L.</given-names>
          </string-name>
          <string-name>
            <surname>Braynard</surname>
          </string-name>
          .
          <article-title>Networking named content</article-title>
          .
          <source>In In Proceedings of the 5th international conference on Emerging networking experiments and technologies</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>Amin</given-names>
            <surname>Karami</surname>
          </string-name>
          and
          <string-name>
            <given-names>Manel</given-names>
            <surname>Guerrero-Zapata</surname>
          </string-name>
          .
          <article-title>A fuzzy anomaly detection system based on hybrid pso-kmeans algorithm in content-centric networks</article-title>
          .
          <source>Neurocomputing</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>Amin</given-names>
            <surname>Karami</surname>
          </string-name>
          .
          <article-title>Data clustering for anomaly detection in content-centric networks</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>81</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>