<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Class Weight based K-NN Classifier (CWK-NN)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohamad Hijazi School of Arts and</string-name>
          <email>MohamadOsamaHijazi@Gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jawad Khalife</string-name>
          <email>jkhalife.khalife@liu.edu.lb</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hussein Al-ghor Faculty</string-name>
          <email>Hussein.ghor@ul.edu.lb</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesus Diaz Verdejo School of</string-name>
          <email>jedv@ugr.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IT and Telecom. Eng. University, of Granada</institution>
          ,
          <addr-line>Granada</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Arts and Sciences, Lebanese International University</institution>
          ,
          <addr-line>Beirut</addr-line>
          ,
          <country country="LB">Lebanon</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Sciences Lebanese International, University Nabatieh</institution>
          ,
          <country country="LB">Lebanon</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>of Technology Lebanese, University Saida</institution>
          ,
          <country country="LB">Lebanon</country>
        </aff>
      </contrib-group>
      <fpage>105</fpage>
      <lpage>112</lpage>
      <abstract>
        <p>- Network traffic identification is the first and most important step in network management and security. Numerous methods introduced by researchers. One solution depends on processing both packet header and payload, which is costly from both time and processing aspects. Another solution depends on the flow of statistical information such as packet header length, flow duration. Blind classifiers are not accurate yet very fast and do not violate privacy. Machine learning fills the gap between accuracy and time by using the blind classifier method and comparing the results with grand truth then adapting and increasing the accuracy. K-NN is used widely for its effectiveness and simplicity. However, a major drawback of K-NN is its dependency on the training set, being a lazy classification algorithm with no classification model to build. In this work, we aim first at assessing the KNN algorithm in traffic classification. Then we introduce a new deficiency, related to the training samples distribution in the n-dimensional space we measure and propose an enhancement for K-NN adapting to the new problem and outperforming native K-NN classifier. We weight the classes, not the instance, based on the intersections of class clusters in the dataset. Finally, we propose a new Class Weight based K-NN Classifier (CWKNN), an enhanced K-NN algorithm that can easily adapt to the newly explored training set deficiency.</p>
      </abstract>
      <kwd-group>
        <kwd>K-NN</kwd>
        <kwd>weighted K-NN</kwd>
        <kwd>traffic classification</kwd>
        <kwd>computer network</kwd>
        <kwd>traffic identification</kwd>
        <kwd>training dataset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The ability to identify network applications is centric to many
network management and security tasks, including quality of
service assignment, traffic engineering, content-dependent
pricing, resource allocation, traffic shaping, and others. With
the proliferation of applications, many of them using different
kinds of obfuscation, traditional port-based classification has
long become obsolete.
inspection as in [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ], or DPI, evaluates the data part
and the header of a packet that is transmitted through an
inspection point. DPI goes beyond examining IP packet
headers, therefore, it raises many privacy concerns and is not
applicable when the traffic is encrypted or tunneled.
However, DPI techniques are considered in the literature as
the most accurate techniques and are used therefore as
reference classifiers to build the Ground Truth or reference
results.
      </p>
      <p>On the other hand, blind classifiers do not inspect the payload
and have the potential ability to deal with these obstacles, at
the expense of an acceptable sacrifice in accuracy. However,
less accurate, the so-called blind methods are preferred in
most
environments
because
they
guarantee
the
users’</p>
      <p>privacy, have the potential to classify encrypted
communications and usually require less computational
power.</p>
      <p>Most of these techniques are based on traffic attributes at the
network and transport layers, such as packet sizes and
interarrival times. Due to the problem dimensionality, Machine
Learning (ML) techniques can be used in the classification
context. ML classification is considered an instance of
supervised learning, i.e., learning where a training set of
correctly identified observations is available. An algorithm
that implements classification, especially in a concrete
implementation, is known as a classifier. In machine learning,
the observations are often known as instances, the
explanatory variables are termed features (grouped into a
feature vector), and the possible categories to be predicted are
classes. The term "classifier" sometimes also refers to the
mathematical function, implemented by a classification
algorithm that maps input data (traffic attributes in the traffic
classification case) to a category (Application type like
HTTP, NTP, DNS…etc.).</p>
      <p>
        There is a relevant research activity in network traffic
classification, employing different Machine learning
approaches. Among them, K-NN was used in several papers
[
        <xref ref-type="bibr" rid="ref3 ref5">4-7</xref>
        ] due to its simplicity and zero-training time, as will be
shown in the next sections.
      </p>
      <p>
        Cover and Hart first proposed K-NN (K Nearest Neighbor)
algorithm in 1968. The algorithm itself is simple and
effective, which is a typical lazy or instance-based learning
algorithm. It is not essential for the classifier to use the
training set to train and the training time complexity is 0.
With K-NN, the function is only approximated locally and all
computation is deferred until classification. The principle of
K nearest neighbor classification algorithm
(Kneighborhood) is that if the K samples in the eigenvalue space
most closed to the sample to be identified almost belong to
the same class, we determine this sample also belongs to such
a class. The key difference between classification and
regression tree in [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ] is that in classification, the task is
predicting a discrete class label while in regression, the task
consists of predicting a continuous quantity.
      </p>
      <p>In section two, we included the most significant papers
available in the literature. In section three, we explained the
known knn problems which is assessing knn in classifying the
datasets in each application. In section four we introduced the
new problem and we proposed and applied a solution. In
section 5 we will describe briefly the Experimental design
(for the lack of space). And provide the results of this work.
In section 6 there is a conclusion.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        A large number of traffic classification approaches exist in the
literature, some are based on deep packet inspection, which
discloses the packet payload and raises privacy concerns,
while others rely on discriminative traffic attributes at the
network and transport layers, such as packet sizes and
interarrival times. This emergent research needs to arise since
IANA Port-based classification in [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ] has become obsolete for
more than one decade. Machine learning techniques have been
widely used in the literature, especially the K Nearest
Neighbor due to its simplicity, zero-time for training, and
adequacy for general classification contexts.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ] authors reduced the flow attributes used in
classification by introducing a relationship between flows in
order to derive an expanding vector of flow which represents
the attributes of the classes. Their algorithm outperforms naïve
base classifier with 3.2% of decrease in error rate and achieved
99% of accuracy.
      </p>
      <p>
        Known DPI methods were enhanced in [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ] by introducing a
methodology that automatically learns signatures for
application. Their method enhances known DPI by identifying
new applications that old dpi cannot identify, has higher
accuracy, handles variation in applications and adapts to
changing network traffic without user intervention.
Many papers in the literature used K-NN for traffic
classification and attempted to enhance the native algorithm
by addressing key challenges. For example, to select the most
appropriate K value for K-NN, authors in [7] used Dynamic
K-NN, distance weighted K-NN, clustering the results, and
building a model. Dynamic K-NN is responsible for selecting
the most appropriate K value for each dataset. In the training
phase, K value is selected as well as calculating the weight of
each sample, then having clusters in order to build a model. In
the test phase, they calculate the distance between the center
of each cluster and the test point X to be classified. After
finding the nearest cluster, the weighted distance between X
and the rest of the cluster samples is calculated. After
determining the class of X they add it to the training dataset
after calculating its weight. As a result, they preserved the
same accuracy but were able to decrease the time to classify
an instance by 50%.
      </p>
      <p>III.</p>
      <p>K-NN CHALLENGES</p>
      <sec id="sec-2-1">
        <title>A. Algorithm Challenges</title>
      </sec>
      <sec id="sec-2-2">
        <title>1) Adjusting the K value</title>
        <p>
          The K value has a major impact in determining the nature of
the instance to be classified, selecting a small value of K
means that noise will have a higher influence on the result and
a large value make it computationally expensive as explained
in [
          <xref ref-type="bibr" rid="ref11">12</xref>
          ]. .The only way to validate a K-NN model is by the
error on test data. As The authors in [
          <xref ref-type="bibr" rid="ref12">13</xref>
          ] applied the manual
way of K value Selection. They also applied other solutions
related to the dataset modelling. They outperformed old K-NN
with performance increase ranging from 0.01 to 0.02 of
average misclassification cost.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2) Distance Method Selection</title>
        <p>
          The distance function plays a major role in K-NN
classification accuracy. Researches constantly enhance
distance functions and add a new attribute to increase
accuracy. To determine the proximity of two objects in the
feature space, any distance type method can be applied.
(Euclidean, Manhattan, Minkowski, etc.) yet researcher must
test different distance types for each dataset application . In
[
          <xref ref-type="bibr" rid="ref13">14</xref>
          ] author introduced a new parameter for the original
distance method enhanced K-NN by focusing on enhancing
the distance measurement method. In [
          <xref ref-type="bibr" rid="ref13">14</xref>
          ] the authors
introduced a new variable to the distance rule that made it
dynamic. This variable represents the radius of a sphere
centered on the new object and excluding all samples that have
a distinct label, the distance between an unknown point X (to
be classified) and samples that have the same
“nature” is less than 1 while the distance between X and
the different elements is 1. They tested the new algorithm on
five different Datasets and values of K ranging from 1 to 50,
with Manhattan, Euclidean distance methods used, they
achieved lowering error rates up to 6%.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>B. Dataset Challenges</title>
      </sec>
      <sec id="sec-2-5">
        <title>1) Outlier Instances</title>
        <p>An outlier is a data point that differs significantly
from other observations. An outlier may be due to
variability in the measurement or it may indicate the
experimental error; the latter are sometimes excluded
from the data set. Outliers, when found in the training
dataset, can decrease the classification accuracy of
KNN voting.</p>
        <p>
          Finding such exception has received much attention
in the data mining field. For example, one of the
solutions used in [
          <xref ref-type="bibr" rid="ref14">15</xref>
          ] is to partition the dataset into
several clusters and then in each cluster, to calculates
the Kth nearest neighborhood for the object to find
outliers.
        </p>
        <p>Another Solution is weighting each Instance which
is quantifying each Instance contribution in the
Classification process. And adding that number to the
voting process.</p>
        <p>
          Authors in [
          <xref ref-type="bibr" rid="ref15">16</xref>
          ] used fuzzy K-NN to enhance
accuracy f score and g mean of pulsar selection
algorithm. They a special weighting algorithm in the
training phase, where they compute the relationship
between the features and the class membership. Nearest
neighbors are given high membership value to the class
they are near from. However, the samples that are far
from a class will be given a small membership value to
this class. The proposed algorithm outperforms other
machine learning algorithms in accuracy metric values
starting from 0.01 to 0.04, G means metric from 0.03 to
0.06, and decreased the FPR value by 0.04 .
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>2) Class Imbalance Problem</title>
        <p>A drawback of the basic "majority voting"
classification occurs when the class distribution is
skewed. That is, examples of a more frequent class tend
to dominate the prediction of the new example because
they tend to be common among the k nearest neighbors
due to their large number.</p>
        <p>One way to overcome this problem is to weight the
classification, taking into account the distance from the
test point to each of its k nearest neighbors. The class
(or value, in regression problems) of each of
the k nearest points is multiplied by a weight
proportional to the inverse of the distance from that
point to the test point.</p>
        <p>
          Another Intersecting application of weighting is
Introduced by authors in [
          <xref ref-type="bibr" rid="ref3">4</xref>
          ] used the sigmoid function to
weight training Instances based on Instances distribution in
the dataset.
        </p>
        <p>WKS algorithm boost the weight of instances surrounded
by “friendly neighbors” and the decreases the weight if the
instance is surrounded mostly with different class members.
They were able to decrease the error rate that ranged from
0.007 to 0.025 for 10 different datasets. We find that their
work was the most significant to our algorithm so we
implemented their algorithm in the experimental setup we
built and compared the results between the native, WKS, and
CWK-NN.</p>
        <p>IV. PROPOSED ALGORITHM (CLASS WEIGHT K-NN)
As mentioned previously, the dataset may have
various deficiencies, most commonly, outliers (samples
with errors) and, imbalance in the dataset classes (some
classes are predominant over the others). To address
these deficiencies, weighing samples (outliers problem),
or clustered samples (imbalance problem) are among the
most common solutions found in the literature.</p>
        <p>To explore and understand further KNN training set
deficiencies, let us consider the ideal training set for
KNN, as derived from the way KNN classifies best any
unknown sample.</p>
        <p>In a very ideal training-set situation, space is made
of one representative sample for each class so that any
unknown point will fall in the n-dimensional (n is the
number of features or attributes)
“exactly” (i.e. coincides) in the location where
only one of the class samples exist.</p>
        <p>In real training-sets, however, the location of the
unknown point and existing sample can rarely coincide.
They might rather be close enough for KNN to be able
to derive the classification decision through majority
voting. When each class has many samples, KNN voting
described earlier will make the assumption about the
eventual class of the unknown point.</p>
        <p>As such, in a less or near-ideal training-set situation,
an unknown point may fall then within the zone in the
ndimensional space where only members of the same
class surround it, in this case, KNN voting will be more
reliable since the unknown point shall undoubtedly
belong to the same class of its surrounding neighbors.
Following the same reasoning, the closer the samples are
to the class centroid (i.e. higher density near the
centroid), the more reliable KNN decision shall be.</p>
        <p>To illustrate this concept, Figure 1 shows, the
distribution of labeled training points. The set of labels
belonging to each class (Classes 1, 2 and 3), are
clustered into circles for n=2 or spheres for n=3. The
point labeled with X is the unknown point. Class cluster
radius is defined by the standard deviation of the
training samples distances from the centroid, and the
distance between two clusters is the distance between
their centroids. For example, D12 denotes the distance
between centroid C1 of class 1 and C2 of class 2
Fig. 2. ideal distribution for a KNN-based classifier</p>
        <p>Fig 2. Shows a near-ideal distribution where the
clusters are disjoint and do not intersect at all,
mimicking the near ideal situation. In this case, a test
point falling into only one of the class clusters will lead
to a highly deterministic KNN voting decision since all
nearest neighbors to X will have the same label, class 2
in this case.</p>
        <p>In the case where unknown points fall out of the
radius of any of the existing clusters, KNN voting
decision still applies however with less reliability. In
other words, the points falling in the intersections or
between zones are those for which KNN decision will
be controversial (i.e. not all of the nearest neighbors
have the same class label), which will decrease the
overall reliability of KNN classification decisions when
compared to the ideal training-set distribution in Fig 2.</p>
        <p>To highlight more on this aspect and for simplicity,
in Fig3. we introduce a dataset having only four classes.
Each instance has only two attributes. Please notice that
the classes are balanced yet two classes is intersecting.</p>
        <p>Consider the reliability of the KNN classification
decision when classifying the instances that belong to
collide classes. The reliability of KNN decisions is very
weak. yet for the classes not intersecting the reliability
is supposed to increase evidently for training sets having
the distribution similar to class one and class 3. Since
the reliability of KNN decisions for points, falling inside
the clusters will be higher than those falling outside any
of the existing clusters or at the intersections. Except for
the differences in the neighbors’ distances
considerations, the reliability of the classification
decision shall be quite the same for points falling
outside any of the existing clusters or at the intersections
in both cases.</p>
        <p>Therefore, when the circles representing class
clusters are more distant and dense as in Fig2, KNN
classification is supposed to become more reliable, at
least for the unknown points falling within class
clusters. On the contrary, having too close cluster
centroids with lower density (i.e. more spread
distribution) will result in additional intersections
between clusters, which will decrease the reliability of
KNN classification for the unknown points falling on
the intersecting zones between classes (e.g. C2 and C3
in Fig 1.).</p>
        <p>In other words, when class clusters have no or fewer
intersections, the training dataset shall become closer to
the “near-ideal distribution” in Fig 2. . Thus
leading for better classification results.</p>
        <p>Despite the fact that imbalanced sets are likely to
have class clusters intersections, having one
characteristic (i.e. class imbalance or class intersections)
does not necessarily imply the other:</p>
        <p>A dataset can be imbalanced but classes do
not intersect
A dataset can be balanced but classes do
intersect</p>
        <p>The latter case, depicted in Fig 1. is the most
relevant to our approach where we measure the new
KNN training set deficiency related to
relative points’ distribution within the
ndimensional space rather than
their relative numbers. To the best of our knowledge,
“Class
explored
clusters
intersections”
is
not
yet</p>
        <p>as a deficiency for KNN training
sets, nor
was it regarded in separate from
the
imbalanced class problem.</p>
        <p>Up to this level, we have theoretically explained
and discussed what we think to be ideal for KNN
training set distribution and speculated about its impact
on KNN classification. However, most of the presented
assumptions
have
to
be
proven
through
experimentation.</p>
        <p>For this reasons, we attempt next to measure and
quantify
the
value
of
“class
clusters
intersections”
inside</p>
        <p>KNN training sets. Then,
we attempt to derive a new KNN classifier in order to
experimentally demonstrate the importance and impact
of considering of class intersections or training points
distribution in the training set on KNN classification.</p>
      </sec>
      <sec id="sec-2-7">
        <title>A. Weighting Classes with CWKNN</title>
        <p>A Class Weight based KNN Classifier or CWKNN is
simply a KNN classifier that takes into consideration the
samples distribution or class intersections described earlier.
As mentioned, intersections should be expressed in terms of
newly defined class weights. The WCKNN algorithm is a
two-phase process that includes the training and the testing
described as follows:
123</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Calculate centroids of each class.</title>
      <p>Calculate the variance instance of each class.</p>
      <p>Calculate the distance between centroids.</p>
      <p>if the distance is larger than the sum of two
radiuses the classes does not intersect and the
weight of the two classes remain the same. Else
if the distance between centroids is smaller than
the sum of radiuses there is an intersection and
the weights of the classes must be decreased.
4</p>
      <p>Assign the weight calculated for all instances.</p>
      <p>Training phase:
For each class we add two instances the first instance is
The centroid, and second instance represents the avarge
distribution of the class in the n dimensions (where n is the
number of attributes).</p>
      <p>The centroid is the instance core instance having the
average value of all attributes. The radius is the distance
between centroid and the instance having the variance of
value for all attributes.</p>
      <p>Thus
it’s
the
allowed
normal margin for class distribution in the n dimensional
space. All the instances outside the threshold is considered as
outliers and they are not significant comparing to the majority
of the class instances. Distances between centroids are
accordingly calculated. To formalize how CWKNN operates
in the training phase refers to simply formalize how class
weights are calculated. The main idea of setting the class
weights is that classes who do not intersect have less error
margin. We boost the weight of all of Instances having the
same class. However, classes
with intersections should
have decreased weights.</p>
      <p>Let M be the total number of classes in the dataset
For each class i, find the distance between the class
centroid and all other centroids.</p>
    </sec>
    <sec id="sec-4">
      <title>Equation 1</title>
      <p />
      <p>= 1 +  −
Equation 2
=</p>
      <p>∑ 
=1
, 1]
(1 + 
aij is a discrete value indicating if there is an intersection
Ai is a value measuring the level of intersections between
class i and the rest of classes.</p>
      <p>As shown in equation 1, if Ai, denoting the number of class
i’s
intersections
with
other classes,
increases
then Wi should decrease.</p>
      <p>1</p>
    </sec>
    <sec id="sec-5">
      <title>Ai has a minimal value of</title>
      <p>(since each class intersects with
itself at least)--when the class is disjoint from all other
classes-- and a maximum value for Ai=1 when theoretically
it intersects with all other classes.</p>
    </sec>
    <sec id="sec-6">
      <title>This mechanism ensures that the interval of Wi is.</title>
      <p>Intersecting classes are not thus explicitly punished, but
rather, they are not rewarded higher weights.</p>
      <p>Test phase: each training sample belonging to class i is
assigned the weight of the class. The weighted voting of K
neighbors is combined, and the weight of each neighbor is the
function of the class to which the neighbor
the training phase.</p>
      <p>As per equation 1, the class weight decreases exponentially
when the value of the number of intersections increases
towards M.</p>
      <p>The contributions of the classification for the
training sample belonging to a good class is regulated thus by
the weight value when the numbers of intersections get
considerably higher.
The
completion
of
traffic
these
accomplishment of some additional objectives related to
building the experimental setup.</p>
      <p>In this experiment we used a server that has 16 GB of ram ,
2 Intel(R) Xeon(R) 2.66 GHz processors with 4 cores each.
With Ubuntu 18.01 OS to capture network traffic and
converting the flows into a dataset. We captured network
flows by using tcp dump. then we used nDPI to detect the real
nature of the flow traffic and to set the flow attributes.
was collected the flow , over an extended period of time
spanning around 3 days and totaling around 1.2 GB of real
traffic. we have excluded the encrypted traffic.
The 3 days span of collecting the data was due to our desire
to collect, a large dataset for testing and classification and
then we were obliged to use a part of the data collected 51752
labeled flow for the lack of resources. The part was selected
randomly. Knn classification algorithm stores all training
dataset in the in the RAM.</p>
      <p>After obtaining the dataset we have used weka
open
source software to add wks and our algorithm.</p>
      <p>To validate the obtained Result we have used many other
datasets from the weka library. We will only include the
results from three dataset for the lack of time.</p>
      <p>CONCLUSION</p>
    </sec>
    <sec id="sec-7">
      <title>In this work, we have explored and measured a new deficiency, related to KNN training samples distribution in the</title>
      <p>n-dimensional
space and classes’
intersections.</p>
      <p>We weighted classes accordingly in the training set. Then,
we suggested class weighting scheme, then we derived a
newly
proposed</p>
      <p>Class</p>
      <p>Weight based</p>
      <p>KNN</p>
    </sec>
    <sec id="sec-8">
      <title>Classifier (CWKNN), an enhanced KNN algorithm that can account for sample distributions or classes intersections in KNN training sets.</title>
      <p>Furthermore, we have assessed KNN algorithm in traffic
classification contexts through an experimental test-bed that
we build for this purpose.</p>
      <p>Our tests on real captured and Internet-based datasets showed
up top 7%</p>
      <p>enhancements in the classification results,
compared to the native and samples’ weighted KNN
approaches.</p>
      <p>Moreover, and as per our testbed, we showed that K=3 and
Euclidean distance type can lead to more than 98% of overall
Tp in computer traffic classification context.</p>
      <p>VII.</p>
      <p>REFRENCES</p>
      <p>J. Khalife J. Verdejo, and A. Hajjar A multilevel taxonomy and
requirements for an optimal traffic-classification model [Journal].
[s.l.] : Int. Journal of Network Management , 2014. - 2 : Vol. 24.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. Khalife J.</given-names>
            <surname>Verdejo</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hajjar</surname>
          </string-name>
          <article-title>On the Performance of OpenDPI in Identifying P2P Truncated traffic</article-title>
          [Book]. - [s.l.]
          <source>: JOURNAL OF NETWORKS</source>
          ,
          <year>2013</year>
          . -
          <fpage>Vols</fpage>
          . VOL.
          <volume>8</volume>
          , NO. 1.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>digitalguardian</surname>
            <given-names>digitalguardian</given-names>
          </string-name>
          [Online] // digitalguardian. - digitalguardian,
          <year>2019</year>
          . - https: //digitalguardian.com/blog/what
          <article-title>-deeppacket-inspection-how-it-works-use-cases-dpi-and-more.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Min</given-names>
            <surname>Zhang</surname>
          </string-name>
          Min Qi ，*,
          <string-name>
            <surname>Kang</surname>
            <given-names>Sun</given-names>
          </string-name>
          , YuJun Niu ,
          <source>Longxiang Shi New Classification Algorithm WKS Based on Weight [Conference] // 2017 IEEE 19th International Conference on e-Health Networking, Applications and Services (Healthcom)</source>
          .
          <article-title>-</article-title>
          <string-name>
            <surname>Dalian</surname>
          </string-name>
          , China : IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Hou</given-names>
            <surname>Kaihu</surname>
          </string-name>
          <string-name>
            <surname>Zhan</surname>
          </string-name>
          , Hui ,
          <source>Zhou Zhou A Weighted KNN Algorithm Based on Entropy [Journal]. - Kunming</source>
          , china : IEEE,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Shweta</given-names>
            <surname>Taneja Charu Gupta</surname>
          </string-name>
          , Kratika Goyal ,
          <string-name>
            <surname>Dharna Gureja Enhanced K-Nearest Neighbor Algorithm Using Information Gain</surname>
          </string-name>
          and Clustering [Conference] // 2014 Fourth International Conference on Advanced Computing &amp; Communication
          <string-name>
            <surname>Technologies</surname>
          </string-name>
          . - [s.l.]
          <source>: IEEE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Kumar</given-names>
            <surname>Nath Syed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <article-title>Akram Salah An Enhancement of k - Nearest Neighbor Classification Using Genetic Algorithm</article-title>
          [Journal].
          <article-title>- 2018.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <surname>machinelearningmastery</surname>
            <given-names>machinelearningmastery</given-names>
          </string-name>
          [Online] // machinelearningmastery. - machinelearningmastery,
          <year>2018</year>
          . - https: //machinelearningmastery.com
          <article-title>/classification-versus-regression-inmachine-learning/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>IANA</given-names>
            <surname>IANA Port</surname>
          </string-name>
          <article-title>Numbers</article-title>
          . [Online] // IANA . - IANA,
          <year>2013</year>
          . - 01
          <year>2013</year>
          . - http: //www.iana.org/assignments/ port-numbers.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Lei</given-names>
            <surname>Dinga Jun Liua</surname>
          </string-name>
          , Tao Qina,
          <source>Haifei Lib Internet Traffic Classification Based on Expanding Vector of Flow [Journal]</source>
          . - China : IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Alok</surname>
            <given-names>Tongaonkar⇑ Ruben</given-names>
          </string-name>
          <string-name>
            <surname>Torres</surname>
          </string-name>
          , Marios Iliofotou, Ram Keralapura,
          <article-title>Antonio Nucci Towards self adaptive network traffic classification</article-title>
          [Journal]. - [s.l.]
          <source>: IEEE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <article-title>pseudo_teetotaler stack overflow</article-title>
          [Online] // stack overflow.
          <source>- 2019</source>
          . - https: //stackoverflow.com/questions/33884325/ideal
          <article-title>-k-value-inKNN-for-classification.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Zhang</given-names>
            <surname>Shichao</surname>
          </string-name>
          Cost-Sensitive KNN Classification [Journal]. - [s.l.]
          <source>: IEEE</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Jigang</given-names>
            <surname>Wang Predrag Neskovic</surname>
          </string-name>
          ,Leon N.
          <article-title>Cooper Improving nearest neighbor rule with a simple adaptive Distance measure</article-title>
          [Conference] // International Conference on Natural Computation. - [s.l.]
          <source>: ICNC 2006: Advances in Natural Computation</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Yang B. Huang KNN Based Outlier</surname>
          </string-name>
          Detection Algorithm in Large Dataset [Conference] // 2008 International Workshop on Education Technology and Training &amp; 2008 International Workshop on Geoscience and
          <string-name>
            <given-names>Remote</given-names>
            <surname>Sensing</surname>
          </string-name>
          . - Shanghai : Research Gate,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>TM</given-names>
            <surname>Mohamed Pulsar Selection Using Fuzzy knn Classifier</surname>
          </string-name>
          [Conference] // Future Computing and
          <string-name>
            <surname>Informatics Journal</surname>
          </string-name>
          (
          <year>2018</year>
          ). - [s.l.]
          <source>: IEEE</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>