<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multi-Layer Model and Training Method for Information-Extreme Malware Traffic Detector</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Sumy State University</institution>
          ,
          <addr-line>Rimsky-Korsakov st., 2, Sumy, 40007</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Model-based on multilayer convolutional sparse coding feature extractor and information-extreme decision rules for malware traffic detection is presented in the paper. Growing sparse coding neural gas algorithms for unsupervised pre-training of the feature extractor are used. Random forest regression model as a student in knowledge distillation from sparse coding layers is proposed for speed up inference mode. Information-extreme learning method based on binary encoding with tree ensembles and class separation with radial basis function in binary Hamming space are proposed. Information-extreme classifier is characterized by low computational complexity and high generalization ability for small labeled training sets. Simulation results with an optimized model on test open datasets confirm the suitability of proposed algorithms for practical application.</p>
      </abstract>
      <kwd-group>
        <kwd>malware detection system</kwd>
        <kwd>convolutional sparse coding network</kwd>
        <kwd>growing neural gas</kwd>
        <kwd>tree ensembles</kwd>
        <kwd>random forest regression</kwd>
        <kwd>information criterion</kwd>
        <kwd>information-extreme machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Existing malware traffic detection systems still do not provide high-reliability
solutions, as there are a constant increase the number and variety of new sources of
malware traffic and a small number of relevant labeled data [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Thus, the use of
handcrafted features for the description of observations leads to a decline the
informativeness of the features description and the effectiveness of learning of the decision rules
of the malware traffic detection system [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. Therefore, the most promising
approach to the synthesis of a features extractor is the use of ideas and methods of
machine learning for the hierarchical (deep) representation of observations for unlabeled
data [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        Conventional approaches to deep supervised machine learning require a significant
amount of labeled training examples and computational resources [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. In addition,
models trained with a supervisor based on gradient descent and its modifications are
vulnerable to adversarial attacks, noise and data novelty. To increase the
informativeness of the feature representation of observations, it is promising to use ideas and
methods of sparse coding and unsupervised competitive learning [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. This allows to
use the large volume of unlabeled data with maximum efficiency. Among the ways to
increase the generalization ability of the decision rules are considered ensemble
algorithms, error correction codes and methods of class separation within the geometric
approach. Also high speed of packet flow in modern networks require high
productivity of traffic analysis algorithms. To reduce computational complexity of data analysis
models, different methods of model pruning and knowledge distillation are used.
However, models hybridization and integrated use of different methods bring some
uncertainties to the final result, so the solution in this approach requires research and
verification. In this case, information criteria are considered the best metrics for
validation and verification of the result, because they directly characterize the reduction
of uncertainty in decision-making and are less sensitive to outlier and imbalances in
the data.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Formal Problem Statement</title>
      <p>
        Let the CTU-Mixed and CTU-13 datasets are given data collections from the real
network environment by CTU researchers from 2011 to 2015, which are formed as
pcap-files [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. The first CTU-Mixed dataset can be used for training a feature
extractor. The second CTU-13 dataset contains labeled flows and it could be used to
train the decision rules for detecting malware network traffic.
      </p>
      <p>It is necessary to build an informative feature extractor and reliable decision rules
using labeled and unlabeled datasets through optimization of model parameters. In the
process of training, it is necessary to maximize the information efficiency criterion of
the malware traffic detector</p>
      <p>*
E 
1 M</p>
      <p> max Em(k) ,
M m1 {k}
where E(k) is information efficiency criterion of recognition the class X mo on k -th
m
step of training; {k} – ordered set of training steps.</p>
      <p>When the malware traffic detector functions in its inference mode, it is necessary
to provide computational efficiency for high speed traffic.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Literature Review</title>
      <p>
        Convolutional multi-layer neural networks allow forming an informative hierarchical
features representation of input observations [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In addition, they have already shown
high efficiency in solving problems of machine vision and analysis of time series [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Meanwhile, supervised training requires a large amount of labeled data, the
labeling of which may be expensive or inaccessible in a reasonable amount of time. The
unsupervised training of convolutional networks is aimed at efficient use of unlabeled
examples, which are usually available quite a lot. It is carried out based on an
autoencoder or Restricted Boltzmann machine, which requires a large amount of training
data and long learning time to obtain an acceptable result [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In work [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] it is
proposed to use alternative approach based on k-means cluster-analysis algorithm to
speed up feature set training. However, k-means is characterized by slow convergence
and sub-optimality of the results due to the hard-competitive nature of its learning
scheme and the sensitivity to initial cluster initialization.
      </p>
      <p>
        In work [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is proposed a combination of the principles of neural gas and sparse
coding for the feature set training on unlabeled data. Given approach is characterized
by soft-competitive learning scheme that facilitates robust convergence to close to
optimal features distributions over the training sample. At the same time, embedding
of sparse coding methods can increase the immunity against interference and
generalization ability of features representation. Also, it is a well-known fact that sparse
representations of the input data are a crucial tool for combating adversarial attacks
and the production of de-correlated features as a result of the explaining-away effect.
However, the size of feature set is unknown beforehand and it is selected by the
developer, which leads to increase the optimization time.
      </p>
      <p>
        The required size of feature set in each layer of hierarchical representation is
difficult to predict in advance, so the promising approach to feature set learning is to use
the principles of growing neural gas, which automatically determines the required
number of neurons (features) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The presence of a mechanism for the adding of
new neurons, as well as the removal of excessive old ones, makes the algorithm more
flexible compared to the classical neuron gas, but it also has serious disadvantages.
The small values of the period between the iterations of the generation of new
neurons  lead to the instability of the learning process and the distortion of the formed
structures, as here observed the excessively frequent adding of new neurons. The high
value of the period  provides the expected effect, but at the same time it leads to a
significant slowdown in the algorithm. However, in the works [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ] it was shown
that achieving stability of learning could be done by setting the “radius of reach” of
the neurons, which involves the replacement of the parameter  on the threshold of
maximum distance of the neuron from each points of the training set attributed to it.
However, the mechanisms for updating neurons and assessing the remoteness of the
points of the input space to the neurons have not yet been reviewed in order to adapt
the learning process to the sparse coding of observations.
      </p>
      <p>
        The main disadvantage of sparse coding in representation learning is the use of an
iterative procedure during the inference which slows down the recognition process.
One of the popular ways to accelerate models is to use the principles of knowledge
distillation, where the redundant model acting as a teacher can be replaced by a
lightweight model acting as a student [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The ensemble of decision trees is a flexible and
computationally efficient model, which can potentially be used as a student model to
approximate the sparse coder [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. However, no such research had been conducted
and the effectiveness of such an approach is unknown, which underscores the
relevance of this issue.
      </p>
      <p>
        In addition the decision rules are important components in the malware detection
systems. As a rule, it represents a trainable classifier. At the same time, the
effectiveness of training a classifier is often considered as a measure of the effectiveness of the
feature extractor [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The most popular algorithm for classification analysis is the
method of support vector machine, where the training of decision rules takes place
within the framework of a geometric approach by constructing linear separable
hypersurface in the secondary features space [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. However, this algorithm requires a lot of
hyper-parameters adjustments and its performance depends on the complexity of the
kernel functions. In work [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], were proposed the construction of decision rules by
adaptive binary encoding of the input features and the optimizing in information sense
the radial-basis based separable hyper-surface in the Hamming binary space. Such a
classifier has high operational efficiency, since it uses low computing complexity
operations as comparison and logical XOR.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Model and Training Method for Malware Traffic Detector</title>
      <p>
        The internal characteristics of the unit of traffic (packet stream or session) are best
displayed in the front part of its bytes, which contains connection data and some
content data. The process of converting a pcap-file into a training data set involves three
main steps: the separation of traffic into discrete units, taking into account some
granularity, clearing traffic by removing empty and duplicate units, forming training
images. When dividing traffic into discrete units, one can consider the following
granularities: TCP connection, flows, session, service, and host. In this paper, it is proposed
to divide the incoming traffic into flows, where a number of packets have the same
tuple of five elements: the source and destination IP address, the source and
destination ports, the protocol number. In this case, the length of the stream is limited to 784
bytes, so longer streams are cropped, and shorter ones are supplemented by zero
bytes. As a result, we have an image of 28x28 pixels, which will be delivered to the
input of the feature extractor. The brightness of each pixel is normalized to the range
[
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ].
      </p>
      <p>
        As a basis for building architecture of features extractor was used a convolutional
network is known as LeNet-5 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the main modification of which relates to use the
unfixed number of convolutional filters, the amount of which is determined during the
layer-wise training. The pixel activation of each channel of features map is offered to
calculate based on greedy-L0 Orthogonal Matching Pursuit algorithm (OMP) or
L1regularized least angle regression algorithm (LARS) with the function of ReLU
activation [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In order to accelerate the model in the inference mode, it is possible to
replace the computationally intensive search for sparse coefficients with a
noniterative approximating encoder (Figure 1). According to distillation knowledge
principle, the training set for approximation encoder is formed from input of the layer and
pseudo-labels from output of the layer. In this case, pseudo-labels are obtained by
OMP or LARS algorithms.
      </p>
      <p>
        It is proposed to implement sparse coding with OMP and LARS algorithms where
stop criterion based on achievement of 30% non-zero entries in sparse code. A Local
Contrast Normalization layer, placed after the sub-sampling layer, before the next
layer, amplifies the informative features and weakens the rest of the pixels of the
feature map.
The dataset for training of feature extraction layer is formed by decomposition of
images or activation maps to patches. These patches are reshaped to 1D vectors,
which put on the input of growing sparse coding neural gas algorithm, main steps of
which are given below [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
1. Initialization of the counter of training vectors t : 0 .
2. Two initial nodes (neurons) wa and wb are assigned by random selection from the
training set. Nodes wa and wb are connected by an edge whose age is zero. These
nodes are considered non-fixed.
3. Selected from the dataset the following vector х , which is normalized to a unit
length (L2-normalization).
4. Normalizing each base vector wk , k  1, M to a unit length (L2-normalization).
5. Calculation of the similarity of the input vector х to the base vectors wsk W for
their sorting
(wsT х)2  ...  (wsT х)2  ...  (wsTM1 х)2
0 k
.
6. The closest node is selected ws0 and the second closest to the node ws1 .
7. Increase the age of all incident edges ws0 by one.
8. If ws0 is fixed, then should move to step 9, otherwise, to step 10.
9. If (wsT х)2  v , then proceed to step 12. Otherwise, should be added a new
non0
fixed neuron wr to a point that coincides with the input vector wr  х , also is
adding a new edge, that connects wr and ws0 , then proceed to step 13.
10. The node ws0 and its topological neighbors (the nodes connected to it by the edge)
are displaced in the direction to the input vector х by the next formulas [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
ws0   bt y0 (x  y0 ws0 ) , y0 : wsT0 x ,.
wsn   n t yn (x  yn wsn ) , yn : wsTn x ,.
      </p>
      <p>0   b  1, 0   n   b ,</p>
      <p>
         t :  0 ( final / 0 )t /tmax ,
where ws0 , wsn – vectors of correction of weight of the neuron-winner and its
topological neighbors, respectively;  b ,  n – the constants of updated forces of
weighting coefficients of the neuron-winner and its topological neighbors
respectively;  0 ,t , final – initial, current and final learning rate respectively.
11. If (wsT0 х)2  v , note the neuron ws0 as fixed.
12. If ws0 and ws1 are connected by edge, then its age is reset, otherwise a new edge
with a zero age is formed between ws0 and ws1 .
13. All edges in the graph with the age more than amax are removed. In the case that
some nodes do not have incident edges (become isolated), they are also removed.
14. If t  tmax then proceed to step 15, otherwise – increment of the counter of steps is
t : t 1 and then proceed to step 3.
15. If all neurons are fixed, the execution of the algorithm stops, otherwise proceed to
step 3 and a new epoch of learning begins (repetition of the training set).
Features extractor can be fine-tuned based on the backpropagation algorithm with a
temporary or permanent neural classifier at the model output [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Since in the
conditions of nonstationarity the informativeness of features in advance cannot be known,
the fine tuning is not provided in our algorithm. The purpose of the feature extractor
is to disentangle explanatory factors.
      </p>
      <p>
        The information-extreme classifier requires binary representation of the input
signal to build error-correction decision rules. The ensemble of decision trees is a
computationally effective method for inducing informative binary features of observations
(Figure 2). Nodes of decision trees are numbered. Numbers of nonzero bits of
resulting binary code correspond to the numbers of nodes through which the decision path
lies [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>Information-extreme classifier under inference mode make decision on belonging
of input datapoint x with appropriate binary representation b to one class from set
{X zo | z  1, Z} according to maximum value of membership function  z (b) through
the expression arg max{z (b)}.</p>
      <p>z</p>
      <p>In this case membership function  z (b) , the optimal container of which has
support vector bz* with dimension N 2 and radius dz*, is derived from formula
 N2
 z (b)  exp   bi  bz*,i / dz* .</p>
      <p> i1 
(1)
Let D  {xj , y j | j  1, n} be a training set, where n is size of dataset and y j is label
of j-th datapoint, which correspond to one class from set of classes {X zo | z  1, Z} . In
this case, classifier evaluates belonging of j-th datapoint x j with N1 features to one
of the Z classes performs feature encoding using decision trees and decision rules
constructed in radial basis of binary Hamming space. The training of
informationextreme classifier is performed according to the following steps.
1. Initialize weight wj  1 / n .
4. Train decision tree Tk</p>
      <p>split.
5. Binary encoding of x j datapoint from
2. For k = 1,…, K do
3. Bootstrap Dk from D using probability distribution P( X  x j )  wj .
on Dk using entropy criterion to measure the quality of</p>
      <p>D
using concatenation of results from
T1 ,...,Tk
trees.</p>
      <p>The
output
of
this
step
is
a
binary
matrix
{bz,s,i | i  1, N2 ; s  1, nz ; z  1, Z} , where N2 is a number of induced binary
features and nz is a number of samples corresponded to class X zo . Hence the equality
n   nz condition is met.</p>
      <p>z
6. Build information-extreme decision rules in radial basis of binary Hamming space
and compute optimal information criterion:</p>
      <p>Ez  max Ez (d ) ,</p>
      <p>d
where {d}  {0,1,...,   bz,i  bc,i 1} is a set of concentric radiuses with center bz

 i 
(support vector) of data distribution in class X zo , which computed using rule
0, otherwise.

1, if
bz,i  
1 nz</p>
      <p> bz,s,i 
nz s1
1 Z 1 nс</p>
      <p>  bс,s,i ; ,</p>
      <p>Z с1 nс s1
Ez </p>
      <p>1 ( z   z )
log2 (2  )  log2 </p>
      <p>
         2  ( z   z )  
 log2   ,
 ( z   z )  
where Ez – training efficiency criterion of decision rule for X zo class, which is
computed as the normalized modification of the S. Kullback’s information measure [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]:
where  z ,  z are the false-positive and false-negative rates of classification of input
vectors as belonging to the X zo class; ς is any small non-negative number, introduced
to avoid uncertainty when dividing by zero.
7. Test obtained information-extreme rules on dataset D and compute error rate for
each sample from D . Under the inference mode, decision on belonging of
datapoint b to one class from set {X zo | z  1, Z} is made according to maximum value
of membership function  z (b) through the expression arg max{z (b)}. In this
z
case membership function  z (b) of binary representation b of input datapoint x to
X zo class, the optimal container of which has support vector bz* and radius dz*, is
derived from formula (2).
8. Update {wj } proportional to errors on datapoint x j :
w j  1  m (x j ), m  y j ;
w j 
      </p>
      <p>w j
 j w j
.
9. If | Ek*  Ek*1 |  and k &lt; K/2 abort loop, where  = 0.001.</p>
      <p>Thus, the resulting model consists of several layers of tree ensemble with optimal in
informational sense decision rules at the output.
(2)
(3)
(4)</p>
    </sec>
    <sec id="sec-5">
      <title>Result and discussion</title>
      <p>The training sample formed with CTU-Mixed for the training of the feature extractor
contains 10,000 instances. To train the information-extreme classifier are formed by
1000 instances per class in the training and test datasets. In growing sparse coding
neural gas algorithm were chosen the following parameters b  0.5 , b  0.05 ,
amax  100 , 0  1 та  final  0.01 . The parameter of the threshold of neuron fixation
v and the parameter of the maximum number of trees K of the classifier are adjusted
by scrolling through the values. Table 1 shows the dependence of the number of
neurons in the first M1 and second M 2 layers of feature extractor, the criterion of the
effectiveness of training averaged over the classes E and accuracy by the validation
sampling of the parameter v . In the tree ensembles, max depth is set to 5 and max
features is set to N1 .</p>
      <p>The analysis of Table 1 shows that increasing the threshold  leads to an increase
in the number of neurons in the process of unsupervised training the features
extractor. At the same time, increasing the threshold from 0.8 to 0.9 practically does not
affect the accuracy of the decision rules. It means, that the value *  0.8 is optimal
and allows to form a more compact features representation (compression), meanwhile
  0.9 allows to form a sparse representation based on overcomplete basis.</p>
      <p>Knowledge distillation is implemented with Random Forest regression as student
model, where the number of decision trees is limited to 150. The obtained model has
equivalent accuracy. In this case, the inference time is reduced by 65 times.</p>
      <p>Figure 3 shows a graph of maxima’s changes of the information criterion (4)
averaged in the set of classes in dependence on the number of decision trees in
information-extreme classifier with v* = 0.8. In this case, the maximum number of trees is
limited, K= 100.</p>
      <p>A graph of the change of the average information criterion (4) in dependence from
the number of decision trees in information-extreme classifier</p>
      <p>The analysis of Figure 3 shows that the optimal value of the hyper parameter K*
is equal to 185. Further increase of the parameter K does not lead to an increase in
the accuracy of the decision rules. At optimal parameters of the extractor and the
classifier, the accuracy of detection of malware traffic is 96.1%. It indicates on the
informative nature of the features descriptive of observation. Figure 4 shows the
dependence of the information criterion (4) on the code radius of the container of each class.</p>
      <p>The analysis of Figure 4 shows that the maximum values of information criterion
of learning for the first and second classes are equal to E1* =0.590 and E2* = 0.597,
respectively, and the optimal values of radii of the corresponding containers of the
classes of recognition d1* = 26, d 2* = 32 (in code units). In this case, the inter-center
Hamming distance is 65 indicating compactness of the feature vector distributions and
the clarity of partition in the binary Hamming space.</p>
      <p>Thus, the proposed training algorithm allows determining automatically the
optimal number of neurons at each layer. At the same time, approximation of the sparse
encoder by the non-iterative model, Random Forest regression, allowed accelerating
the inference mode.</p>
      <p>
        The results of simulation on data from CTU-Mixed and CTU-13 datasets show that
obtained result is superior to result from [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and it is acceptable for practical
applications.
      </p>
      <p>Conclusions
10. The scientific novelty of the obtained results is as follows:
─ the algorithm of growing sparse coding neural gas is proposed for the first time,
which allows unsupervised learning the optimal set of neurons for each layer of the
convolution sparse coding model of feature extractor model;
─ for the first time it was proposed to apply the principle of knowledge distillation to
reduce computational costs in the algorithms of sparse coding through the
application of approximation by the random forest model, which in the inference mode is
non-iterative and computationally efficient;
─ for the first time, an information-extreme algorithm of supervised learning is
proposed for constructing the decision rules of the detector of malware network
traffic.
11. The practical value of obtained results obtained for malware traffic detection
systems is a developing a new learning method that effectively uses both labeled and
unlabeled training sets. The results of simulation with using the CTU-Mixed and
CTU-13 datasets confirm the effectiveness of the obtained decision rules in
identifying the malware in test samples of traffic. In this case, the accuracy of the
decision rules of the malware traffic detector is 96.1%.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment</title>
      <p>The work was performed in the laboratory of intellectual systems of the computer
science department at Sumy State University with the financial support of the
Ministry of Education and Science of Ukraine in the framework of state budget
scientific and research work of DR No. 0117U003934.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Skrzewski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Flow Based Algorithm for Malware Traffic Detection</article-title>
          .
          <source>Computer Networks</source>
          .
          <fpage>271</fpage>
          -
          <lpage>280</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Berkay</given-names>
            <surname>Celik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Walls</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>McDaniel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Swami</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Malware traffic detection using tamper resistant features</article-title>
          .
          <source>MILCOM 2015 - 2015 IEEE Military Communications Conference</source>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Iglesias</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zseby</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Analysis of network traffic features for anomaly detection</article-title>
          .
          <source>Machine Learning</source>
          .
          <volume>101</volume>
          ,
          <fpage>59</fpage>
          -
          <lpage>84</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Yousefi-Azar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varadharajan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamey</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tupakula</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Autoencoder-based feature learning for cyber security applications</article-title>
          .
          <source>2017 International Joint Conference on Neural Networks (IJCNN)</source>
          .
          <article-title>(</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wei</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Ming Zhu, Xuewen Zeng, Xiaozhou Ye, Yiqiang Sheng:
          <article-title>Malware traffic classification using convolutional neural network for representation learning</article-title>
          .
          <source>2017 International Conference on Information Networking (ICOIN)</source>
          .
          <article-title>(</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Yangqing Jia, Sermanet,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Anguelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Erhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Rabinovich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <article-title>(</article-title>
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks for time series classification</article-title>
          .
          <source>Journal of Systems Engineering and Electronics</source>
          .
          <volume>28</volume>
          ,
          <fpage>162</fpage>
          -
          <lpage>169</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Qiying</given-names>
            <surname>Feng</surname>
          </string-name>
          , Chen,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Long</surname>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          :
          <article-title>Compressed auto-encoder building block for deep learning network</article-title>
          .
          <source>2016 3rd International Conference on Informative and Cybernetics for Computational Social Systems (ICCSS)</source>
          .
          <article-title>(</article-title>
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Weed identification based on Kmeans feature learning combined with convolutional neural network</article-title>
          .
          <source>Computers and Electronics in Agriculture</source>
          .
          <volume>135</volume>
          ,
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Labusch</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barth</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinetz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Sparse Coding Neural Gas: Learning of overcomplete data representations</article-title>
          .
          <source>Neurocomputing</source>
          .
          <volume>72</volume>
          ,
          <fpage>1547</fpage>
          -
          <lpage>1555</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mrazova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kukacka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Image Classification with Growing Neural Networks</article-title>
          .
          <source>International Journal of Computer Theory and Engineering</source>
          .
          <volume>422</volume>
          -
          <fpage>427</fpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Palomo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez-Rubio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The Growing Hierarchical Neural Gas Self-Organizing Neural Network</article-title>
          .
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          .
          <volume>1</volume>
          -
          <fpage>10</fpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Layer-Level Knowledge Distillation for Deep Neural Network Learning</article-title>
          .
          <source>Applied Sciences. 9</source>
          ,
          <issue>1966</issue>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hooker</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Approximation Trees: Statistical Stability in Model Distillation</article-title>
          , https://arxiv.org/abs/
          <year>1808</year>
          .07573. (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kil</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Deep learning of support vector machines with class probability output networks</article-title>
          .
          <source>Neural Networks</source>
          .
          <volume>64</volume>
          ,
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Moskalenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moskalenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korobov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semashko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The Model and Training Algorithm of Compact Drone Autonomous Visual Navigation System</article-title>
          .
          <source>Data. 4</source>
          ,
          <issue>4</issue>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gwon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cha</surname>
          </string-name>
          , H. T. Kung:
          <article-title>Deep Sparse-coded Network (DSN</article-title>
          ).
          <source>2016 International Conference on Pattern Recognition</source>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>