<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mapping and Compressing a Convolutional Neural Network through a Multilayer Network</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessia Amelio</string-name>
          <email>a.amelio@unich.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianluca Bonifazi</string-name>
          <email>g.bonifazi@univpm.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Corradini</string-name>
          <email>e.corradini@pm.univpm.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Marchetti</string-name>
          <email>m.marchetti@pm.univpm.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domenico Ursino</string-name>
          <email>d.ursino@univpm.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Virgili</string-name>
          <email>luca.virgili@univpm.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DII, Polytechnic University of Marche</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INGEO, University “G. D'Annunzio” of Chieti-Pescara</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper falls in the context of the interpretability of the internal structure of deep learning architectures. In particular, we propose an approach to map a Convolutional Neural Network (CNN) into a multilayer network. Next, to show how such a mapping helps to better understand the CNN, we propose a technique for compressing it. This technique detects if there are convolutional layers that can be removed without reducing the performance too much and, if so, removes them. In this way, we obtain lighter and faster CNN models that can be easily employed in any scenario.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Deep Learning</kwd>
        <kwd>Convolutional Neural Networks</kwd>
        <kwd>Multilayer Networks</kwd>
        <kwd>Convolutional Layer Pruning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In recent years, we have witnessed a massive spread of deep learning models in many contexts
[1, 2] with the goal of solving increasingly complex problems. The growing complexity of the
problems to solve requires increasingly sophisticated models to achieve the best performance.
However, more and more researchers realize the need to reduce the size and complexity of deep
networks [3, 4]. As a result, enormous eforts are being made to introduce new architectures of
deep learning networks that are more eficient and less complex. In parallel, several methods are
being proposed to reduce the size of existing networks without afecting their performance too
much [5, 6]. To this end, it is extremely important to be able to explore the various layers and
components of a deep learning model. In fact, we could identify the most important components,
the most interesting patterns and features, the information flow, and so on.</p>
      <p>In this paper, we want to make a contribution in this setting. In particular, we start from the
assumption that complex networks, and in particular multilayer ones, can significantly support
the representation, analysis, exploration and manipulation of deep learning networks. Based on
this insight, we first propose an approach to map deep learning networks into multilayer ones
and then use the latter to explore and manipulate the former.</p>
      <p>We focus on one family of deep learning networks, namely Convolutional Neural Networks
(hereafter, CNNs) [7]. Multilayer networks [8] are a type of complex networks sophisticated
enough to represent all aspects of a CNN. In fact, through their fundamental components (i.e.,
nodes, arcs, weights and layers), they are able to represent all the typical concepts of a CNN
(i.e., nodes, connections, filters, weights, etc.). Once we have mapped a CNN into a multilayer
network, we can use the latter to study and manipulate the former. To give an idea of its
potential, in the paper we will use it to support an approach to prune convolutional layers [9]
from a CNN. This approach aims to identify if there are layers in the CNN that can be pruned
without reducing the performance too much, and, if so, proceeds with pruning and returns a
new CNN without those layers.</p>
      <p>The outline of this paper is as follows: In Section 2, we describe our approach to mapping a
CNN into a multilayer network. In Section 3, we describe the use of the multilayer network
to prune one or more convolutional layers of the CNN. Finally, in Section 4, we draw our
conclusion and take a look at some possible future developments of our research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Mapping a Convolutional Neural Network into a multilayer network</title>
      <p>In this section, we present our approach to mapping a CNN into a multilayer network. It
represents the first contribution of this paper.</p>
      <sec id="sec-2-1">
        <title>2.1. Class network definition</title>
        <p>In this subsection, we describe a CNN by means of a single-layer network, referred to as a class
network. It is a weighted directed graph  = (, ,  ), where  is the set of nodes,  is the
set of arcs, and  is the set of arc weights.</p>
        <p>A CNN consists of  convolutional layers, each having  filters (also called “kernels”). In a
convolutional layer, each filter slides over the input with a given stride and creates a feature
map. The input ℐ of the first convolutional layer is given by the original image, while the input
of the next convolutional layer is given by a feature map. Applying a filter on the element
ℐ(, ) of the input provides a new element (, ) of the output feature map  of the output.
Based on this, the set  of the nodes of  consists of a set {1, 2, ...,  } of node subsets.
The subset  indicates the contribution of the ℎ convolutional layer . This is obtained by
applying the  filters of this layer to the input of . Therefore, a node  ∈  represents the
output obtained by applying the  filters of  to some position (, ) of the input.</p>
        <p>Since by applying a filter on ℐ(, ) we get a new element (, ), it is straightforward that
there is a direct connection between ℐ(, ) and (, ). Actually, there is a direct connection
not only between ℐ(, ) and (, ), but also between ℐ(, ) and each element adjacent to
(, ) within the filter area. As each convolutional layer  has a set of  filters, there are 
sets of direct connections between ℐ(, ) and (, ), one for each filter. Since the  filters
are applied to the input with a given stride, a set of similar connections towards the feature
maps is generated for diferent positions of the input.</p>
        <p>In a CNN, besides the convolutional layer there are pooling layers. A pooling layer shrinks
the input feature map and leads to an increase in the number of connections between the input
and the output. This is achieved by sliding the filter over the input with a given stride and
determining an aggregate value (e.g., the maximum) for each filter window. As aggregate values
are still elements of the feature map provided as input, the next application of a convolutional
layer to the feature map returned by a pooling layer generates connections between the aggregate
values and the elements of the feature map returned by the convolutional layer. In particular,
there will be direct connections between the aggregate values of the feature map provided as
input to the pooling layer and the adjacent elements of the feature map generated by the next
convolutional layer.</p>
        <p>Therefore, based on the previous reasoning, we can say that the application of a filter to
the element ℐ(, ) generates a new element (, ), whose value is obtained by the following
convolution operation:
(, ) =  (, ) * ℐ (, ) = ∑︀ =−   (, )ℐ( + ,  + )</p>
        <p>=−  ∑︀
Here,  is a filter of size (2 + 1) × (2 + 1).</p>
        <p>The direct connections generated between ℐ(, ) and (+, +), −  ≤  ≤ , −  ≤  ≤ ,
are labeled with the same weight (, ), representing the convolution result. In presence of
 filters, the direct connections between ℐ(, ) and ℎ( + ,  + ), − 1 ≤  ≤ 1, − 1 ≤
 ≤ 1, 1 ≤ ℎ ≤ , are weighted with the values of the corresponding convolution results
1(, ), 2(, ), . . . , (, ).</p>
        <p>
          Figure 1 illustrates the application of three filters of size 3 × 3 (green, blue and yellow
colored, respectively) to ℐ(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ). For each filter ℎ, it also shows the weighted direct connections
generated between ℐ(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) and ℎ(8+, 8+), − 1 ≤  ≤ 1, − 1 ≤  ≤ 1, 1 ≤ ℎ ≤ 3. The three
weights are obtained as: 1(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = 1(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) · ℐ (
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = 5; 2(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = 2(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) · ℐ (
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = − 2;
3(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = 3(
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) · ℐ (
          <xref ref-type="bibr" rid="ref8 ref8">8, 8</xref>
          ) = 11.
        </p>
        <p>To obtain the arcs of  from the weights of the direct connections of the  filters, we
adopt some statistical descriptors. Our objective is to have only one set of arcs from the node
corresponding to ℐ(, ) to the node corresponding to ( + ,  + ), −  ≤  ≤ , −  ≤  ≤ 1.
The weight of an arc from ℐ(, ) to ( + ,  + ) is obtained by applying a suitable statistical
descriptors to the weights of ℎ(, ), 1 ≤ ℎ ≤ . The two statistical descriptors we have chosen
to adopt are the mean and the median, which achieved the best performance results.</p>
        <p>As a consequence of the previous reasoning, the set  of the arcs of  consists of a set
of subsets  = {1, 2, ..., − 1}. Here,  denotes the set of arcs connecting nodes of
 to nodes of +1. Analogously, the set  of the weights of  consists of a set of subsets
 = {1, 2, ..., − 1}. Here,  is the set of weights associated with the arcs of .</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Mapping a CNN into a multilayer network</title>
        <p>After showing how it is possible to build a class network representing a CNN, in this section we
illustrate how, with the support of the class network model and a dataset containing the training
data, it is possible to build a multilayer network capable of fully mapping both the CNN and
its behavior. Roughly speaking, a multilayer network is a set of  class networks, one for each
target class in the dataset. Formally speaking, let  be a dataset consisting of  target classes
1, 2, . . . , , and let  be a Convolutional Neural Network. The multilayer network
 = {1, 2, ...} corresponding to  is a set of  class networks such that ℎ, 1 ≤ ℎ ≤ ,
corresponds to the ℎℎ target class of .</p>
        <p>
          Figure 2 shows a multilayer network  characterized by three layers, 1, 2 and 3 each
consisting of a generated class network. Figure 3 shows the generation of a portion of  obtained
by extracting the three feature maps of three target classes from . In this last network, a
set of filters of size 3 × 3 with stride 1, sliding over positions (
          <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
          ) and (
          <xref ref-type="bibr" rid="ref2 ref4">4, 2</xref>
          ) of the feature
maps, generates some arcs for the class networks of . The weights of the arcs are computed
by applying the rules described in Section 2.1. In particular, for the class network 1 (resp., 2,
3), two sets of arcs of weights 5 (resp., 1, 3) and 4 (resp., 7, 6) are generated between the first
and second feature maps. Similarly, a set of arcs of weights 2 (resp., -2, 4) and 8 (resp., 9, -1) are
generated between the second and the third feature maps.
        </p>
        <p>Our algorithm for constructing  from  and  consists of two steps. During the first
step, it creates a support data structure consisting of a list of patch lists. A patch list is a part of
a feature map having the same size as the filters applied to the next convolutional layer and
giving rise to a node in the multilayer network. The list of patch lists is then used during the
second step to construct .</p>
        <p>The function corresponding to the first step, called CREATE_PATCHES, receives  and
a target class ℎ, 1 ≤ ℎ ≤ . It operates on the feature maps provided in input to each
convolutional layer of  and proceeds as follows. It uses a list conv_layers that initially
contains the convolutional layers of . Afterwards, it iterates over the elements of conv_layers
and provides them with the images of ℎ. During each iteration, it treats the current element
as the source and the next element as the target. The feature map returned from source is given
as input to target, which processes it as specified below.</p>
        <p>At the beginning of each iteration, CREATE_PATCHES determines the starting and ending
points of the source output. After that, it iterates over the source output and creates a patch for
1Here and in the following, we employ the symbols ℐ(, ) and (, ) for indicating both the elements of the
feature maps and the corresponding nodes of the class network.
each application of the convolutional filter on an element. At the end of the iteration, it stores
the patches corresponding to a convolutional layer in a list called patch_list. Finally, it returns
in output the lists corresponding to all convolutional layers of .</p>
        <p>The function corresponding to the second step, called CREATE_LAYER_NETWORK receives 
and a target class ℎ, 1 ≤ ℎ ≤  of the multilayer network . It first calls CREATE_PATCHES
to receive the list of patch lists corresponding to  and ℎ. Then, it creates an initially
empty network ℎ. Afterwards, it scrolls the list of patch lists returned by CREATE_PATCHES
considering the current list as source and the next one as target.</p>
        <p>For each iteration, it adds to ℎ a node for each patch present in source or target, if they are
not already present in ℎ. Then, it determines the size of the area covered by the filter in target.
This size is equal to the one covered in source, if we are in presence of a convolutional layer, or
it is greater, if we have a pooling layer. After that, it iterates over the source and target nodes on
which the filter acts. Given a node, it considers all source nodes that can be processed by the
iflter whose center falls in the rectangle defined by the coordinates of the patch of  and, for
each of them, adds an arc from it to  in ℎ.</p>
        <p>Once this arc has been entered, it computes the corresponding weights by applying the
formula seen above (see Figure 3 for an example of its application). At the end of its iterations,
CREATE_LAYER_NETWORK has created the ℎℎ layer ℎ of , corresponding to the target class
ℎ. Applying this function  times, one for each target class in the dataset , we obtain the
ifnal multilayer network.</p>
        <p>(a) Class network 1
(b) Class network 2
(c) Class network 3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Applying the multilayer network model to compress a CNN</title>
      <p>Once a CNN is mapped into a multilayer network, the latter can be used to analyze and
manipulate the former. There are several operations that can be performed on the CNN thanks
to its mapping in a multilayer network. To give an idea of such operations, in this section we
examine one of them, namely the compression of a CNN. This represents the second main
contribution of this paper.</p>
      <p>Let  be a multilayer network and let ℎ be its ℎℎ layer. ℎ is a weighted directed graph.
Given a node  of ℎ, we can define: (i) the indegree of , as the sum of the weights of the arcs
of ℎ incoming into ; (ii) the outdegree of , as the sum of the weights of the arcs outgoing
from ; (iii) the degree of , as the sum of its indegree and its outdegree. We adopt the symbol
ℎ() to denote the degree of  in ℎ and the symbol  () to represent the overall (weighted)
degree of  in . As we will see below,  () is an important indicator of the efectiveness of the
iflter represented by .</p>
      <p>As previously pointed out,  = {1, 2, . . . , } has a layer for each target class. As a
consequence, the overall degree  () of the node  in  can be obtained by suitably aggregating
the degrees 1(), 2(), · · · , () of  in the  layers of . More specifically, if we indicate
with ℱ a suitable aggregation function, we have that:</p>
      <p>() = ℱ (1(), 2(), ..., ())</p>
      <p>
        Let () = ∑︀ℎ=1 ℎ() be the sum of the degrees of  in the  layers of . We adopt the
following entropy-based aggregation function [10, 11] for determining  ():
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
 () = −
ℎ=1
∑︁ ℎ()
()

︂( ℎ() )︂
()
,
      </p>
      <p>This function refers to the famous concept of information entropy introduced by Shannon
[12]. It favors the presence of a uniform distribution of the degree of  in the diferent layers,
while it penalizes the presence of a high degree of  in few layers and a low degree of  in many
layers. In this way, we favor those nodes whose feature extraction is balanced for diferent
target classes [13] and penalize those nodes that do not favor such a balance.</p>
      <p>Our approach for the compression of  selects a subset of the nodes of  with the highest
values of  . This is equivalent to selecting the convolutional layers of  that contribute the
most to the quality of the classification and, therefore, cannot be discarded in any way. In
particular, our approach selects the nodes of  whose values of  are higher than a certain
threshold ℎ =  ·  .</p>
      <p>Here,  denotes a statistical descriptor of the values of the overall degree  of all nodes in .
Our approach allows the usage of the mean or the median as statistical aggregators because
they are the ones that allowed us to obtain the best experimental results.  is a scaling factor
that allows us to tune the contribution of  . Its value belongs to the real interval [0, +∞).</p>
      <p>After selecting the subset of the nodes of  with an overall degree  higher than ℎ , our
approach determines the set of the convolutional layers of  from which these nodes were
extracted. Finally, it keeps these layers in the compressed version of  while discarding the
others. At this point, our approach terminates by training the pruned network.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In this paper, we have seen how it is possible to map a deep learning architecture, specifically a
CNN, into a multilayer network, which facilitates its analysis, exploration and manipulation.
To this end, we proposed an approach to map each element of a CNN into the elements of the
multilayer networks, namely nodes, arcs, arc weights and layers. Then, to give an idea of what
can be done thanks to such a mapping, we proposed an approach to compress a CNN based on
the pruning of the layers that least contribute to the quality of the classification.</p>
      <p>The approach proposed in this paper is a starting point for further research eforts in this area.
For example, our approach does not consider the residual connections, typical of ResNet. In the
future, we plan to extend it in order to handle this type of architecture. Also, at the moment, our
approach is not able to prune single filters from a convolutional layer. Doing so would require a
much more detailed mapping of the CNN into the multilayer network. At the moment, we have
not done this to keep the size of the multilayer network limited. However, conceptually there
are no limitations in performing such a task. In the future, we plan to proceed in that direction
so that we can perform the pruning of a CNN with finer granularity.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dargan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Ayyagari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>A survey of deep learning and its applications: A new paradigm to machine learning</article-title>
          ,
          <source>Archives of Computational Methods in Engineering</source>
          <volume>27</volume>
          (
          <year>2020</year>
          )
          <fpage>1071</fpage>
          -
          <lpage>1092</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Corradini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Porcino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Scopelliti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ursino</surname>
          </string-name>
          , L. Virgili,
          <article-title>Fine-tuning SalGAN and PathGAN for extending saliency map and gaze path prediction from natural images to websites</article-title>
          ,
          <source>Expert Systems With Applications</source>
          <volume>191</volume>
          (
          <year>2022</year>
          )
          <fpage>116282</fpage>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Deep neural network acceleration based on low-rank approximated channel pruning</article-title>
          ,
          <source>IEEE Transactions on Circuits and Systems I: Regular Papers</source>
          <volume>67</volume>
          (
          <year>2020</year>
          )
          <fpage>1232</fpage>
          -
          <lpage>1244</lpage>
          . doi:
          <volume>10</volume>
          .1109/TCSI.
          <year>2019</year>
          .
          <volume>2958937</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Discrimination-aware network pruning for deep model compression</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          . doi:
          <volume>10</volume>
          .1109/TPAMI.
          <year>2021</year>
          .
          <volume>3066410</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Choudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sarangapani</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey on model compression and acceleration</article-title>
          ,
          <source>Artificial Intelligence Review</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.-R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <article-title>Joint architecture and knowledge distillation in CNN for Chinese text recognition</article-title>
          ,
          <source>Pattern Recognition</source>
          <volume>111</volume>
          (
          <year>2021</year>
          )
          <fpage>107722</fpage>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sohail</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Zahoora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Qureshi</surname>
          </string-name>
          ,
          <article-title>A survey of the recent architectures of deep convolutional neural networks</article-title>
          ,
          <source>Artif. Intell. Rev</source>
          .
          <volume>53</volume>
          (
          <year>2020</year>
          )
          <fpage>5455</fpage>
          -
          <lpage>5516</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s10462-020-09825-6.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kivela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arenas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Barthelemy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Gleeson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Porter</surname>
          </string-name>
          ,
          <article-title>Multilayer networks</article-title>
          ,
          <source>Journal of Complex Networks</source>
          <volume>2</volume>
          (
          <year>2014</year>
          )
          <fpage>203</fpage>
          -
          <lpage>271</lpage>
          . doi:
          <volume>10</volume>
          .1093/comnet/cnu016.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Shallowing deep networks: Layer-wise pruning based on feature representations</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>41</volume>
          (
          <year>2019</year>
          )
          <fpage>3048</fpage>
          -
          <lpage>3056</lpage>
          . doi:
          <volume>10</volume>
          .1109/TPAMI.
          <year>2018</year>
          .
          <volume>2874634</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>R. K. Manel Hmimida</surname>
          </string-name>
          ,
          <article-title>Community detection in multiplex networks: A seed-centric approach</article-title>
          ,
          <source>Networks &amp; Heterogeneous Media</source>
          <volume>10</volume>
          (
          <year>2015</year>
          )
          <fpage>71</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Battiston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nicosia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Latora</surname>
          </string-name>
          ,
          <article-title>Structural measures for multiplex networks</article-title>
          ,
          <source>Phys. Rev. E</source>
          <volume>89</volume>
          (
          <year>2014</year>
          )
          <article-title>032804</article-title>
          . doi:
          <volume>10</volume>
          .1103/PhysRevE.89.032804.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Shannon</surname>
          </string-name>
          ,
          <source>A mathematical theory of communication</source>
          ,
          <source>The Bell System Technical Journal</source>
          <volume>27</volume>
          (
          <year>1948</year>
          )
          <fpage>379</fpage>
          -
          <lpage>423</lpage>
          . doi:
          <volume>10</volume>
          .1002/j.1538-
          <fpage>7305</fpage>
          .
          <year>1948</year>
          .tb01338.x.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N.</given-names>
            <surname>Gowdra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sinha</surname>
          </string-name>
          , S. MacDonell,
          <string-name>
            <given-names>W. Q.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Mitigating severe over-parameterization in deep convolutional neural networks through forced feature abstraction and compression with an entropy-based heuristic</article-title>
          ,
          <source>Pattern Recognition</source>
          (
          <year>2021</year>
          )
          <fpage>108057</fpage>
          . Elsevier.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>