<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SCL-UMD at the Medico Task-MediaEval 2017: Transfer learning based Classification of Medical Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Taruna Agrawal</string-name>
          <email>taruna3@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rahul Gupta</string-name>
          <email>rahul.1987iit@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saurabh Sahu</string-name>
          <email>ssahu89@umd.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carol Espy Wilson</string-name>
          <email>espy@isr.umd.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Image classification</institution>
          ,
          <addr-line>transfer learning, Convolutional Neural Networks</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Speech and Communication Lab, University of Maryland</institution>
          ,
          <addr-line>College Park</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Detecting landmarks in medical images can aid medical diagnosis and is a widely researched problem. The Medico task at MediaEval 2017 addresses the problem of detecting gastrointestinal landmarks, keeping into consideration the amount of training data as well as the speed of the detection system. Since medical data is obtained from real-world patients, access to large amounts of data for training the models can be restricted. We therefore focus on a transfer learning approach, where we can borrow image representations yielded by other image classification/detection systems and then train a supervised learning schemes on the available annotated medical data. We borrow the state of the art deep learning classification schemes (VGGNet and Inception-V3 networks) to obtain representations for the medical images and use them in addition to the provided set of features. A joint model trained on all these features yields a Matthew's Correlation Coeficient (MCC) of 0.826 with an accuracy and F1-score values of 0.961 and 0.847, respectively.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The Medico task addresses the problem of detecting diseases based
on image signals from the gastrointestinal (GI) tract. The goal of
the task is to advance the application of machine learning tools
within the medical domain, while specifically focusing on the
detection of GI landmarks from images. Our approach in this task
involves leveraging the established frameworks for the detection
of real-world objects from images. Specifically, we borrow the state
of the art deep learning models in object classification to aid the
classification of medical images. Models such as VGGNet [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and
Inception-V3 [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] contain several convolutional, pooling and fully
connected layers and are typically trained on large amounts of
datasets. Training on these datasets yield models that can capture
various geometrical patterns in the input images and translate them
into features vectors that are then consumed by the final soft-max
layer for class prediction. We aim to harness the capability of such
deep networks by retaining the initial convolution filters and
pooling layers in these networks. We then obtain the representations
yielded by these networks towards the final layers of these
networks. This approach can be particularly useful in the cases with
limited amount of training data. Since medical domain data is
often obtained from real-world patients, training models on limited
resources is a requirement. We motivate our approach by
discussions of some related work in the next section, followed by the
description of the database and the methodology.
2
Transfer learning involves borrowing knowledge from related
domains to aid classification in a domain of interest [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Transfer
learning has been successfully applied in tasks such as human
behavioral understanding [
        <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
        ], developing deep architectures [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
and autonomous shaping [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Several recent works have leveraged
advances in image recognition and detection techniques to improve
diferent but related tasks. Shin et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] provide an overview of
CNN architectures, data characteristics and transfer learning. Li
et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] perform domain adaptation for object localization, using
VGGNet. Zheng et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] provide good practices for CNN
feature transfer, as we have used in our work. Other applications that
have used VGGNet and Inception-V3 based architectures include
Alzheimer’s disease classification [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], disambiguation for large
scene classification [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and plant classification [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The success of
these CNN based transfer learning inspire our experiments.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>DATABASE</title>
      <p>
        We use the dataset provided as part of the The 2017 Multimedia for
Medicine Task (Medico) task during MediaEval benchmarking
initiative 2017 [
        <xref ref-type="bibr" rid="ref10 ref13">10, 13</xref>
        ]. The dataset consists of 8000 images of the GI tract
which are annotated and verified by experienced medical doctors
into eight diferent anatomical landmarks. We use the suggested
split of 4000 images as training set and the remaining images as
the testing test for the purpose of our experiments. The training
dataset contains a balanced number of instances per class. More
details regarding the dataset can be found in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHODOLOGY</title>
      <p>
        Deep learning models have achieved state of the art performance
in several image classification related tasks. In particular,
Convolutional Neural Networks (CNN) have provided the best
performances on tasks such as object classification [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], detection [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
and tracking [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Inspired from these developments, we obtain a set
of features from popular CNN designs, in addition to the provided
set of features. First, we describe the set of features used in our
experiments, followed by the classification setup.
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Features</title>
      <p>∗Independent authors.</p>
      <p>
        We use an assembly of features provided as part of the challenge as
well as a few CNN based features. We discuss these features below.
Baseline features The task provides a set of features extracted
on the images such as Tamura, ColorLayout, EdgeHistogram and,
AutoColorCorrelogram. Each of these features is a global descriptor
of the image. Note that these features quantify a specific property
of each image, which may or may not be associated with the final
classification task. On the other hand, CNN architectures learn
to extract features relevant to the task at hand, although it may
be hard to interpret those features. More details regarding these
features can be found in the task paper [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        VGGNet based features We use the VGGNet [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] pre-trained on
ImageNet dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] as a feature extractor. Since the ImageNet
dataset contains a large number of training samples, we expect
the VGGNet dataset to be able to model a large variety of shape
patterns in the images. We hypothesize that this characteristic of
the VGGNet can be useful in the Medico task. We use the 16 layer
configuration of VGGNet to predict the outcomes on the ImageNet
dataset (configuration D in Table 1, [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). After this pre-training,
we provide the Medico task images as input to the trained VGGNet.
Note that each image is of a diferent size and is rescaled to 244 ×244
in order to be fed to the VGGNet network. We use the outputs
from the first fully connected layer as features for the classification
task at hand. The dimensionality of the outputs from the first fully
connected layer is 4096.
      </p>
      <p>
        Inception-V3 features Similar to the VGGNet based features, we
extract features from the Inception-V3 network [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Inception-V3
consists of a stack of convolutional layers and pooling stacked
together and the features we obtain are obtained from the penultimate
layer. We again resize the Medico task images to 139×139 pixels.
The dimensionality of the features obtained from the penultimate
layer is 2048. We next describe the classification setup to predict
the anatomical landmarks.
4.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Classification setup</title>
      <p>We evaluate three diferent classification setups, with diferent
combinations of features. We train a multi-class Support Vector
Machine (SVM) classifier on the following combinations:
• Baseline + Inception-V3 features
• Baseline + VGGNet features and
• Baseline + Inception-V3 + VGGNet features
The hyper-parameters for the SVM classifier was tuned using five
fold cross-validation framework on the training dataset. We tuned
the SVM box-constraint parameter as well as the kernel. The linear
kernel performs the best, suggesting that further non-linear
transformation of the features is not required. In the next section, we
present the obtained results.</p>
    </sec>
    <sec id="sec-6">
      <title>5 RESULTS</title>
      <p>We present our results for each set of features in Table 1. The
evaluation metric used in the challenge is a multi-class generalization
of Matthew’s Correlation Coeficient ( Rk ). From the results, we
observe that the combination including all sets of features performs
the best. Since the evaluation metric also takes into account the
amount of training data used, we consistently use only 3200 samples
out of the 4000 samples for training. We also provide the class-wise
confusion matrix in Table 2. We observe that most of the confusion
lies between the classes Normal z-line and Esophagitis. We aim to
investigate this class confusion in future to reduce the error rate.</p>
      <p>In order to further understand the complementarity of the three
feature sets used in our experiments, we performed another
crossvalidation experiments. We used only one out of the baseline,
Inception-V3 based and VGGNet based feature sets and evaluate
their performance. We observed that the Inception-v3 and VGGNet
based features outperform the baseline features, indicating that
these CNN based features capture better representation in the
images. This may be due to the fact that they are trained on a larger
(albeit mismatched) corpus and can model a larger number of
geometrical shapes in the images, as compared to the baseline features.
6</p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSION</title>
      <p>The Medico task at MediaEval-2017 challenge addresses the problem
of detecting GI landscapes from images. The task focuses on training
limited amount of dataset with fast evaluation. We address this
problem by adopting a transfer learning method, borrowing
pretrained CNN architectures, successfully applied to other image
detection and classification problems. We borrow features extracted
from VGGNet and Inception-V3 models and train a supervised
algorithm along with the provided baseline features. With only
3200 training samples, we obtain an MCC value of 0.826.</p>
      <p>
        In the future, we aim to add more sources of transfer learning
to this task. Recently further modifications have been proposed to
deep CNN architectures such as GoogLeNet [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], ResNet [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and
generative adversarial networks [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We also aim to experiment
with ensemble methods to fuse the prediction from these network
based features, along with the simple feature fusion in this paper.
Since each of the CNN networks carry a diferent methodology for
convolution and pooling, we also aim to understand the
discriminative power of each of these feature sets independently. Finally,
we also aim to extend the proposed methodology to more image
classification tasks within the medical domain.
      </p>
      <p>Medico challenge, MediaEval 2017</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sabyasachee</given-names>
            <surname>Baruah</surname>
          </string-name>
          , Rahul Gupta, and
          <string-name>
            <surname>Shrikanth</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Narayanan</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Knowledge Transfer and Boosting Approach to the Prediction of Afect in Movies</article-title>
          .
          <source>In Proceedings of IEEE International Conference on Audio, Speech and Signal Processing (ICASSP).</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          and others.
          <year>2009</year>
          .
          <article-title>Learning deep architectures for AI. Foundations and trends</article-title>
          ® in
          <source>Machine Learning</source>
          <volume>2</volume>
          ,
          <issue>1</issue>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>127</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          .
          <article-title>IEEE Conference on</article-title>
          . IEEE,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>George</given-names>
            <surname>Konidaris</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Barto</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Autonomous shaping: Knowledge transfer in reinforcement learning</article-title>
          .
          <source>In Proceedings of the 23rd international conference on Machine learning. ACM</source>
          ,
          <volume>489</volume>
          -
          <fpage>496</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Matej</given-names>
            <surname>Kristan</surname>
          </string-name>
          , Jiri Matas, Ales Leonardis, Michael Felsberg, Luka Cehovin, Gustavo Fernández, Tomas Vojir, Gustav Hager, Georg Nebehay, and
          <string-name>
            <given-names>Roman</given-names>
            <surname>Pflugfelder</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The visual object tracking vot2015 challenge results</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision workshops</source>
          . 1-
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Sue</given-names>
            <surname>Han</surname>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          ,
          <source>Yang Loong Chang, Chee Seng Chan, and Paolo Remagnino</source>
          .
          <year>2016</year>
          .
          <article-title>Plant Identification System based on a Convolutional Neural Network for the LifeClef 2016 Plant Classification Task.</article-title>
          .
          <source>In CLEF (Working Notes)</source>
          .
          <fpage>502</fpage>
          -
          <lpage>510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Dong</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jia-Bin</surname>
            <given-names>Huang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yali</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shengjin</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ming-Hsuan Yang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Weakly supervised object localization with progressive domain adaptation</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>3512</fpage>
          -
          <lpage>3520</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Qinyi</given-names>
            <surname>Luo</surname>
          </string-name>
          , Rahul Gupta, and
          <string-name>
            <given-names>Shrikanth</given-names>
            <surname>Narayanan</surname>
          </string-name>
          .
          <article-title>Transfer Learning between Concepts for Human Behavior Modeling: An Application to Sincerity and Deception Prediction.</article-title>
          . In Interspeech,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Sinno</given-names>
            <surname>Jialin</surname>
          </string-name>
          Pan and
          <string-name>
            <given-names>Qiang</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A survey on transfer learning</article-title>
          .
          <source>IEEE Transactions on knowledge and data engineering 22</source>
          , 10 (
          <year>2010</year>
          ),
          <fpage>1345</fpage>
          -
          <lpage>1359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato,
          <string-name>
            <surname>Duc-Tien</surname>
          </string-name>
          Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt, and others.
          <source>2017</source>
          .
          <article-title>Kvasir: a multi-class image dataset for computer aided gastrointestinal disease detection</article-title>
          .
          <source>In Proceedings of the 8th ACM on Multimedia Systems Conference. ACM</source>
          ,
          <volume>164</volume>
          -
          <fpage>169</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Alec</surname>
            <given-names>Radford</given-names>
          </string-name>
          , Luke Metz, and
          <string-name>
            <given-names>Soumith</given-names>
            <surname>Chintala</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Unsupervised representation learning with deep convolutional generative adversarial networks</article-title>
          .
          <source>arXiv preprint arXiv:1511.06434</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Shaoqing</surname>
            <given-names>Ren</given-names>
          </string-name>
          , Kaiming He,
          <string-name>
            <surname>Ross Girshick</surname>
            , and
            <given-names>Jian</given-names>
          </string-name>
          <string-name>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <string-name>
            <surname>Faster</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          :
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>91</volume>
          -
          <fpage>99</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Konstantin Pogorelov, PÃěl Halvorsen, Carsten Griwodz, Thomas de Lange, Kristin Ranheim Randel, Sigrun Losada Eskeland,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen,
            <given-names>Mathias</given-names>
          </string-name>
          <string-name>
            <surname>Lux</surname>
            , and
            <given-names>Concetto</given-names>
          </string-name>
          <string-name>
            <surname>Spampinato</surname>
          </string-name>
          .
          <source>Multimedia for Medicine: The Medico Task at MediaEval</source>
          <year>2017</year>
          ,. In MediaEval,
          <fpage>13</fpage>
          -15
          <source>September</source>
          <year>2017</year>
          , Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Saman</surname>
            <given-names>Sarraf</given-names>
          </string-name>
          , John Anderson, Ghassem Tofighi, and others.
          <year>2016</year>
          .
          <article-title>DeepAD: AlzheimerâĂš s Disease Classification via Deep Convolutional Neural Networks using MRI and fMRI</article-title>
          . bioRxiv (
          <year>2016</year>
          ),
          <fpage>070441</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Hoo-Chang</surname>
            <given-names>Shin</given-names>
          </string-name>
          , Holger R Roth,
          <string-name>
            <given-names>Mingchen</given-names>
            <surname>Gao</surname>
          </string-name>
          , Le Lu, Ziyue Xu, Isabella Nogues, Jianhua Yao, Daniel Mollura, and
          <string-name>
            <surname>Ronald M Summers</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning</article-title>
          .
          <source>IEEE transactions on medical imaging 35</source>
          ,
          <issue>5</issue>
          (
          <year>2016</year>
          ),
          <fpage>1285</fpage>
          -
          <lpage>1298</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Richard</surname>
            <given-names>Socher</given-names>
          </string-name>
          , Brody Huval, Bharath Bath,
          <string-name>
            <surname>Christopher D Manning</surname>
          </string-name>
          , and Andrew Y Ng.
          <year>2012</year>
          .
          <article-title>Convolutional-recursive deep learning for 3d object classification</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>656</volume>
          -
          <fpage>664</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Christian</surname>
            <given-names>Szegedy</given-names>
          </string-name>
          , Wei Liu, Yangqing Jia,
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Sermanet</surname>
          </string-name>
          , Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Rabinovich</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition. 1-9.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Christian</surname>
            <given-names>Szegedy</given-names>
          </string-name>
          , Vincent Vanhoucke, Sergey Iofe, Jon Shlens, and
          <string-name>
            <given-names>Zbigniew</given-names>
            <surname>Wojna</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Rethinking the inception architecture for computer vision</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>2818</fpage>
          -
          <lpage>2826</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Limin</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Sheng Guo, Weilin Huang,
          <string-name>
            <given-names>Yuanjun</given-names>
            <surname>Xiong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yu</given-names>
            <surname>Qiao</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Knowledge guided disambiguation for large-scale scene classification with multi-resolution CNNs</article-title>
          .
          <source>IEEE Transactions on Image Processing 26</source>
          ,
          <issue>4</issue>
          (
          <year>2017</year>
          ),
          <fpage>2055</fpage>
          -
          <lpage>2068</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Liang</surname>
            <given-names>Zheng</given-names>
          </string-name>
          , Yali Zhao,
          <string-name>
            <surname>Shengjin</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jingdong</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Qi</given-names>
            <surname>Tian</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Good practice in CNN feature transfer</article-title>
          .
          <source>arXiv preprint arXiv:1604.00133</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>