<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Inception-like CNN Architecture for GI Disease and Anatomical Landmark Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefan Petscharnig</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Klaus Schöfmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mathias Lux</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alpen-Adria-Universität Klagenfurt</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>In this working note, we describe our approach to gastrointestinal disease and anatomical landmark classification for the Medico task at MediaEval 2017. We propose an inception-like CNN architecture and a fixed-crop data augmentation scheme for training and testing. The architecture is based on GoogLeNet and designed to keep the number of trainable parameters and its computational overhead small. Preliminary experiments show that the architecture is able to learn the classification problem from scratch using a tiny fraction of the provided training data only.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        With recent developments in computer vision, it seems only natural
to transfer the progress to the domain of medical imaging. However,
in this domain where neither machines nor domain-experts provide
lfawless results [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], there is a need for specialized methods in order
to make a significant leap forward at tasks such as computer aided
diagnosis. The Medico task at MediaEval 2017 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] aims to improve
methods for multimedia-assisted diagnosis in the domain of
endoscopic imaging for the special case of the gastrointestinal (GI) tract.
More precisely, participants of this task should develop approaches
for GI disease and anatomical landmark detection. The goal of the
task is (eficient) classification of diseases with as little training data
as possible. To cope with this task, a rich training dataset
comprising of 4000 images (500 per class) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] was made available by the
task organizers. The individual performance is benchmarked on
a similarily sized test dataset. Recently, much efort in the field of
deep learning in medical image analysis has been conducted. For
the use case of surgical action and anatomical structure recognition
in laparoscopic interventions, diferent of-the-shelf architectures
have been investigated by our research group [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. For the use
case of cholecystectomy, Twinanda et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] altered a well-known
CNN architecture to suit the domain and use-case of temporal
segmentation. Within medical image analysis of the gastrointestinal
(GI) tract, Pogorelov et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] provide a system for disease detection.
Automated polyp detection in colonoscopy videos was proposed by
Tajbakhsh et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. For more information on automated please
refer to Litjens et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this paper, we propose an eficient
inception-like CNN architecture, which is capable of learning the
classification problem from scratch using only as small amount
of training data. To achieve this goal, we propose a crop-based
data augmentation scheme, which is tailored to the use case of GI
disease detection. We furthermore propose a variant of our
architecture increasing predictive performance at the expense of increased
computational cost and number of trainable parameters (model
size).
(a) Overview on the two variants of the proposed CNN
Architecture with max-pooling input preparation, stacked
inceptionlike modules for feature extraction and classification part.
      </p>
      <p>(b) Inception-like module structure.</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        We approach this problem by defining a CNN architecture that is
capable of learning the distinction of a relatively small amount of
classes from a small training set. We base our work on GoogLeNet [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
an already existing CNN architecture, which yields a decent
performance in many tasks. Its prominent architectural feature is the
inception module. The basic idea behind the inception module is
that the network may select at training time whether pooling, small
convolution, or wider convolution suits the underlying data best.
Therefore, the aforementioned operations are calculated in parallel
and their results are merged and the feature dimensionality is
reduced. Our inception module is depicted in Figure 1b and consists
of three main branches: small convolution, large convolution, and
pooling. We also use 1x1 convolutions before the main
convolutional layers in order to reduce computational cost. Furthermore, we
use padding in order to preserve the size of the feature maps. After
the main branches are merged channel-wise, we use max-pooling
with a stride of 2 for dimensionality reduction, reducing feature map
size by a factor of 4. After pooling, we use a LeakyReLU [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] with a
negative slope of 0.01 as non-linear activation function. We refer to
such inception-like modules as MX where X denotes the number of
S. Petscharnig, K. Schoefmann, M. Lux
learned filers per convolutional layer. Whereas GoogLeNet features
a 1x1 convolution branch, our preliminary experiments showed
that there is no performance gain for this specific domain, thus we
do not use a fourth branch. Furthermore, we only use activation
functions at the end of each module and not after each
convolutional layer and skip the 1x1 convolution after the pooling branch.
A graphical overview on our proposed architecture is given in
Figure 1a. The first layer of the proposed architecture is overlapping
max-pooling. Please note that we propose two diferent variants
of the architecture: Model A and Model B. The diference between
the variants lies in the first (max pooling) layer: Model A uses a
stride of 4 and max-pooling window size 5, whereas Model B uses
stride 2 window size 3. We then use three stacked inception-like
modules as feature extraction stage. The number of convolutions
per convolutional layer is doubled the deeper we delve into the
network to compensate for the spatial reduction. The feature
extraction stage is followed by a 1x1 convolution with 92 learned filters.
This convolutional step is a further dimensionality reduction in
the channels axis (whereas the inception-like modules reduce the
dimensionality in the spatial axis). The second convolution reduces
the feature map to a spatial size of 1x1. We also experiment with
the number of channels in the late stages and therefore, the output
of this layer has - depending on the actual choice - 1024 or 2048
channels. The convolution size for this layer is dependent of the
model variant (4x4 and 8x8 for Model A and Model B respectively).
The only regularization technique we use is a dropout layer (with
a dropout chance of 0.2). This is followed by a further 1x1
convolution layer and a dense layer with softmax activation. Henceforth,
we refer to our architecture variant with a combination of model
identifier (A or B) number of neurons in deep layers (1024 or 2048)
and percent of training data used, e.g., B1024−10 refers to model
B (3x3 pooling at low network size and 8x8 convolution before
dropout) with 1024 neurons in the deep layers and was trained on
10 percent of the training data. Input to our proposed network is a
128x128 image patch. We augment the training set by extracting
seven diferent image patches according to Figure 2 and resizing
them to the input shape. The patch selection was motivated from
the consideration that the most important image details are in the
center of the image. We furthermore extract the patches from three
diferent scales. Furthermore, we randomly mirror the patches at
training time. We standardize the training set by subtracting the
mean image pixel. For testing, we extract the a set image patches
from each testing image the same way as we augment our test data
(see Figure 2). We classify each of these patches and aggregate the
results by using a simple average.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RESULTS AND ANALYSIS</title>
      <p>
        We used three diferent variants for the tasks: Model A as well
as Model B with 1024 and 2048 neurons in the deep layers. An
overview on the individual results is given in Table 1. All in all,
we observe the trend that generalization performance increases
with more training data available. Generally, we discover that all
the models are confusing dyed resection margins with
dyed-liftedpolyps as well as polyps with ulcerative-colitis. We argue that this
weakness originates in the choice of training data augmentation:
polyps and resection margins are not always visible on center-like
crops. Our models also show minor weaknesses at distinguishing
normal-z-line from esophagitis. In preliminary experiments, we
also tried to distinguish these selected classes with a binary CNN
classifier and the fusion of global features at a deep level, but this did
not improve results. Model A is used in the speed runs. Its strength
is the small number of parameters (2.8M against 7.3M for Model
B1024, 16.5M for Model B2048) and its small computational cost.
We measure forward passes over 1000 iterations using a GeForce
GTX Titan X (maxwell) graphics card. The model takes 2.25ms per
forward pass, in contrast to 2.91ms and 3.42ms for Model B1024 and
B2048 respectively. Cafenet [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], an AlexNet variant takes 3.27ms per
forward pass, the computation time GoogLeNet 14.16ms. Variants
from Model B are used in the detection runs and As baseline in the
speed run. Interestingly, we observe that the lower capacity model
B1024 is superior to B2048. These results were surprising on first
view, as our preliminary evaluations indicated the opposite. We
conclude that the larger model tends to better adapt to the training
data. Thus, the model sufers from over-fitting and yields smaller
generalization performance.
      </p>
    </sec>
    <sec id="sec-4">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>We provide a CNN architecture capable of learning classification
with little training data. The proposed architecture is able to provide
acceptable results with even as little as 50 training examples per
class. The presented data augmentation and testing method using
a fixed set of selected crops per image is beneficial for the overall
performance. In preliminary experiments, we also investigated late
fusion of global feature to the CNN architecture which did not
lead to significant performance improvement. For future work, we
want to investigate two main aspects: (1) whether the architecture
is performing well in the domain of laparoscopic surgery, and (2)
whether the model is capable of eficiently dealing with more input
channels for early fusion of temporal information which can be
used in action recognition in laparoscopic surgery.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported by Universität Klagenfurt and Lakeside
Labs GmbH, Klagenfurt, Austria and funding from the European
Regional Development Fund and the Carinthian Economic Promotion
Fund (KWF) under grant KWF 20214 u. 3520/ 26336/38165.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jef</given-names>
            <surname>Donahue</surname>
          </string-name>
          .
          <year>2014</year>
          . BVLC Cafenet.
          <article-title>(</article-title>
          <year>2014</year>
          ). https://github.com/BVLC/cafe/tree/ master/models/bvlc_reference_cafenet Online, Accessed:
          <fpage>2017</fpage>
          -09-54.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Geert</given-names>
            <surname>Litjens</surname>
          </string-name>
          , Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian,
          <string-name>
            <surname>Jeroen A.W.M. van der Laak</surname>
          </string-name>
          , Bram van Ginneken, and
          <string-name>
            <surname>Clara</surname>
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sánchez</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A survey on deep learning in medical image analysis</article-title>
          .
          <source>Medical Image Analysis</source>
          <volume>42</volume>
          (
          <year>2017</year>
          ),
          <fpage>60</fpage>
          -
          <lpage>88</lpage>
          . https://doi.org/10. 1016/j.media.
          <year>2017</year>
          .
          <volume>07</volume>
          .005
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Andrew</surname>
            <given-names>L Maas</given-names>
          </string-name>
          , Awni Y Hannun, and Andrew Y Ng.
          <year>2013</year>
          .
          <article-title>Rectifier nonlinearities improve neural network acoustic models</article-title>
          .
          <source>In Proc. ICML</source>
          , Vol.
          <volume>30</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Petscharnig</surname>
          </string-name>
          and
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Schoefmann</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep Learning of Shot Classification in Gynecologic Surgery Videos</article-title>
          . In International Conference on Multimedia Modeling, Laurent Amsaleg, Gylfi Þór Guð mundsson,
          <source>Cathal Gurrin, Björn Þór Jónsson, and Shin'ichi Satoh (Eds.)</source>
          . Springer, Cham,
          <fpage>702</fpage>
          -
          <lpage>713</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Petscharnig</surname>
          </string-name>
          and
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Schoefmann</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning laparoscopic video shot classification for gynecological surgery</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          (apr
          <year>2017</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          . https://doi.org/10.1007/s11042-017-4699-5
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Sigrun Losada Eskeland, Thomas de Lange, Carsten Griwodz, Kristin Ranheim Randel, Håkon Kvale Stensland,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen, Concetto Spampinato, Dag Johansen,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Holistic Multimedia System for Gastrointestinal Tract Disease Detection</article-title>
          .
          <source>In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys'17)</source>
          . ACM, New York, NY, USA,
          <fpage>112</fpage>
          -
          <lpage>123</lpage>
          . https://doi.org/10.1145/3083187.3083189
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Kvasir: A Multi-Class Image Dataset for Computer Aided Gastrointestinal Disease Detection</article-title>
          .
          <source>In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys'17)</source>
          . ACM, New York, NY, USA,
          <fpage>164</fpage>
          -
          <lpage>169</lpage>
          . http://dx. doi.org/10.1145/3083187.3083212
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Mathias Lux, Carsten Griwodz, Concetto Spampinato, Thomas de Lange, Sigrun L. Eskeland, Konstantin Pogorelov, Wallapak Tavanapong, Peter T. Schmidt, Cathal Gurrin, Dag Johansen, Håvard Johansen, and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Multimedia and Medicine: Teammates for Better Disease Detection and Survival</article-title>
          .
          <source>In Proceedings of the 2016 ACM on Multimedia Conference (MM '16)</source>
          . ACM, New York, NY, USA,
          <fpage>968</fpage>
          -
          <lpage>977</lpage>
          . https://doi.org/10.1145/2964284.2976760
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Konstantin Pogorelov, Pål Halvorsen, Carsten Griwodz, Thomas de Lange, Kristin Ranheim Randel, Sigrun Losada Eskeland,
          <string-name>
            <surname>Duc-Tien</surname>
            <given-names>DangNguyen</given-names>
          </string-name>
          , Mathias Lux, and
          <string-name>
            <given-names>Cibcetti</given-names>
            <surname>Spampinato</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Multimedia for Medicine: The Medico Task at MediaEval 2017</article-title>
          .
          <source>In Proc of the MediaEval 2017 Workshop</source>
          . Dublin, Ireland, Sept.
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Christian</surname>
            <given-names>Szegedy</given-names>
          </string-name>
          , Wei Liu, Yangqing Jia,
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Sermanet</surname>
          </string-name>
          , Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Rabinovich</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 1-9.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tajbakhsh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Gurudu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <source>Automated Polyp Detection in Colonoscopy Videos Using Shape and Context Information. IEEE Transactions on Medical Imaging</source>
          <volume>35</volume>
          ,
          <issue>2</issue>
          (Feb
          <year>2016</year>
          ),
          <fpage>630</fpage>
          -
          <lpage>644</lpage>
          . https://doi.org/10.1109/TMI.
          <year>2015</year>
          . 2487997
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Twinanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shehata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Marescaux</surname>
          </string-name>
          , M. de Mathelin, and
          <string-name>
            <given-names>N.</given-names>
            <surname>Padoy</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos</article-title>
          .
          <source>IEEE Transactions on Medical Imaging</source>
          <volume>36</volume>
          ,
          <issue>1</issue>
          (Jan
          <year>2017</year>
          ),
          <fpage>86</fpage>
          -
          <lpage>97</lpage>
          . https: //doi.org/10.1109/TMI.
          <year>2016</year>
          .2593957
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>