<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Generative Adversarial Networks for Automatic Polyp Segmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Awadelrahman M. A. Ahmed</string-name>
          <email>aahmed@ifi.uio.no</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper aims to contribute in bench-marking the automatic polyp segmentation problem using generative adversarial networks framework. Perceiving the problem as an image-to-image translation task, conditional generative adversarial networks are utilized to generate masks conditioned by the images as inputs. Both generator and discriminator are convolution neural networks based. The model achieved 0.4382 on Jaccard index and 0.611 as F2 score.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Developing an automated computer-aided diagnosis system for
polyp segmentation is indeed one of the potential solutions that can
assist colonoscopy, the current gold-standard medical procedure
for examining the colon, in the sense that can lessen the percentage
of the overlooked polyps. Motivated by that, this paper contributes
by examining a model based on generative adversarial networks
(GANs) to represent one of the benchmark methods that other
cutting edge new methods can be compared to. This paper discusses a
submission to Medico automatic polyp segmentation challenge [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
for task 1 which asks the participants to develop algorithms for
segmenting polyps on a comprehensive data set. In this work we
perceive the polyp segmentation problem as an image-to-image
translation problem, as we are given a gastrointestinal polyp
image and the task is to translate it to the corresponding mask that
locates the polyps. GANs framework have been successfully
implemented to solve image-to-image translation problem in many
application fields. Authors in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] investigated conditional GANs as
a general-purpose solution to image-to-image translation problems
and evaluated in diferent application fields such as translating
aerial images to maps and reconstructing objects from edge maps
or templates. The challenge task can be seen in the same way, hence
we adapted the model architecture suggested in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to fit the polyp
segmentation problem. The following sections will illustrate the
model details, evaluation and results followed by the conclusion
and future work.
      </p>
    </sec>
    <sec id="sec-2">
      <title>GANS FOR POLYP SEGMENTATION</title>
      <p>
        The GANs framework is a generative model scheme based on a
game-theoretic formulation for training data-synthesis models [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
A GAN consists of two models (e.g., neural networks) trained in
opposition to each other. A generator network  takes a random
noise vector as an input and maps it to an output which represents
fake data. The other network is the discriminator , which receives
that fake data and classifies it as fake data (i.e., generated by  ) or as
real data from a real dataset. Nevertheless, our particular task is to
map each gastrointestinal polyp image to the corresponding mask
that defines the polyps area, this description does not fit within
the standard GAN setting described above where the input is a
noise and the output is an arbitrary (but realistic-looking) image.
Therefore, for that we utilize a prominent variation of GANs which
produces an output that is conditioned by its input and hence called
conditional GAN.
      </p>
      <p>The polyp segmentation GAN-based model consists of two
networks. A generator takes the images as input and tries to produce
realistic-looking masks conditioned by this input, and a
discriminator which is basically a classifier which has the access to the ground
truth masks and tries to classify whether the generated masks are
real or not. To stabilize the training, the images are concatenated
with the masks (generated or real) before being fed to the
discriminator. The ultimate goal is to find the generator parameter set ∗
that satisfies the minimax optimization problem is shown in (1).
We particularly draw a batch of images  and their corresponding
masks  from their distributions  and  respectively. We feed
 to the generator  to produce a fake mask set  ( ). We then
concatenate both masks (real and fake) with the input images, this
is represented by . These concatenated pairs are then fed to the
discriminator . To learn these models, both networks are updated
in an adversarial fashion based on the discriminator output;
meaning that the discriminator aims to maximize the exact function
that the generator aims to minimize. Through training time, the
generator will be able to generated realistic-looking masks that are
good enough to deceive the discriminator, i.e. it can classify as real.
∗ =   E,∼, [ ( (, )]
+ E∼ [(1 −  ( (,  ( ))]
(1)</p>
      <p>The model block diagram is shown in Figure 1 and the model
details are shown in Table 1. Both networks are based on
convolution neural networks (CNNs). The generator has two segments,
one is based on convolution operations folllowed by deconvolution
layers with using skip connections.
3</p>
    </sec>
    <sec id="sec-3">
      <title>EXPERIMENTAL EVALUATION</title>
      <p>
        As required by the challenge, the data set used is the Kvaris-SEG [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
training data set which consists of 1000 image-mask pairs. We
ifrstly split the dataset to 800 training set and 200 validation set
to optimize the parameters. After we got satisfied with the model,
we re-trained the model on the whole 1000 data pairs and used to
produce the masks of a separate 160 test set that we submitted. The
model is implemented using Python[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and PyTorch[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] framework
on a 2×14-core Intel/128Gib machine.
As the images have diferent dimensions, we fit the images to the
mean width and height by cropping the larger images and padding
the smaller images with zeros. The pixel values are normalized
between -1 and 1. No data augmentation were used. The networks
in Table 1 were trained by optimizing the loss function in equation
1. The discriminator minimizes the negative log likelihood and the
generator minimizes the negative of that. We tried other options
like using feature matching, i.e. minimizing the loss between the
features of the discriminator pre-last layer and the original masks
instead of using the last classification layer output, but this seems to
hurt the performance. We used Adam optimizer for both networks
with learning rates of 0.002, for 12 epochs and batch size of 4. Some
samples of the generated masks for some epochs during training is
shown in Figure 2.
To quantify the overlap percentage between the ground truth mask
and our generated masks, we report both Jaccard index the Dice
similarity coeficient (DSC). We also report the per pixel recall,
precision, accuracy and F2 (giving more weight to recall). Results
on the test set are in Table 2. However it is dificult to comment on
those results objectively because of the uniqueness of the dataset
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], i.e. no previous publications to compare with, but in general
the higher recall than precision reflects the model accountability
for false-negatives which is desirable for this application. The test
throughput is 16 frames/sec on a 2×14-core Intel/128Gib machine.
      </p>
      <p>
        We report the model outputs for some samples in Figure 3. Even
though we do not have the access to the ground truth of the test
set, we may observe that the model incorrectly identified the polyp
location of the bottom two samples; whereas it did far better in
locating the polyps area of other samples (top and middle rows).
We do not have a clear explanation for that, but we speculate that
the small receptive field of the convolution layers could be a reason.
In other words, the convolution layers pay attention to the close by
area to each pixel and if this area is rich enough with features (e.g.
has sharp edges) the polyp can be more distinguishable. This why
incorporating attention mechanism [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] might help the model to
attend to fine details and larger ranges; hence we suggest studying
it as an extended model.
This paper aimed to benchmark the polyp segmentation problem
using GANs. The problem has been perceived as an image-to-image
translation task and we utilized conditional GANs architecture. The
model was able to learn the masks however higher performance can
be achieved by trying some improvements such as adding
reconstruction loss and increasing the dataset with data augmentation.
An interesting extension for this model is to try incorporating an
attention layer [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] which can help the convolution layers in both
generator and discriminator to attend to fine details and expand
the receptive field.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ian</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          , Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Generative adversarial nets</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>2672</volume>
          -
          <fpage>2680</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Phillip</given-names>
            <surname>Isola</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jun-Yan</surname>
            <given-names>Zhu</given-names>
          </string-name>
          ,
          <source>Tinghui Zhou, and Alexei A Efros</source>
          .
          <year>2017</year>
          .
          <article-title>Image-toimage translation with conditional adversarial networks</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          .
          <volume>1125</volume>
          -
          <fpage>1134</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          , Steven A.
          <string-name>
            <surname>Hicks</surname>
          </string-name>
          , Krister Emanuelsen, Håvard Johansen, Dag Johansen, Thomas de Lange,
          <article-title>Michael A</article-title>
          .
          <string-name>
            <surname>Riegler</surname>
          </string-name>
          , and Pål Halvorsen.
          <fpage>14</fpage>
          -
          <issue>15</issue>
          <year>December 2020</year>
          . Medico Multimedia Task at MediaEval 2020:
          <article-title>Automatic Polyp Segmentation</article-title>
          .
          <source>In Proc. of the MediaEval 2020 Workshop</source>
          , Online.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <surname>Pia H Smedsrud</surname>
          </string-name>
          ,
          <article-title>Michael A Riegler, Pål Halvorsen</article-title>
          , Thomas de Lange, Dag Johansen, and
          <string-name>
            <given-names>Håvard D</given-names>
            <surname>Johansen</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Kvasir-SEG: A segmented polyp dataset</article-title>
          .
          <source>In Proc. of International Conference on Multimedia Modeling</source>
          .
          <fpage>451</fpage>
          -
          <lpage>462</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Adam</given-names>
            <surname>Paszke</surname>
          </string-name>
          , Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang,
          <string-name>
            <surname>Zachary</surname>
            <given-names>DeVito</given-names>
          </string-name>
          , Zeming Lin, Alban Desmaison, Luca Antiga, and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Lerer</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automatic diferentiation in PyTorch</article-title>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Guido</given-names>
            <surname>Van</surname>
          </string-name>
          Rossum and
          <string-name>
            <surname>Fred L. Drake</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Python 3 Reference Manual</article-title>
          . CreateSpace, Scotts Valley, CA.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Han</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Ian Goodfellow, Dimitris Metaxas, and
          <string-name>
            <given-names>Augustus</given-names>
            <surname>Odena</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Selfattention generative adversarial networks</article-title>
          .
          <source>In International conference on machine learning. PMLR</source>
          ,
          <fpage>7354</fpage>
          -
          <lpage>7363</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>