<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Medico 2021: Medical Image Augmentation and Segmentation using Combination of Segmentation Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Zeshan Khan</institution>
          ,
          <addr-line>Mubasher Khan, Mubashir Yasin, Muhammad Hassan</addr-line>
          ,
          <institution>Muhammad Atif Tahir FAST School of Computing, National University of Computer and Emerging Sciences</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Polyp identification is a critical task for the pre-detection of colon cancer. Proper removal of a polyp requires accurate estimation of the size and shape of the polyp. The identification of the size and shape of the polyp can be done using polyps segmentation. This research investigated various polyps segmentation approaches evaluated on the benchmark dataset of Kvasir. The best results by the use of UNet++ architecture on the augmented data resulted in an accuracy of 0.92 with a Dice coeficient of 0.53.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Capsule endoscopy has been used for endoscopic abnormalities
diagnostics for more than 10 years. Endoscopic images provide
diagnosis capability for the detection of several abnormalities including
various types of cancers in the Gastrointestinal Tract (GI-Tract),
ulcer, and polyps detection. The analysis of such video frames takes
a lot of time of medical experts, which can be reduced by the use of
computer-aided diagnostics. with the increase in processing powers
of computational machines, deep learning (DL) based automated
diagnostics can result in good accuracy and eficiency.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>The GI-Tract disease detection is an active area of research with
the benefits of computer-aided diagnostics of various endoscopic
diseases. There are various works on the segmentation of the polyps
using neural network approaches.</p>
      <p>
        Jha et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] investigated the semantic segmentation of polyps
in the GI-Tract. An auto-encoder-based architecture of ResUNet
is used in the research for the segmentation of polyp. A modified
version of the ResUNet with the name ResUNet++ was proposed
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Trinh et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] used an auto-encoder-based approach with the
replacement of RelU with the leaky ReLU. The Leaky-ReLU is a
ReLU with some dead neurons enhances the results by ignoring
some of the neurons in the computation of ReLU. The network for
the encoder and decoder used by the authors is based on resnet50
trained on imagenet [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The approach resulted in 0.95 accuracies
when tested on the MediaEval 2020 challenge dataset [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Brandao et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] converted Convolution Neural Network (CNN)
to Fully Connected Network (FCN) in their architecture for the
segmentation of the polyps. The basic idea of the detection was
deconvolution of the images before providing to the neural network
for the detection of pixels if these are polyps or not.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>The methodology of the research is data augmentation and
segmentation. The data augmentation is done by applying several
noises and reshaping methodologies on the images. The sequence
of the image operations for the noise and reshape are crop, adding
noise, horizontal and vertical flipping, mirroring, scaling, brightness
change, contrast, and sharpness. The parameters for the various
augmentation operations were random in a range such that the
output image should be of size 224 × 224 with three channels. These
operations resulted in augmented images. The augmentation is
applied to generate 1300 images from the training set and the same
operation is applied on the ground truths as well for the
augmentation.</p>
      <p>The segmentation of the image is done using various neural
networks and clustering techniques. The methodologies were
evaluated on the evaluation data, which was taken as 20% from training
data. The methodologies of the neural network approaches worked
better on the evaluation data than clustering techniques.</p>
      <p>
        The neural network approaches for the segmentation of the
images were based on auto-encoder architectures inspired by
UNet [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], U-Net++ [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], ResUNet++ [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and SegNet [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The various
auto-encoder approaches were evaluated on the validation dataset
and the approaches of UNet++ architecture gave good accuracy and
the Dice score. Based on the validation data, the best approaches
used for the segmentation of the test data were a 10-layered UNet++
based auto-encoder and a ResUNet++ auto-encoder.
      </p>
      <p>The UNet++ auto-encoder has 4 CNN layers in the encoder and
then a CNN layer as a Bridge then another set 4 of CNN with the
same filters as in the encoder but in reverse. The auto-encoder
output is again passed through a CNN layer to compute the masks
of the provided image. The architecture of the Unet++ is shown
in Figure 1. The auto-encoder is trained for the 100 epochs with a
batch size of 25 based on the availability of the resources. These
100 epochs training with loss function of binary cross-entropy and
learning rate of 0.0001 resulted in the best validation Dice and
accuracy.</p>
      <p>
        The auto-encoder build on the inspiration of the ResUNet++
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] consists of encoder and decoder using residual blocks in the
network. The encoder used one stem and three residual blocks of
three convolutions in each of the residual and a stem block with
some pooling and the fully connected layers. There is a similar
decoder block with the same number of convolution and pooling
layers with the addition of the attention layers in between them.
The decoder’s convolution layers are in a similar pattern as in
the encoder with the reverse order. The model is trained for the
200 epochs with the learning rate of 0.0001 with a batch size of 8
images. Similar to the ResNet++, the loss function of the binary
cross-entropy was used with the Adam optimizer.
      </p>
      <p>
        The overall methodology of the training can be expressed as the
data augmentation and the segmentation by the Figure 2.
The dataset for the research is provided by the Simula Lab as Hyper
Kvasir dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for the training and the test data as MediaEval
2021 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This data consists of 1360 RGB images for the segmentation
of the polyps with the ground truth masks for these images. The
images were of various dimensions from 352 × 449 to 1072 × 1024
with RGB channels. The test data consists of 200 unlabeled RGB
images of sizes varying from 576 × 576 to 1072 × 1072.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND ANALYSIS</title>
      <p>The various approaches for the polyps segmentation were done
on the training dataset with the 30% validation data from training.
The various image segmentation networks and the pixel
clustering approaches were investigated with the 70% training and 30%
validation dataset for the evaluation measure of accuracy and Dice
coeficient. The task rules were the submission of the 5 best runs.
So, the top 5 methodologies based on validation results were
selected for the submission. The oficial results from organizers show
the evaluations as in Table 1. The best results were received by
the application of the Unet++ with Adam optimizer on augmented
images.</p>
      <p>The P in Table 1 is used for the precision and R for the recall.
Table 1 shows a comparison of the test results for the top 5 approaches
applied for the polyps segmentation. There are three evaluation
measures used in the comparison. The Dice coeficient, pixel
accuracy, and the Jaccard Index. The Dice coeficient is the two times
ratio of the common area of both images over the total area of both
the images. The incorrect segmentation reduces the common area
a lot and afects the Dice coeficient a lot. So, the Dice coeficient
of the various approaches is 0.34 to 0.53 only. The pixel accuracy
is the proportion of the pixels classified correctly over the total
number of pixels. The polyps segmentation task has less than 20%
of the image as polyps and the rest 80% image is not the polyps.
So, most of the approaches where the polyps were present but not
detected correctly caused a lesser impact on accuracy and the
accuracy remained higher than 0.88. The Jaccard is a measure of the
intersection over the union. The incorrectly segmented images can
miss the huge portion of the intersection, polyps, that causes a low
Jaccard Index.
6</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSION AND FUTURE WORK</title>
      <p>
        The research is conducted using augmentation of the polyps images
and segmentation using CNN-based auto-encoder architecture. The
results of the approach show a good segmentation accuracy for
the validation results. The results of the segmentation for some of
the endoscopic images are not correct and the segmented region
for polyps in those images is of zero pixels. This problem of no
detection can be resolved by the application of detection before
segmentation. In the future, we will investigate the detection as a
pre-step for the segmentation of the polyps images based on Kvasir
polyps detection datasets [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The second problem in the results
is low segmentation accuracy for the images with light reflections
on the polyp [
        <xref ref-type="bibr" rid="ref1 ref14">1, 14</xref>
        ] which will be investigated for reflection removal
methodologies to improve the segmentation accuracy.
Medico: Transparency in Medical Image Segmentation
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Mojtaba</given-names>
            <surname>Akbari</surname>
          </string-name>
          , Majid Mohrekesh, Kayvan Najariani, Nader Karimi, Shadrokh Samavi, and
          <string-name>
            <given-names>SM Reza</given-names>
            <surname>Soroushmehr</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Adaptive specular reflection detection and inpainting in colonoscopy video frames</article-title>
          .
          <source>In 2018 25th IEEE International Conference on Image Processing (ICIP)</source>
          . IEEE,
          <fpage>3134</fpage>
          -
          <lpage>3138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Vijay</given-names>
            <surname>Badrinarayanan</surname>
          </string-name>
          , Alex Kendall, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Cipolla</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Segnet: A deep convolutional encoder-decoder architecture for image segmentation</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>39</volume>
          , 12 (
          <year>2017</year>
          ),
          <fpage>2481</fpage>
          -
          <lpage>2495</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Borgli</surname>
          </string-name>
          , Vajira Thambawita, Pia H Smedsrud, Steven Hicks, Debesh Jha, Sigrun L Eskeland, Kristin Ranheim Randel, Konstantin Pogorelov, Mathias Lux,
          <source>Duc Tien Dang Nguyen</source>
          , et al.
          <year>2020</year>
          .
          <article-title>HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy</article-title>
          .
          <source>Scientific data 7</source>
          ,
          <issue>1</issue>
          (
          <year>2020</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Brandao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Odysseas</given-names>
            <surname>Zisimopoulos</surname>
          </string-name>
          , et al.
          <year>2018</year>
          .
          <article-title>Towards a computedaided diagnosis system in colonoscopy: automatic polyp segmentation using convolution neural networks</article-title>
          .
          <source>Journal of Medical Robotics Research</source>
          <volume>3</volume>
          ,
          <issue>02</issue>
          (
          <year>2018</year>
          ),
          <fpage>1840002</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          .
          <volume>770</volume>
          -
          <fpage>778</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Steven</given-names>
            <surname>Hicks</surname>
          </string-name>
          , Debesh Jha, Vajira Thambawita, Hugo Hammer, Thomas de Lange, Sravanthi Parasa,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2021</year>
          . Medico Multimedia Task at MediaEval 2021:
          <article-title>Transparency in Medical Image Segmentation</article-title>
          .
          <source>In Proceedings of MediaEval 2021 CEUR Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          , Sharib Ali, Nikhil Kumar Tomar,
          <string-name>
            <surname>Håvard D Johansen</surname>
          </string-name>
          , Dag Johansen, Jens Rittscher,
          <article-title>Michael A Riegler,</article-title>
          and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Real-time polyp detection, localization and segmentation in colonoscopy using deep learning</article-title>
          .
          <source>Ieee Access</source>
          <volume>9</volume>
          (
          <year>2021</year>
          ),
          <fpage>40496</fpage>
          -
          <lpage>40510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          , Steven A Hicks, Krister Emanuelsen, Håvard Johansen, Dag Johansen, Thomas de Lange,
          <article-title>Michael A Riegler,</article-title>
          and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Medico multimedia task at mediaeval 2020: Automatic polyp segmentation</article-title>
          . arXiv preprint arXiv:
          <year>2012</year>
          .
          <volume>15244</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <surname>Pia H Smedsrud</surname>
          </string-name>
          ,
          <article-title>Michael A Riegler, Dag Johansen</article-title>
          , Thomas De Lange, Pål Halvorsen, and
          <string-name>
            <given-names>Håvard D</given-names>
            <surname>Johansen</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Resunet++: An advanced architecture for medical image segmentation</article-title>
          .
          <source>In 2019 IEEE International Symposium on Multimedia (ISM)</source>
          . IEEE,
          <fpage>225</fpage>
          -
          <lpage>2255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Michael Riegler, Pål Halvorsen, Steven Hicks, Kristin Ranheim Randel, Duc Tien Dang Nguyen, Mathias Lux, Olga Ostroukhova, and Thomas de Lange.
          <year>2018</year>
          .
          <article-title>Medico multimedia task at mediaeval 2018</article-title>
          .
          <source>In CEUR Workshop Proceedings</source>
          , Vol.
          <volume>2283</volume>
          . Technical University of Aachen, 1-
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Konstantin Pogorelov, Pål Halvorsen, Carsten Griwodz, Thomas Lange, Kristin Randel, Sigrun Eskeland,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen,
            <given-names>Mathias</given-names>
          </string-name>
          <string-name>
            <surname>Lux</surname>
            , and
            <given-names>Concetto</given-names>
          </string-name>
          <string-name>
            <surname>Spampinato</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Multimedia for medicine: the medico task at mediaeval</article-title>
          <year>2017</year>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Olaf</surname>
            <given-names>Ronneberger</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>
          . In International Conference on
          <article-title>Medical image computing and computer-assisted intervention</article-title>
          . Springer,
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Quoc-Huy</surname>
            <given-names>Trinh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minh-Van Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Thiet-Gia Huynh</surname>
          </string-name>
          , and
          <string-name>
            <surname>Minh-Triet Tran</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>HCMUS-Juniors 2020 at Medico Task in MediaEval 2020: Refined Deep Neural Network and U-Net for Polyps Segmentation</article-title>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Rui</surname>
            <given-names>Yao</given-names>
          </string-name>
          , Yilun Wu,
          <string-name>
            <surname>Wei</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiaolin Lin</surname>
          </string-name>
          , Shidan
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>and Su</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Specular reflection detection on gastroscopic images</article-title>
          .
          <source>In 2010 4th International Conference on Bioinformatics and Biomedical Engineering</source>
          . IEEE, 1-
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Zhou</surname>
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahman Siddiquee M.M.</surname>
          </string-name>
          ,
          <string-name>
            <surname>Tajbakhsh</surname>
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Liang</surname>
            <given-names>J.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>UNet++: A Nested U-Net Architecture for Medical Image Segmentation</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>11045</volume>
          (
          <year>2018</year>
          ). https://doi.org/10. 1007/978-3-
          <fpage>030</fpage>
          -00889-
          <issue>5</issue>
          _
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>