<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HCMUS-Juniors at Medico Polyp Segmentation Task 2021: Eficient U-Net for Polyps Segmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quoc-Huy Trinh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trong-Hieu Nguyen Mau</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Van Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Van-Son Ho</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tan-Cong Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Triet Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Information Technology, University of Science</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>John von Neumann Institute</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Social Sciences and Humanities</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh city</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Medico task in the Mediaeval with the target to segment the Polyps in the endoscopic images. In this paper, we propose methods that use Eficient Unet and propose the Multiscale Eficient Unet to deal with this task. In the experiment, we also benchmark our method with others previous methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>With the developing of bio-medical and the information
technology, medical images now are stored on a digital database. Moreover,
with the increase of cases that have abnormal findings and
symptoms on the digestive system, it is necessary to have a system to
help the doctor accurately diagnose and detect the position of the
abnormal in the medical images. That is why many methods have
been proposed to help diagnose the polyps or the abnormal in the
digestive system through endoscopic images.</p>
      <p>
        On the other hand, the improvement of the Convolutional Neural
Network architecture leads to improving the task for the
segmentation of the medical images, and several architectures have been
proposed such as U-Net[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], PSP-Net[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], PraNet[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], etc. However,
there are many drawbacks in each method and need to be improved
and many challenges to the researchers to improve the performance
of their methods.[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
The goal of the Medico automatic polyp segmentation challenge
is to evaluate various methods for automatic polyp segmentation
that can be used to detect and mask out various types of polyps
(including irregular, small or flat polyps) with high accuracy[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In
this challenge, our goals are to segment the mask of all types of
polyps in the dataset.[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
      </p>
    </sec>
    <sec id="sec-2">
      <title>DATASET</title>
      <p>
        To evaluate our proposed method, we use the Hyper Kavsir dataset
proposed in 2020. This open dataset includes a comprehensive
multi-class image and video dataset for gastrointestinal endoscopy,
including the ground truth with mask and the bounding boxes value
for the multi-task on the endoscopic images.[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
In this task, we use the segmentation part of the Hyper Kvasir
dataset. This dataset consists of 1000 Ground Truth images with
masks to experiment on the segmentation tasks. With the test
dataset, we evaluate our on the test dataset of Medico task
organizer of Mediaeval.
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
      <sec id="sec-3-1">
        <title>We consider five solutions corresponding to our five submitted</title>
        <p>runs. To evaluate the performance of the proposed method, we also
compare the results with those from other methods, such as are</p>
      </sec>
      <sec id="sec-3-2">
        <title>ResUNet and PraNet.</title>
        <p>4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Propose Architecture</title>
      <p>In our proposed methods, we propose the architecture that uses the</p>
      <sec id="sec-4-1">
        <title>Eficient Net for the encoder block. Moreover, we propose using the</title>
        <p>low scale feature to get a better mask of the output and improve the
model’s performance, which is Multiscale Eficient U-Net. This
architecture includes three main blocks Multiscale Block (MC Block),</p>
      </sec>
      <sec id="sec-4-2">
        <title>EficientNet Encoder block (EEN Block), and Decoder Block.</title>
        <p>
          Initially, the input with the shape (, ℎ,  ) passes through the MC
Block; this block includes 3 Max Pooling layers, 2 Convolution 2D
layers, and 1 Batch Normalization layer. MC layers to create the
new low-scale feature for the model[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The parameter of the layers
can be changed to adapt to the feature representation of images.
From the output of the MC block, there are two types of features:
low scale and high scale features. Then these features pass the
encoder block created by the Eficient Encoder Block, with high scale
features, they map to the decoder block while low scale features
pass the encoder block.
        </p>
        <p>After passing the EEN Blocks, the feature will be concatenated at
the block Encoder 4, then continue to the Decoder Block.</p>
        <sec id="sec-4-2-1">
          <title>Input</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>Ground Truth</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Output</title>
          <p>t
e
N
t
n
e
iiff
c
E</p>
          <p>MC-Block</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>E-Block 1</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>E-Block 2</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>E-Block 3</title>
          <p>Jaccard
Loss</p>
        </sec>
        <sec id="sec-4-2-7">
          <title>D-Block 1</title>
        </sec>
        <sec id="sec-4-2-8">
          <title>Concat</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>D-Block 2</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>Concat</title>
        </sec>
        <sec id="sec-4-2-11">
          <title>D-Block 3</title>
        </sec>
        <sec id="sec-4-2-12">
          <title>Concat</title>
        </sec>
        <sec id="sec-4-2-13">
          <title>E-Block 4</title>
          <p>Max Pool
2D</p>
          <p>Multi-scale Block (MC-Block)
Conv2D NormBaalticzhation Max2DPool</p>
          <p>Conv2D</p>
          <p>Max
Pool 2D
By using MC Block, the model can use the low-scale feature to
enrich the feature in the learning process, particularly when the
model has to adapt to the small dataset.</p>
          <p>There is a limitation for this architecture because there are two
types of features to the encoder block; this is why this architecture
costs more computing resources than the traditional U-Net.
4.2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Loss Function</title>
      <p>To use the proposed architecture, we propose using the Jaccard</p>
      <sec id="sec-5-1">
        <title>Loss Function with the following formula [1]:</title>
        <p>+ Í  ∗ ˆ
  (, ˆ) =  ∗ (1 −  +
Í  + ˆ −  ∗ ˆ )
This loss function enable the segmentation process better and can
control the performance of model on the pitch of the tissues.
(1)
4.3</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Data Augmentation</title>
      <p>To enrich the dataset, we propose some augmentation methods. We
use Center Crop, Random Rotate, GridDistortion, Horizontal, and</p>
      <sec id="sec-6-1">
        <title>Vertical Flip to improve the quantity of the dataset.</title>
      </sec>
      <sec id="sec-6-2">
        <title>Following is the sample of the data after augmentation:</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>RESULTS</title>
      <p>Evaluation is on the test set of Mediaeval- Medico task, which
includes 200 images of Polyps in endoscopic images. The Benchmark
table shows that the Eficient Unet model with low features
performs better than the original Eficient Unet and PraNet overall.
However, if we compare Precision or Recall metrics, the model gets
a lower performance than the other two models.
The performance of the proposed architecture is positive. The mask
can cover almost all the tissue on the images, and it can cover cases
with dificult shapes.</p>
      <p>Method
EficientUnet
MCEU
PraNet
ResUnet</p>
      <p>Jaccard
0.6572
0.7059
0.6929
0.6739</p>
      <p>Dice
0.7425
0.7961
0.7774
0.7737</p>
      <p>Recall
0.7264
0.8167
0.8204
0.8371</p>
      <p>Precision
0.8442
0.8295
0.8160
0.7766</p>
      <p>Accuracy
0.9529
0.9565
0.9511
0.9495</p>
      <p>With the benchmark table, the Eficient Unet that uses the low
feature achieves the high score. The reason is the data for the
training and validation is the limitation, and augmentation can
be used to enrich the quantity of data. However, some features
can be as similar as the original sample. That is why the lower
scale feature can help the model adapt better to the low quantity of
sample dataset.
6</p>
    </sec>
    <sec id="sec-8">
      <title>CONCLUSION</title>
      <p>In general, we propose the Multiscale Eficient U-Net to deal with
the segmentation task. MCEU has the merit that can enrich the
feature for the training process. Moreover, this architecture can
help normalize the high-scale feature to help the model adapt to
the small dataset; however, some limitations exist. Regarding the
evaluation of the experiment, the result we achieved is quite
positive, compared to the PraNet, Res-UNet, and Eficient-UNet, our
model achieves better performance. This positive impact can help
the later architecture have another approach to deal with this task.</p>
    </sec>
    <sec id="sec-9">
      <title>ACKNOWLEDGMENT</title>
      <p>This research is funded by Vietnam National University Ho Chi</p>
      <sec id="sec-9-1">
        <title>Minh City (VNU-HCM) under grant number DS2020-42-01.</title>
        <p>Medico: Transparency in Medical Image Segmentation</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jeroen</given-names>
            <surname>Bertels</surname>
          </string-name>
          , Tom Eelbode, Maxim Berman, Dirk Vandermeulen, Frederik Maes, Raf Bisschops, and
          <string-name>
            <surname>Matthew</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Blaschko</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Optimizing the Dice Score and Jaccard Index for Medical Image Segmentation: Theory and Practice</article-title>
          .
          <source>Medical Image Computing and Computer Assisted Intervention - MICCAI</source>
          <year>2019</year>
          (
          <year>2019</year>
          ),
          <fpage>92</fpage>
          -
          <lpage>100</lpage>
          . https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -32245-8_
          <fpage>11</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Borgli</surname>
          </string-name>
          , Vajira Thambawita, Pia H Smedsrud, Steven Hicks, Debesh Jha,
          <string-name>
            <surname>Eskeland Sigrun</surname>
            <given-names>L</given-names>
          </string-name>
          , Kristin Ranheim Rand l, Konstantin Pogorelov, Mathias Lux, Duc Tien Dang Nguyen, Dag Johansen, Carsten Griwodz,
          <string-name>
            <surname>Stensland Håkon</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enrique</surname>
          </string-name>
          Garcia-Ceja, Peter T Schmidt, Hugo L Hammer,
          <article-title>Michael A Riegler, Pål Halvorsen</article-title>
          , and Thomas de Lange.
          <year>2020</year>
          .
          <article-title>HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy</article-title>
          .
          <source>Scientific Data</source>
          <volume>7</volume>
          ,
          <issue>1</issue>
          (
          <year>2020</year>
          ),
          <volume>283</volume>
          . https://doi.org/10.1038/s41597-020-00622-y
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Deng-Ping</surname>
            <given-names>Fan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge-Peng</surname>
            <given-names>Ji</given-names>
          </string-name>
          , Tao Zhou, Geng Chen, Huazhu Fu,
          <string-name>
            <given-names>Jianbing</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Ling</given-names>
            <surname>Shao</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>PraNet: Parallel Reverse Attention Network for Polyp Segmentation</article-title>
          .
          <source>In Medical Image Computing and Computer Assisted Intervention - MICCAI</source>
          <year>2020</year>
          , Anne L. Martel, Purang Abolmaesumi, Danail Stoyanov, Diana Mateus, Maria A.
          <string-name>
            <surname>Zuluaga</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kevin</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Daniel Racoceanu, and Leo Joskowicz (Eds.). Springer International Publishing, Cham,
          <fpage>263</fpage>
          -
          <lpage>273</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Steven</given-names>
            <surname>Hicks</surname>
          </string-name>
          , Debesh Jha, Vajira Thambawita, Hugo Hammer, Thomas de Lange, Sravanthi Parasa,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2021</year>
          . Medico Multimedia Task at MediaEval 2021:
          <article-title>Transparency in Medical Image Segmentation</article-title>
          .
          <source>In Proceedings of MediaEval 2021 CEUR Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Debesh</given-names>
            <surname>Jha</surname>
          </string-name>
          , Steven A.
          <string-name>
            <surname>Hicks</surname>
          </string-name>
          , Krister Emanuelsen, Håvard Johansen, Dag Johansen, Thomas de Lange,
          <article-title>Michael A</article-title>
          .
          <string-name>
            <surname>Riegler</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2020</year>
          . Medico Multimedia Task at MediaEval 2020:
          <article-title>Automatic Polyp Segmentation</article-title>
          .
          <source>In Proc. of the MediaEval 2020 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Pardha</given-names>
            <surname>Saradhi</surname>
          </string-name>
          Mittapalli and
          <string-name>
            <surname>Thanikaiselvan V.</surname>
          </string-name>
          <year>2021</year>
          .
          <article-title>Multiscale CNN with compound fusions for false positive reduction in lung nodule detection</article-title>
          .
          <source>Artificial Intelligence in Medicine</source>
          <volume>113</volume>
          (
          <year>2021</year>
          ),
          <volume>102017</volume>
          . https://doi.org/10.1016/j.artmed.
          <year>2021</year>
          . 102017
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Olaf</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>U-Net: Convolutional Networks for Biomedical Image Segmentation</article-title>
          .
          <source>LNCS 9351</source>
          ,
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          . https: //doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -24574-4_
          <fpage>28</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Olaf</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>U-Net: Convolutional Networks for Biomedical Image Segmentation</article-title>
          .
          <source>In Medical Image Computing and Computer-Assisted Intervention - MICCAI</source>
          <year>2015</year>
          ,
          <string-name>
            <given-names>Nassir</given-names>
            <surname>Navab</surname>
          </string-name>
          , Joachim Hornegger,
          <string-name>
            <surname>William M. Wells</surname>
          </string-name>
          , and Alejandro F.
          <source>Frangi (Eds.)</source>
          . Springer International Publishing, Cham,
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Hengshuang</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jianping</given-names>
            <surname>Shi</surname>
          </string-name>
          , Xiaojuan Qi,
          <string-name>
            <given-names>Xiaogang</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jiaya</given-names>
            <surname>Jia</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Pyramid Scene Parsing Network</article-title>
          .
          <source>CoRR abs/1612</source>
          .01105 (
          <year>2016</year>
          ). arXiv:
          <volume>1612</volume>
          .01105 http://arxiv.org/abs/1612.01105
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>