<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DETECTION AND SEGMENTATION OF ENDOSCOPIC ARTEFACTS AND DISEASES USING DEEP ARCHITECTURES Nhan T. Nguyen , Dat Q. Tran , Dung B. Nguyen</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Medical Imaging Department, Vingroup Big Data Institute (VinBDI)</institution>
          ,
          <addr-line>Hanoi</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We describe in this paper our deep learning-based approach for the EndoCV2020 challenge, which aims to detect and segment either artefacts or diseases in endoscopic images. For the detection task, we propose to train and optimize EfficientDet-a state-of-the-art detector-with different EfficientNet backbones using Focal loss. By ensembling multiple detectors, we obtain a mean average precision (mAP) of 0.2524 on EDD2020 and 0.2202 on EAD2020. For the segmentation task, two different architectures are proposed: UNet with EfficientNet-B3 encoder and Feature Pyramid Network (FPN) with dilated ResNet-50 encoder. Each of them is trained with an auxiliary classification branch. Our model ensemble reports an sscore of 0.5972 on EAD2020 and 0.701 on EDD2020, which were among the top submitters of both challenges.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Disease detection and segmentation in endoscopic imaging
play an important role in the early detection of numerous
cancers, such as gastric, colorectal, and bladder cancers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Meanwhile, the detection and segmentation of endoscopic
artefacts is necessary for image reconstruction and quality
assertion [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Many approaches [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ] have been proposed
to detect and segment artefacts and diseases in endoscopy.
This paper describes our solution for the EndoCV2020
challenge, which consists of two tracks1: one deals with artefacts
(EAD2020) and the other one is for diseases (EDD2020)
. Each track is divided into two tasks: detection and
segmentation. We tackle both tasks in both tracks by exploiting
state-of-the-art deep architectures like EfficientDet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and
U-Net [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] with variants of EfficientNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and ResNet [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] as
backbones. In the next sections, we provide a short
description of the datasets, the details of the proposed approach, and
experimental results.
Fig. 1. The number of bounding boxes for each disease class
in training set provided by the EDD2020 dataset.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. DATASETS</title>
      <p>
        EDD2020 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a comprehensive dataset established to
benchmark algorithms for disease detection and segmentation
in endoscopy. It is annotated for 5 different disease classes,
including BE, Suspicious, HGD, Cancer, and Polyp. The
dataset comes with bounding boxes for disease detection and
with masked image annotations for semantic segmentation.
The training set includes total 386 endoscopy frames, each
of which is annotated with either single or multiple diseases.
Regions of the same class are merged into a single mask,
while a bounding box of multiple classes is treated as
separate boxes with the same location. Figure 1 shows the number
of bounding boxes for each disease class. EAD2020 [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ],
on the other hand, is used for the track of endoscopy artefact
detection and segmentation. The training set contains 2,531
annotated frames for 8 artefact classes, including
specularity, bubbles, saturation, contrast, blood, instrument, blur, and
imaging artefacts. Note that only first 5 classes are used for
the segmentation task.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. PROPOSED METHODS</title>
    </sec>
    <sec id="sec-4">
      <title>3.1. Multi-class detection task</title>
      <p>
        Detection network: For the detection task, we deployed
EfficientDet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], currently a state-of-the-art architecture for
object detection. It employs EfficientNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] as the backbone
network, BiFPN as the feature network, and shared class/box
prediction network. Both BiFPN layers and class/box net
layers are repeated multiple times based on different resource
constraints. Figure 3 illustrates the EfficientDet architecture.
Training procedure: Due to the limited training data
available (386 images in EDD2020 and 2531 images in EAD2020),
we use various data augmentation techniques, including
random shift, random crop, rotation, scale, horizontal flip,
vertical flip, blur, Gauss noise, sharpen, emboss, and contrast. In
particular, we found that the use of mixup could significantly
reduce the overfitting. Given x1 and x2 as input images, the
mixup image x~ is constructed as
x~ =
x1 + (1
      </p>
      <p>)x2;
x~ Netwo!rk y^:
During training, our goal is to minimize the MixLoss Lmixup,
which is expressed as</p>
      <p>Lmixup =</p>
      <p>
        L(y^; y1) + (1
)L(y^; y2):
(1)
where the symbol L denotes the Focal loss [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and is drawn
from (0:75; 0:75) distribution; y1 and y2 are the
groundtruth labels, while y^ is the predicted label produced by the
network. Fig. 4 visualizes a mixup example with being
fixed to 0.5.
      </p>
      <p>
        Our detectors are optimized by the gradient decent using
Adam update rule [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] with weight decay. In addition,
cyclical learning rate [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] with restarts is also used. The ensemble
of 6 models with different backbones (D0, D1, D2, D3, D4,
and D5) using weighted box fusion [15] serves as our final
model. Additionally, we search for the non-maximum
suppression (NMS) threshold and the confidence threshold for
different categories so that the resulting score (0:5 mAP +
0:5 IOU) is maximized.
      </p>
    </sec>
    <sec id="sec-5">
      <title>3.2. Multi-class segmentation task</title>
      <p>Segmentation network: We propose two different
architectures for this task: U-Net with EfficientNet encoders and
BiFPN with ResNet encoders.</p>
      <p>U-Net: Our first network design makes use of U-Net with
EfficientNetB3/B4 as backbones. We keep the original strides
between blocks in EfficientNet and extract the feature maps
from the last 5 blocks for the segmentation. A classification
branch is used to provide the label predictions. The overall
framework is depicted in Figure 2.</p>
      <p>BiFPN: To generate the segmentation output from the
BiFPN features, we combine all levels of the BiFPN pyramid
by following the design illustrated in Figure 5. Starting with
the deepest BiFPN level (stride-32 output), we apply three
upsampling stages to obtain the feature map of the stride-4
output. An upsampling stage consists of a 3 3 Convolution,
BatchNorm, ReLU and a 2 2 bilinear upsampling. This
strategy is repeated for other BiFPN levels with strides of 16,
8, and 4. The result is a set of feature maps at the same scale,
which are then channel-wise concatenated. Finally, a 1 1
Convolution, 4 4 bilinear upsampling and Sigmoid
activation are used to generate the mask at the image resolution.
Training procedure: All models are trained end-to-end
with additional supervision from the multi-label
classification task. The image labels are obtained directly from the
segmentation masks. For example, if an image has B.E. mask
annotation then the B.E. label is 1. Due to class imbalance
in the training dataset, we use Focal loss for the
classification task. Our final loss is L = Lseg + Lcls where = 0:4.
Inference: Relying solely on segmentation branch to
predict masks will result in high false positives. Hence, we make
use of the class predictions to remove masks. We search
optimal classification thresholds to maximize the macro F1 score
on the validation set. For every image, if the class probability
is less than the optimal threshold then its predicted mask is
completely removed.</p>
    </sec>
    <sec id="sec-6">
      <title>4. EXPERIMENTAL RESULTS</title>
      <p>Table 1 summarizes the detection and segmentation results
of our submissions for both challenges. We describe the
results of each sub-task below. Results on the validation set of
EDD2020 for the detection task are detailed in Table 2. Our
best single model (i.e. EfficientDet-D5) obtained a detection
score (dScore) of 0.41. The best detection performance was
provided by the ensemble model, which reported a dScore of
0.44, a mean mAP of 0.36 0.05, and an IoU of 0.52. As
shown in Table 1, our ensemble model yielded dScores of
0.2524 0.0948 and 0.2202 0.1029 on the hidden test sets of
EDD2020 and EAD2020, respectively.</p>
      <p>Results on validation sets for the segmentation task are
provided in Table 3 and Table 4. On the EDD2020 validation
set, our best single model achieved a Dice score of 0.854 and
an IoU of 0.832. On the EAD2020 validation set, we obtained
a Dice score of 0.732 and an IoU of 0.578. As shown in
TaChallenge
EAD2020
EDD2020
dscore
0.2202
0.2524
dstd
0.1029
0.0948
sscore
0.5972
0.7008
sstd
0.2765
0.3211
ble 1, our ensemble achieved a segmentation score (sscore) of
0.5972 in the EAD2020 challenge and an sscore of 0.7008 in
the EDD2020 challenge, both of which were among the top
results for the segmentation task of both tracks.
We have described our solutions for the detection and
segmentation tasks on both tracks of EndoCV2020: EAD for
artefacts and EDD for diseases. By using EfficientDet for
detection and U-Net/BiFPN for segmentation, we obtained
significant results on both datasets, especially for the
segmentation task. These results suggest that some of the deep
architectures that are effective for natural images can also be
useful for medical images like endoscopic ones, even with a
small-size training datasets.</p>
    </sec>
    <sec id="sec-7">
      <title>6. REFERENCES</title>
      <p>Adam: A
arXiv preprint
[15] Roman Solovyev and Weimin Wang. Weighted boxes
fusion: ensembling boxes for object detection models.
arXiv preprint arXiv:1910.13302, 2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sharib</given-names>
            <surname>Ali</surname>
          </string-name>
          , Noha Ghatwary, Barbara Braden, Lamarque Dominique, Adam Bailey, Stefano Realdon, Cannizzaro Renato, Jens Rittscher,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Daul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>James</given-names>
            <surname>East</surname>
          </string-name>
          .
          <article-title>Endoscopy disease detection challenge 2020</article-title>
          . CoRR, abs/
          <year>2003</year>
          .03376,
          <year>February 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Sharib</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>Felix Zhou</surname>
            , Christian Daul, Barbara Braden, Adam Bailey, Stefano Realdon, James East, Georges Wagnieres, Victor Loschenov,
            <given-names>Enrico</given-names>
          </string-name>
          <string-name>
            <surname>Grisan</surname>
          </string-name>
          , et al.
          <article-title>Endoscopy artifact detection (ead 2019) challenge dataset</article-title>
          .
          <source>arXiv preprint arXiv:1905.03209</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>PS</given-names>
            <surname>Hiremath</surname>
          </string-name>
          , BV Dhandra, Iranna Humnabad, Ravindra Hegadi, and
          <string-name>
            <given-names>GG</given-names>
            <surname>Rajput</surname>
          </string-name>
          .
          <article-title>Detection of esophageal cancer (necrosis) in the endoscopic images using color image segmentation</article-title>
          .
          <source>In Proceedings of second National Conference on Document Analysis and Recognition (NCDAR-2003)</source>
          , Mandya, India, pages
          <fpage>417</fpage>
          -
          <lpage>422</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Szczypin</surname>
          </string-name>
          ´ski, Artur Klepaczko, Marek Pazurek, and Piotr Daniel.
          <article-title>Texture and color based image segmentation and pathology detection in capsule endoscopy videos</article-title>
          .
          <source>Computer methods and programs in biomedicine</source>
          ,
          <volume>113</volume>
          (
          <issue>1</issue>
          ):
          <fpage>396</fpage>
          -
          <lpage>411</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Eva</given-names>
            <surname>Tuba</surname>
          </string-name>
          , Milan Tuba, and
          <string-name>
            <given-names>Raka</given-names>
            <surname>Jovanovic</surname>
          </string-name>
          .
          <article-title>An algorithm for automated segmentation for bleeding detection in endoscopic images</article-title>
          .
          <source>In 2017 International Joint Conference on Neural Networks (IJCNN)</source>
          , pages
          <fpage>4579</fpage>
          -
          <lpage>4586</lpage>
          . IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mingxing</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ruoming</given-names>
            <surname>Pang</surname>
          </string-name>
          , and Quoc V Le.
          <article-title>Efficientdet: Scalable and efficient object detection</article-title>
          .
          <source>arXiv preprint arXiv:1911.09070</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Olaf</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>
          .
          <source>In Medical Image Computing and Computer-Assisted Intervention - MICCAI</source>
          <year>2015</year>
          , pages
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          . Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Mingxing</given-names>
            <surname>Tan and Quoc V Le.</surname>
          </string-name>
          <article-title>Efficientnet: Improving accuracy and efficiency through automl and model scaling</article-title>
          .
          <source>arXiv preprint arXiv:1905.11946</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In IEEE CVPR</source>
          , pages
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Sharib</surname>
            <given-names>Ali</given-names>
          </string-name>
          , Felix Zhou, Adam Bailey, Barbara Braden, James East, Xin Lu, and
          <string-name>
            <given-names>Jens</given-names>
            <surname>Rittscher</surname>
          </string-name>
          .
          <article-title>A deep learning framework for quality assessment and restoration in video endoscopy</article-title>
          .
          <source>arXiv preprint arXiv:1904.07073</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Sharib</surname>
            <given-names>Ali</given-names>
          </string-name>
          , Felix Zhou, Barbara Braden, Adam Bailey, Suhui Yang, Guanju Cheng, Pengyi Zhang, Xiaoqiong Li,
          <string-name>
            <given-names>Maxime</given-names>
            <surname>Kayser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Roger D.</given-names>
            <surname>Soberanis-Mukul</surname>
          </string-name>
          , Shadi Albarqouni, Xiaokang Wang,
          <string-name>
            <surname>Chunqing</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Seiryo Watanabe, Ilkay Oksuz, Qingtian Ning, Shufan Yang, Mohammad Azam Khan, Xiaohong W. Gao, Stefano Realdon, Maxim Loshchenov, Julia A.
          <string-name>
            <surname>Schnabel</surname>
          </string-name>
          , James E. East, Geroges Wagnieres, Victor B.
          <string-name>
            <surname>Loschenov</surname>
            , Enrico Grisan, Christian Daul, Walter Blondel, and
            <given-names>Jens</given-names>
          </string-name>
          <string-name>
            <surname>Rittscher</surname>
          </string-name>
          .
          <article-title>An objective comparison of detection and segmentation algorithms for artefacts in clinical endoscopy</article-title>
          .
          <source>Scientific Reports</source>
          ,
          <volume>10</volume>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Tsung-Yi</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dolla´r.
          <article-title>Focal loss for dense object detection</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          , pages
          <fpage>2980</fpage>
          -
          <lpage>2988</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Diederik</surname>
            <given-names>P Kingma</given-names>
          </string-name>
          and
          <article-title>Jimmy Ba. method for stochastic optimization</article-title>
          .
          <source>arXiv:1412.6980</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Leslie</surname>
            <given-names>N</given-names>
          </string-name>
          <string-name>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Cyclical learning rates for training neural networks</article-title>
          .
          <source>In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV)</source>
          , pages
          <fpage>464</fpage>
          -
          <lpage>472</lpage>
          . IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hutter</surname>
          </string-name>
          .
          <article-title>Decoupled weight decay regularization</article-title>
          .
          <source>arXiv preprint arXiv:1711.05101</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>