<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A U-NET++ WITH PRE-TRAINED EFFICIENTNET BACKBONE FOR SEGMENTATION OF DISEASES AND ARTIFACTS IN ENDOSCOPY IMAGES AND VIDEOS Le Duy Huy`nh, Nicolas Boutry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>EPITA Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Development Laboratory (LRDE)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rue Voltaire</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Le Kremlin-Biceˆtre</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Endoscopy is a widely used clinical procedure for the early detection of numerous diseases. However, the images produced are usually heavily corrupted with multiple artifacts that reduce the visualization of the underlying tissue. Moreover, the localization of actual diseased regions is also a complex problem. For that reason, EndoCV2020 challenges aim to make progress in the state-of-the-art in the detection and segmentation of artifacts and diseases in endoscopy images. In this work, we propose approaches based on U-Net and UNet++ architecture to automate the segmentation task of EndoCV2020. We use the EfficientNet as our encoder to extract powerful features for our decoders. Data augmentation and pre-trained weights are employed to prevent overfilling and improve generalization. Test-time augmentation also helps in improving the results of our models. Our methods performs well in this challenge and achieves a score of 60.20% for the EAD2020 semantic segmentation task and 59.81% for the EDD2020's.</p>
      </abstract>
      <kwd-group>
        <kwd>Endoscopy</kwd>
        <kwd>U-Net++</kwd>
        <kwd>EfficientNet</kwd>
        <kwd>Test-time augmentation</kwd>
        <kwd>Segmentation</kwd>
        <kwd>Detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Endoscopy is a widely used clinical procedure for the early
detection of cancers, therapeutic procedures, and minimally
invasive surgery in hollow-organs. Computer-assisted
methods would improve diagnosis and assist in surgical planning.
These endoscopic applications face two main challenges: 1)
Endoscopy video frames are usually heavily corrupted with
multiple artifacts that reduce the visualization of the
underlying tissue and affect post-analysis. Accurate detection of
these artifacts allows corrections of frames and is, therefore,
a core challenge in a wide range of endoscopic applications
[1]. 2) Even with uncorrupted video frames, temporally
consistently localizing and segmenting of disease ROIs is a
challenge due to non-planar geometries, variation in imaging
modalities, and deformations of organs. For these reasons,
after the success of EAD2019 [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ], the endoscopy computer
      </p>
      <p>
        Copyright c 2020 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
(a)
(e)
(b)
(f)
(c)
(g)
(d)
(h)
vision challenges on segmentation and detection 2020
(EndoCV 2020) are held to make progress in the state-of-the-art
further. It consists of two sub-challenges: Endoscopy
Artefact Detection and Segmentation (EAD2020), and Endoscopy
Disease Detection and Segmentation (EDD2020) [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ]. Each
challenge is further divided into the detection and the
semantic segmentation tasks.
      </p>
      <p>
        In this paper, we introduce our works on these challenges.
First, Sec. 2 presents our observations of the datasets. Then,
we describe our methods in Sec. 3. We participate in all tasks
of both challenges but focus mainly on the two semantic
segmentation ones using fully convolutional networks such as
U-Nets [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ] and U-Net++ [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] with EfficientNet [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ] backbones
(Sec. 3.1). We also did some experiments with the detection
tasks using our segmentation results or the YOLOv3 model
[7] (Sec. 3.2). Finally, we will present our results on the
testing data in Sec. 4 and conclude in Sec. 5.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. DATASETS</title>
    </sec>
    <sec id="sec-3">
      <title>2.1. About the dataset</title>
      <p>
        The EndoCV2020 challenge consists of two datasets. The
EAD2020 dataset is an extended version of last year’s
EAD2019 challenge [
        <xref ref-type="bibr" rid="ref6">8</xref>
        ] with annotations corrected, and a
new class added. The EAD2020 train data were released
in four subsets. They consist of a total of 3005 images,
among which 2531 are annotated with bounding boxes for
eight classes (specularity, saturation, artifact, blur, contrast,
bubbles, instrument, and blood), and 643 come with
segmentation ground truth for five classes (instrument, specularity,
artifact, bubbles, saturation, and blood). The EDD2020
train data contains 386 images with boxes and segmentation
ground-truth for five disease classes (normal dysplastic
Barrett’s oesophagus (BE), suspicious area, high-grade dysplasia
(HGD), adenocarcinoma (cancer), and polyp).
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.2. Data correction</title>
      <p>The ground-truth for these challenges contain some issues.
They can be classified into annotator disagreement and
systematic error.</p>
      <p>Some annotator disagreement was spotted in the EAD2020
detection dataset. For example, an artifact region is identified
in one frame but is not in the next frame, or one connected
region is marked with two bounding boxes in one frame but
only one in the other. There is also misclassified segmentation
mask. We consider this type of error as noises in the dataset
since we do not have the resource to make adjustments.</p>
      <p>We corrected all the systematic errors that we noticed.
There are some small anomalies in bounding box ground
truths, likely due to rounding when converting from
PascalVOC to Yolo format. There is also a line at the bottom
of some EAD2020 segmentation ground-truths that do not
correspond to any artifact.</p>
    </sec>
    <sec id="sec-5">
      <title>3. OUR APPROACHES</title>
      <p>We focus mainly on the two segmentation tasks with models
in the U-Net family. In the case of disease detection, we will
take advantage of our segmentation results. For the artifacts
detection tasks, we train a separate YOLOv3 [7] network due
to the difference in the number of classes (eight classes for
detect and five classes for segmentation), and due to
disagreement between EAD2020 segmentation and detection
groundtruths (e.g, Fig 1(a) and Fig 1(b)).</p>
    </sec>
    <sec id="sec-6">
      <title>3.1. Segmentation of Diseases and Artifacts</title>
      <p>3.1.1. Models
The state-of-the-art of semantic segmentation are
methods based on encoder-decoder architecture such as U-Net,
U-Net++. These encoder-decoder architectures use
skipconnections to combine low-resolution, semantically-rich
at deeper feature maps with shallower, fine-grained ones to
recover fine detail of region of interest.</p>
      <p>
        Instead of using the original U-Net encoder, we use
EfficientNet [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ], which claims to be balanced between network
depth, width, and resolution. This architecture achieved
better accuracy on ImageNet [
        <xref ref-type="bibr" rid="ref7">9</xref>
        ] with fewer parameters and
requires fewer FLOPS than other networks such as ResNet [
        <xref ref-type="bibr" rid="ref8">10</xref>
        ]
or DenseNet [
        <xref ref-type="bibr" rid="ref9">11</xref>
        ]. EfficientNet is available with different
versions, starts from B0 at 5.3 million parameters to B7 at 66
million. We extract five feature maps at different scales from
EfficientNet as the input of our decoders. An illustration of
the EfficientNet B1 encoder and where intermediate feature
maps are extracted is in Fig. 2(a).
      </p>
      <p>We start our experimentations with the U-Net architecture
and the EfficientNet B5. Subsequent tests show that a deeper
encoder is not needed, so we transit to a larger decoder
(UNet++) with smaller encoder (EfficientNet B2 and latter B1).
An illustration of our networks could be found in Fig. 2</p>
      <p>We speed up the training process with pre-trained weights.
Although it was trained on ImageNet, which is a database of
natural images, the pre-trained weights do improve the
training and local validation scores.</p>
      <sec id="sec-6-1">
        <title>3.1.2. Data augmentation</title>
        <p>
          Input images are resized (while keeping their aspect ratio) and
zeros-padded to fit into 512x512 pixels. We randomly apply
these augmentation techniques with 50% probability:
Rotation with random angle, RGB value shift, horizontal or
vertical flipping, random scaling, elastic deformation [
          <xref ref-type="bibr" rid="ref10">12</xref>
          ],
cropping.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>3.1.3. The loss function</title>
        <p>
          The semantic segmentation tasks are evaluated with four
metrics: F1 (i.e., Dice score), F2, precision, and recall. The
semantic segmentation score is the average of these four metrics
[
          <xref ref-type="bibr" rid="ref2">3</xref>
          ]. Let us remind that F = (1 + 2) ( 2pprerceicsiisoionnr)e+carellcall
weights recall times as important as precision. Therefore
the final score places more emphasis on recall. As a result,
we will train our network using F2-loss, which is defined
similarly to Dice-loss: F2-loss = 1-F2. We also apply L2
regularization with a factor of 0.0001.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>3.1.4. Training</title>
        <p>We train on 80% of the train set and use the other 20% for
validation. We start by fixing the pre-trained encoder and train
the decoder part for 80 epochs using Adam optimizer with a
learning rate (LR) of 10 3. From the 41st epochs, we train
the whole network, starting at LR=10 3 and decrease with a
factor of 0.5 if the validation score does not decrease after 40
epochs. We trained for a total of 1000 epochs. The training is
early stopped if the training score could not be increased after
88 epochs. We select the weights that maximized the score on
the validation set for evaluation of the test set.</p>
      </sec>
      <sec id="sec-6-4">
        <title>3.1.5. Prediction and Post-processing</title>
        <p>Test images are resized and zero-padded to 512x512. We keep
the aspect ratio and do not enlarge small images. Since our
prediction sometimes contains small holes which rarely
appear in the EAD2020 train set, a small hole-filling operation
is applied at the end of the pipeline. This hole-filling process
did not make a significant improvement in the final score.
(a) EfficientNet B1 encoder and extracted intermediate feature maps (Xi)
(b) Building blocks
(c) U-Net
(d) U-Net++</p>
      </sec>
      <sec id="sec-6-5">
        <title>3.1.6. Test-time Augmentation</title>
        <p>
          The segmentation results could be further improved with
Testtime augmentation (TTA). This approach has been
demonstrated in the literature (e.g., in semantic segmentation of
brain tumor [
          <xref ref-type="bibr" rid="ref11">13</xref>
          ]). In short, we will make predictions on
the test image and several of its transformed versions and
then combine these results. We use five transformations:
horizontal, vertical flipping, rotations of 90 , 180 , and 270 .
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>3.2. Detection of Diseases and Artifacts</title>
      <sec id="sec-7-1">
        <title>3.2.1. Detection of Diseases (EDD2020)</title>
        <p>For this task, we will take advantage of our segmentation
results. The bounding boxes of connected components (CCs)
larger than 0.5% of the image area are presented as our
detection results. This approach is not optimal because it cannot
handle slit-and-merge of detection bounding boxes, that is the
cases where a single CC is marked with more than one box,
or a single box marks several CCs.</p>
      </sec>
      <sec id="sec-7-2">
        <title>3.2.2. Detection of Artifacts (EAD2020)</title>
        <p>
          For reasons stated earlier, we have to train a separate detector
for this task. We choose YOLOv3 because this model’s
effectiveness has been shown in this type of data [1]. It is also
relatively faster than two-states detectors such as RetinaNet
[
          <xref ref-type="bibr" rid="ref12">14</xref>
          ]. We have tried to address the class imbalance issue with
focal loss [
          <xref ref-type="bibr" rid="ref12">14</xref>
          ]. However, it did not improve our local
validation mAP, similar to the remark in [7].
        </p>
        <p>We train our standard YOLOv3 on 416x416 inputs. We
start by training the last three detection layers for 20 epochs
at LR=10 2, then the upscaling part of the network for 40
epochs at LR=10 3. Finally, we train the whole network at
LR =10 5 and reduce the LR by a factor of 0.5 if validation
loss does not decrease after five epochs.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>4. EVALUATION</title>
      <p>As presented in Sec. 3.1.3, the semantic segmentation are
evaluated by the mean of F1, F2, precision, and recall. The
detection tasks are evaluated by a combination of mean
average precision (mAP), and intersection over union (IoU).
However, how these metrics are weighted was not disclosed until
after the challenge ends.</p>
      <p>The evaluation was done online at
https://endocv.grandchallenge.org/. It is divide into two phases. The first
test-phase only evaluate 50% of the test data. The size of
EAD2020 and EDD2020 test data is summarized in Tab. 1.
The EAD2020 detection task includes an out-of-sample set,
which contains images provided exclusively from the training
or other test datasets.</p>
      <p>The quantitative results of our models are presented in
Tab. 2 and Tab. 3. Our segmentation models performed well
with double the number of decoder filters, i.e., Dec(x; 2)
becomes Dec(2x; 2), (-): information is not available because
that method was not submitted to that test phase.</p>
      <p>Full Test Set
Recall Score Final Score
and consistently on both test subset. In Tab. 2, we can observe
that the hole-filling operator adds a small boost while TTA
improves the final score significantly. Our best approach is
U-Net++ with B1 encoder, ran with TTA and post-processed
with holes-filling. We could also see that our detection model
needs improvements.</p>
    </sec>
    <sec id="sec-9">
      <title>5. CONCLUSION AND PERSPECTIVE</title>
      <p>In this work, we demonstrate the effectiveness of U-Net and
U-Net++ with pre-trained EfficientNet backbone for the
segmentation of disease and artifact in endoscopy images. Our
experiments also show that TTA can provide better
segmentation results compared to only predicting on original images.</p>
      <p>For further works, we intend to improve the
generalization of our segmentation and detection models by applying
more data augmentation techniques and using synthetic data.
We are also considering to improve further the decoder of our
models with an attention mechanism.</p>
    </sec>
    <sec id="sec-10">
      <title>6. REFERENCES</title>
      <p>[1] Sharib Ali, Felix Zhou, Adam Bailey, Barbara Braden,
James East, Xin Lu, and Jens Rittscher. A deep
learning framework for quality assessment and restoration
in video endoscopy. arXiv preprint arXiv:1904.07073,
2019.
[7] Joseph Redmon and Ali Farhadi. YOLOv3: An
incremental improvement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Sharib</given-names>
            <surname>Ali</surname>
          </string-name>
          , Felix Zhou, Barbara Braden, Adam Bailey, Suhui Yang, Guanju Cheng, Pengyi Zhang, Xiaoqiong Li,
          <string-name>
            <given-names>Maxime</given-names>
            <surname>Kayser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Roger D.</given-names>
            <surname>Soberanis-Mukul</surname>
          </string-name>
          , Shadi Albarqouni, Xiaokang Wang,
          <string-name>
            <surname>Chunqing</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Seiryo Watanabe, Ilkay Oksuz, Qingtian Ning, Shufan Yang, Mohammad Azam Khan, Xiaohong W. Gao, Stefano Realdon, Maxim Loshchenov, Julia A.
          <string-name>
            <surname>Schnabel</surname>
          </string-name>
          , James E. East, Geroges Wagnieres, Victor B.
          <string-name>
            <surname>Loschenov</surname>
            , Enrico Grisan, Christian Daul, Walter Blondel, and
            <given-names>Jens</given-names>
          </string-name>
          <string-name>
            <surname>Rittscher</surname>
          </string-name>
          .
          <article-title>An objective comparison of detection and segmentation algorithms for artefacts in clinical endoscopy</article-title>
          .
          <source>Scientific Reports</source>
          ,
          <volume>10</volume>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Sharib</given-names>
            <surname>Ali</surname>
          </string-name>
          , Noha Ghatwary, Barbara Braden, Lamarque Dominique, Adam Bailey, Stefano Realdon, Cannizzaro Renato, Jens Rittscher,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Daul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>James</given-names>
            <surname>East</surname>
          </string-name>
          .
          <article-title>Endoscopy disease detection challenge 2020</article-title>
          . CoRR, abs/
          <year>2003</year>
          .03376,
          <year>February 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Olaf</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>
          .
          <source>In Medical Image Computing and Computer-Assisted Intervention - MICCAI</source>
          <year>2015</year>
          , pages
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          . Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Zongwei</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh, and
          <string-name>
            <given-names>Jianming</given-names>
            <surname>Liang</surname>
          </string-name>
          . Unet++
          <article-title>: A nested unet architecture for medical image segmentation. In Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support</article-title>
          , pages
          <fpage>3</fpage>
          -
          <lpage>11</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mingxing</given-names>
            <surname>Tan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Quoc V.</given-names>
            <surname>Le</surname>
          </string-name>
          . Efficientnet:
          <article-title>Rethinking model scaling for convolutional neural networks</article-title>
          .
          <source>In ICML</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Sharib</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>Felix Zhou</surname>
            , Christian Daul, Barbara Braden, Adam Bailey, Stefano Realdon, James East, Georges Wagnieres, Victor Loschenov,
            <given-names>Enrico</given-names>
          </string-name>
          <string-name>
            <surname>Grisan</surname>
          </string-name>
          , et al.
          <article-title>Endoscopy artifact detection (ead 2019) challenge dataset</article-title>
          .
          <source>arXiv preprint arXiv:1905.03209</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Olga</given-names>
            <surname>Russakovsky</surname>
          </string-name>
          , Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein,
          <string-name>
            <surname>Alexander C. Berg</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <article-title>ImageNet Large Scale Visual Recognition Challenge</article-title>
          .
          <source>International Journal of Computer Vision (IJCV)</source>
          ,
          <volume>115</volume>
          (
          <issue>3</issue>
          ):
          <fpage>211</fpage>
          -
          <lpage>252</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Kaiming</surname>
            <given-names>He</given-names>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In 2016 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2016</year>
          ,
          <string-name>
            <surname>Las</surname>
            <given-names>Vegas</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          , USA, June 27-30,
          <year>2016</year>
          , pages
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . IEEE Computer Society,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Van Der</given-names>
            <surname>Maaten</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          .
          <article-title>Densely connected convolutional networks</article-title>
          .
          <source>In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , pages
          <fpage>2261</fpage>
          -
          <lpage>2269</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.Y.</given-names>
            <surname>Simard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Steinkraus</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.C.</given-names>
            <surname>Platt</surname>
          </string-name>
          .
          <article-title>Best practices for convolutional neural networks applied to visual document analysis</article-title>
          .
          <source>In ICDAR</source>
          ,
          <year>2003</year>
          ., volume
          <volume>1</volume>
          , pages
          <fpage>958</fpage>
          -
          <lpage>963</lpage>
          .
          <source>IEEE Comput. Soc.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Guotai</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Wenqi</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Se´bastien Ourselin, and</article-title>
          <string-name>
            <given-names>Tom</given-names>
            <surname>Vercauteren</surname>
          </string-name>
          .
          <article-title>Automatic brain tumor segmentation using convolutional neural networks with test-time augmentation</article-title>
          . In Brainlesion: Glioma, Multiple Sclerosis,
          <source>Stroke and Traumatic Brain Injuries</source>
          , pages
          <fpage>61</fpage>
          -
          <lpage>72</lpage>
          , Cham,
          <year>2019</year>
          . Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollr</surname>
          </string-name>
          .
          <article-title>Focal loss for dense object detection</article-title>
          .
          <source>In 2017 IEEE International Conference on Computer Vision</source>
          (ICCV), pages
          <fpage>2999</fpage>
          -
          <lpage>3007</lpage>
          ,
          <year>Oct 2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>