<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DETECT ARTEFACTS OF VARIOUS SIZES ON THE RIGHT SCALE FOR EACH CLASS IN VIDEO ENDOSCOPY</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xiaokang Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chunqing Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biomedical Engineering, University of California</institution>
          ,
          <addr-line>Davis, Davis, CA 95616</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Ultrasound Imaging, Tiantan Hospital</institution>
          ,
          <addr-line>Beijing, 100050</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Detecting artefacts in video filmed in endoscopy is an important problem for downstream computer-assisted diagnosis. When tackling this problem, one challenge is that the size of an artefact varies in a wide range. The other challenge is that labeling endoscopic images is labor- extensive and is hard to outsource the labeling task to untrained people without the aid of doctors. In this report, we demonstrate how the performance of a Faster R-CNN model can be improved by scaling an image to the right scale before training and testing. The training method overcomes the issue that a convolution neural network trained on one scale barely works when detecting the same category of objects on a different scale. The method is totally independent of the model and can be easily adapted with other models. Besides, it saves time and memory by focusing on the patches that include objects when training the model. The source code? for this report will be made public upon the publishing of my solution.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Endoscopy is a widely used clinical procedure for the early
detection of numerous cancers (e.g., nasopharyngeal, gastric,
colorectal cancers, bladder cancer etc.), therapeutic
procedures and minimally invasive surgery. Video taken by the
camera of an endoscope is usually heavily corrupted with
multiple artefacts (e.g., pixel saturations, motion blur,
defocus, specular reflections, bubbles, fluid, debris etc.).
Accurate detection of artefacts is a core challenge in a wide range
of endoscopic applications addressing multiple different
disease areas. The importance of precise detection of these
artefacts is essential for high-quality endoscopic frame
restoration and crucial for realizing reliable computer assisted
endoscopy tools for improved patient care.</p>
      <p>In the last few years, convolution neural network (CNN)
has outperformed previous non-CNN based methods in
solving the object detection problem. The dominant CNN-based
methods fall into two categories, one-stage approach and
twostage approach, with the former method shining in speed and
the latter in accuracy. These two methods meet the demand
in different fields. For example, in the field of self-driving
car, speed is a prerequisite given an acceptable detection
performance as a self-driving car has to react instantly. In the
case of diagnosis in biomedical engineering, we can bear with
slightly more computation time for higher accuracy. Since
the advent of R-CNN [1], which is a two stage approach, this
region-based detection method has become increasingly
mature. Along the line, Fast R-CNN [2] introduced a RoI
pooling operation that does forward pass on all the object
proposals in an image simultaneously. Faster R-CNN [3]
further speeds up R-CNN by training a region proposal network
(PRN) using the feature maps generated by the convolution
operations at the low level, without introducing much cost.
Thus, I chose Faster R-CNN as the base framework in this
challenge.</p>
      <p>However, two challenges have to be solved when
developing the model. One special challenge is that the size of the
artefacts varies in a wide range and the other one is a limited
number of labeled images (2,192 in total). The scale-related
challenge is associated with the architecture of a CNN. The
low level feature maps of a CNN capture features like edges
and have a small receptive field, whereas the high-level
features capture more semantic features, and have a larger
receptive field [4, 5, 6]. Thus, the high-level features of small
objects (e.g. less than 32 pixels) get mixed with features for
background or objects nearby if the features do not
disappear due to dimension reduction caused in feature extraction.
e.g. For a feature stride of 32, the highest level features were
shrunk 32 times compared to the raw image. For very large
objects, the deeper layers suffer from extracting high-level
semantic features due to failing to integrate low-level features
given a limited feature stride.</p>
      <p>
        To alleviate the problem caused by the wide range of
object size, various solutions have been proposed. One category
of solution focused on designing new CNN architectures to
exploit the features at different levels. Under this paradigm,
SSD [
        <xref ref-type="bibr" rid="ref1">7</xref>
        ] and MSCNN [
        <xref ref-type="bibr" rid="ref2">8</xref>
        ], use feature maps from different
layers to detect objects at different scales. Although the
features for small objects survive in the low-layer features, they
lack semantic information which is supposed to be encoded
in high level features. FPN [
        <xref ref-type="bibr" rid="ref3">9</xref>
        ], DSSD [
        <xref ref-type="bibr" rid="ref4">10</xref>
        ], STDN [
        <xref ref-type="bibr" rid="ref5">11</xref>
        ]
integrate features at different layers. Another solution is to train
a neural network on a multi-scale image pyramid, resulting
in a scale-invariant predictor [
        <xref ref-type="bibr" rid="ref6">12</xref>
        ]. Nevertheless, the previous
solutions do not change the fact that high-level feature maps
for small objects are mixed and the receptive fields for large
objects are limited given an image and a CNN. Recently, a
new training method that detects all objects at a proper scale
by scaling up small objects and scaling down large objects has
been reported in the state-of-the-art models, SNIPER [
        <xref ref-type="bibr" rid="ref7">13</xref>
        ] and
TridentNet [
        <xref ref-type="bibr" rid="ref8">14</xref>
        ].
      </p>
      <p>In this study, we demonstrated the successful application
of the idea of detecting objects of various size at the right
scale in detecting artefacts in endoscopy. The report is
organized in such an order: datasets, methods, results, discussion
and conclusion.</p>
    </sec>
    <sec id="sec-2">
      <title>2. DATASETS</title>
      <p>
        The training dataset consists of 2,192 endoscopic images (Fig.
1 A), in which seven categories of artefacts were labeled [
        <xref ref-type="bibr" rid="ref10 ref9">15,
16</xref>
        ]. The seven categories are pixel saturation, motion blur,
specular reflections, bubbles, strong contrast, instrument and
other artefacts. The size of an artefact varies in a range from
a few pixels to one thousand pixels (Fig. 1 B). The number
of objects in each category is from 453 to 5835 (Fig. 1 C).
The performance of a model was tested on two datasets, one
collected by the same endoscope and the other collected by
a different endoscope to test the generalization ability of a
model. The former and latter testing datasets comprise 195
and 51 images, respectively.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. METHODS</title>
      <p>
        The model we built is a Faster R-CNN with a FPN as the
backbone. An FPN consists of mainly two parts, an encoder
and an decoder, which is very similar to a U-Net [
        <xref ref-type="bibr" rid="ref11">17</xref>
        ]
architecture developed for image segmentation tasks. Considering the
memory capacity of our GPU (GTX 1070 16GB), We chose
ResNet-50 as the workhorse of the encoder [
        <xref ref-type="bibr" rid="ref12">18</xref>
        ]. The
implementation was based on a modularized implementation of
mask R-CNN [
        <xref ref-type="bibr" rid="ref13">19</xref>
        ] in Pytorch [
        <xref ref-type="bibr" rid="ref14">20</xref>
        ]. The weights of the model
were initialized with the weights trained on the COCO dataset
except that the weights for the classification and regression
head were initialized with random weights.
      </p>
      <p>
        When training the model, we adapted the method
proposed in [
        <xref ref-type="bibr" rid="ref7">13</xref>
        ] considering the class imbalance in our dataset
and introduced data augmentation by strategically cutting a
patch from an image for training. In specific, given all the
bounding boxes (bboxes) in an image, k+1 bboxes were
sampled from all the bboxes in this image (Fig. 2 A). k is the
number of categories of objects in this image and 1 represents
a random bbox. Such operation is to alleviate the class
imbalance problem (Fig. 1 C) in our dataset. Then one bbox was
sampled from the k+1 bboxes. Finally a patch of the image
was cut out and scaled up or down depending on the size of
the object in that patch.
      </p>
      <p>When cutting a patch of the image (Fig. 2 B, step 1),
the size of the patch and location of the patch was jittered,
which allows us to generate not exactly the same patch every
time even though the patch with the same object is selected.
Note that the size of the patch is always larger than the object
and a larger patch was cut if a smaller object exists in that
patch. Otherwise, if the object were always in the center or
same location in the patch, the model would not learn to detect
objects but learn to localize objects assuming there is always
an object, which is not true. The setting for the size of a patch
(spatch) is defined by this equation: spatch = r sbbox, where
r = 4.5, 2, 1.5, and 1.2, respectively, for the cases, sbbox &lt; 80,
160, 350, and &gt; 350. If no object exists in a random patch, a
fixed size of patch was cut from an image.</p>
      <p>After cutting a patch from an image, the patch was scaled
up or down depending on the size of the object in that patch
(Fig. 2 B, step 2). The scaling provides a zoomed-in view for
small object objects and a zoomed-out view for large objects.
Thus, both high-level and low-level feature maps will exist
after an image passes a CNN. The scaling ratio (r) is inversely
proportional to the size (s) of the object in the patch: r = 160/s
and the size of all the objects are grouped into 6 bins. So
the settings used here are (r=4, s &lt; 40), (r=2, s¡80), (r=1,
s &lt; 160), (r=0.5, s &lt; 250), (r=0.25, s &lt; 640), and (r=0.13,
s &gt; 640). The reason for choosing such settings is that there
will be 5 pixels in the last layer of the encoder if a raw input
of 160 pixels is fed into the model, which has a feature stride
of 32. A patch, which has no objects in it, was scaled with a
random ratio on the fly.</p>
      <p>After scaling up or down the patch cut from each image in
a batch, objects that are too large/small were excluded. The
thresholds for too small and too large objects are 32 and 2000,
respectively. Choosing 32 as the threshold is because of the
feature stride of the model is 32 and choosing 2000 is just
because it is large enough.</p>
      <p>In each training iteration, multiple patches from multiple
images were cut and normalized as a batch by padding the
patch with with the channel mean and concatenated, resulting
in a batch of images, whose width and height are a multiple
of 32 (Fig. 2 B, step 3). For a patch that is already a multiple
of the stride size of the encoder, no padding was added. Since
the padding is always on the bottom and right side if
necessary, the coordinates of the bounding box does not change.
Collating multiple samples and unifying them in size can be
easily implemented in Pytorch1. For other details like how
the bounding boxes were adjusted accordingly when cutting
a patch from an image, see our code on Github.</p>
    </sec>
    <sec id="sec-4">
      <title>4. RESULTS</title>
      <p>In inference, we tested one image on all the scales used in
training (scales: 4, 2, 1, 0.5, 0.25 and 0.13). The
coordi1https://pytorch.org
bubbles
Instrument
artifact
specularity
specularity
saturation
artifact
blur
contrast
bubbles
instrument</p>
      <p>C6000
5000
nates of all the detected objects were transformed back to the
original scale. To remove redundant bounding boxes,
nonmaximum suppression were conducted on all the detected
bounding boxes for each category of object. It was observed
that many false positive, which were small objects, were
detected on the scale of 4, so I finally chose not to include the
prediction on that scale.</p>
      <p>The performance of our model was evaluated by a hybrid
metric which was a weighted score of mean average
precision (mAP) and Intersection over Union (IoU): 0.6 mAP
+ 0.4 IoU. We compared the performance of two ways for
training the same model. One way is training the model on
the whole image every time and the other is on patches
generated following the method described above. The threshold
for the probability when determining an object is 0.65. In the
former case, the best overall score on the two testing sets was
0.221. For the latter case, we trained the network for 27,105
iterations (batch size is 8 in each iteration) and a significant
boost in performance was observed. The score we reached
was 0.293, which was among the top 10 teams on the
leaderboard.</p>
    </sec>
    <sec id="sec-5">
      <title>5. DISCUSSION &amp; CONCLUSION</title>
      <p>The training method boosted the performance of Faster
RCNN in two ways. First, as we discussed in the
Introduction section, it alleviates the scale variation problem by
scaling up/down an object to the right scale to detect. Second,
randomly cutting a patch which includes an object allows us
to generate far more different training images compared to
feeding the whole image to the model. Thus, cutting a patch
serves as a data augmentation technique. Besides, it offers
flexibility to deal with the class imbalance as we can choose
which patch to cut from a training image, considering the
distribution of the counts of all the classes.</p>
      <p>
        Since the detected bboxes on all the scales were merged, a
bbox was called if it was detected on any of the scales. Such
an integration approach tends to report more false positive.
One failure case we observed is that a false positive object
does look like a true object because the model decides
without considering the context of that object. We run into the
case when an image is scaled up by 4 times. Thus the
context information does matter and a context refinement
probably corrects such kind of errors [
        <xref ref-type="bibr" rid="ref15">21</xref>
        ]. An alternative solution
to solve this bias of this method can be feeding the detected
bboxes as the input for the RoI pooling layer and merging the
features generated by the model on an image pyramid. Since
Sample a bbox from
each class and add a
random bbox
      </p>
      <p>Balanced bounding boxes</p>
      <p>True bounding boxes
the scaled down images include more context information, we
expect the problem to be solved in this way.</p>
      <p>
        To further boost the detection performance, the regular
convolution operation in the FPN backbone can be replaced
by deformable convolution operation will enhance the
transformation modeling capacity of CNNs [
        <xref ref-type="bibr" rid="ref16">22</xref>
        ] or a newly
proposed backbone designed for object detection task [?]. In
conclusion, there is still room for improvement and we have
demonstrated the performance of a Faster R-CNN model can
be improved significantly by training and detecting the
objects on the right scale.
      </p>
      <p>6. REFERENCES
[1] Ross Girshick, Jeff Donahue, Trevor Darrell, and
Jitendra Malik. Rich feature hierarchies for accurate object
detection and semantic segmentation. In Proceedings
of the IEEE conference on computer vision and pattern
recognition, pages 580–587, 2014.
[2] Ross Girshick. Fast r-cnn. In Proceedings of the IEEE
international conference on computer vision, pages
1440–1448, 2015.
[3] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian
Sun. Faster r-cnn: Towards real-time object detection
with region proposal networks. In Advances in neural
information processing systems, pages 91–99, 2015.
[4] Matthew D Zeiler and Rob Fergus. Visualizing and
understanding convolutional networks. In European
conference on computer vision, pages 818–833. Springer,
2014.
[5] Wenjie Luo, Yujia Li, Raquel Urtasun, and Richard
Zemel. Understanding the effective receptive field in
deep convolutional neural networks. In Advances in
neural information processing systems, pages 4898–
4906, 2016.
[6] Bharat Singh and Larry S Davis. An analysis of scale
invariance in object detection snip. In Proceedings of the
IEEE conference on computer vision and pattern
recognition, pages 3578–3587, 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Wei</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Dragomir Anguelov, Dumitru Erhan,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          , Scott Reed, Cheng-Yang
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , and
          <string-name>
            <surname>Alexander C Berg. Ssd</surname>
          </string-name>
          :
          <article-title>Single shot multibox detector</article-title>
          .
          <source>In European conference on computer vision</source>
          , pages
          <fpage>21</fpage>
          -
          <lpage>37</lpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Zhaowei</given-names>
            <surname>Cai</surname>
          </string-name>
          , Quanfu Fan, Rogerio S Feris, and
          <string-name>
            <given-names>Nuno</given-names>
            <surname>Vasconcelos</surname>
          </string-name>
          .
          <article-title>A unified multi-scale deep convolutional neural network for fast object detection</article-title>
          .
          <source>In European conference on computer vision</source>
          , pages
          <fpage>354</fpage>
          -
          <lpage>370</lpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Tsung-Yi</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Piotr Dolla´r, Ross Girshick, Kaiming He,
          <string-name>
            <surname>Bharath Hariharan</surname>
            , and
            <given-names>Serge</given-names>
          </string-name>
          <string-name>
            <surname>Belongie</surname>
          </string-name>
          .
          <article-title>Feature pyramid networks for object detection</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>2117</fpage>
          -
          <lpage>2125</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Cheng-Yang</surname>
            <given-names>Fu</given-names>
          </string-name>
          , Wei Liu, Ananth Ranga, Ambrish Tyagi, and
          <string-name>
            <surname>Alexander C Berg. Dssd</surname>
          </string-name>
          :
          <article-title>Deconvolutional single shot detector</article-title>
          .
          <source>arXiv preprint arXiv:1701.06659</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Peng</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Bingbing Ni, Cong Geng, Jianguo Hu, and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <article-title>Scale-transferrable object detection</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>528</fpage>
          -
          <lpage>537</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Edward</surname>
            <given-names>H Adelson</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Charles H Anderson</surname>
          </string-name>
          , James R Bergen,
          <string-name>
            <surname>Peter J Burt</surname>
          </string-name>
          , and
          <string-name>
            <surname>Joan M Ogden.</surname>
          </string-name>
          <article-title>Pyramid methods in image processing</article-title>
          .
          <source>RCA engineer</source>
          ,
          <volume>29</volume>
          (
          <issue>6</issue>
          ):
          <fpage>33</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1984</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Bharat</surname>
            <given-names>Singh</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mahyar</given-names>
            <surname>Najibi</surname>
          </string-name>
          , and
          <string-name>
            <surname>Larry S Davis. Sniper:</surname>
          </string-name>
          <article-title>Efficient multi-scale training</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>9310</fpage>
          -
          <lpage>9320</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Yanghao</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Yuntao</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Naiyan</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Zhaoxiang</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <article-title>Scale-aware trident networks for object detection</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .
          <year>01892</year>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Sharib</surname>
            <given-names>Ali</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Felix Zhou</surname>
            , Christian Daul, Barbara Braden, Adam Bailey, Stefano Realdon, James East, Georges Wagnires, Victor Loschenov, Enrico Grisan, Walter Blondel, and
            <given-names>Jens</given-names>
          </string-name>
          <string-name>
            <surname>Rittscher</surname>
          </string-name>
          .
          <article-title>Endoscopy artifact detection (EAD 2019) challenge dataset</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1905</year>
          .03209,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Sharib</surname>
            <given-names>Ali</given-names>
          </string-name>
          , Felix Zhou, Adam Bailey, Barbara Braden, James East, Xin Lu, and
          <string-name>
            <given-names>Jens</given-names>
            <surname>Rittscher</surname>
          </string-name>
          .
          <article-title>A deep learning framework for quality assessment and restoration in video endoscopy</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1904</year>
          .07073,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Olaf</surname>
            <given-names>Ronneberger</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>
          . In International Conference on
          <article-title>Medical image computing and computer-assisted intervention</article-title>
          , pages
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          . Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Kaiming</surname>
            <given-names>He</given-names>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , pages
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Kaiming</surname>
            <given-names>He</given-names>
          </string-name>
          , Georgia Gkioxari, Piotr Dolla´r, and
          <string-name>
            <given-names>Ross</given-names>
            <surname>Girshick. Mask</surname>
          </string-name>
          r-cnn.
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          , pages
          <fpage>2961</fpage>
          -
          <lpage>2969</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [20]
          <article-title>Francisco Massa and Ross Girshick. maskrcnnbenchmark: Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch</article-title>
          . https://github.com/facebookresearch/maskrcnnbenchmark,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Roozbeh</surname>
            <given-names>Mottaghi</given-names>
          </string-name>
          , Xianjie Chen, Xiaobai Liu, NamGyu Cho,
          <string-name>
            <surname>Seong-Whan</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Sanja Fidler, Raquel Urtasun, and
          <string-name>
            <given-names>Alan</given-names>
            <surname>Yuille</surname>
          </string-name>
          .
          <article-title>The role of context for object detection and semantic segmentation in the wild</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>891</fpage>
          -
          <lpage>898</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Jifeng</surname>
            <given-names>Dai</given-names>
          </string-name>
          , Haozhi Qi, Yuwen Xiong,
          <string-name>
            <given-names>Yi</given-names>
            <surname>Li</surname>
          </string-name>
          , Guodong Zhang, Han Hu, and
          <string-name>
            <given-names>Yichen</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <article-title>Deformable convolutional networks</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          , pages
          <fpage>764</fpage>
          -
          <lpage>773</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>