<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparative Study of Deep Learning Models for Automatic Coronary Stenosis Detection in X-ray Angiography</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>nilov</string-name>
          <email>viacheslav.v.danilov@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kirill Klyshnikov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ny Ov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ro Fr</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research Institute for Complex Issues of Cardiovascular Diseases</institution>
          ,
          <addr-line>Kemerovo, 650002</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tomsk Polytechnic University</institution>
          ,
          <addr-line>Tomsk, 634050</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Leeds</institution>
          ,
          <addr-line>Leeds, LS2 9JT</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>V.Danilov</institution>
          ,
          <addr-line>A.Frangi</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article explores the application of machine learning approach to detect both single-vessel and multivessel coronary artery disease from X-ray angiography. Since the interpretation of coronary angiography images requires interventional cardiologists to have considerable training, our study is aimed at analysing, training, and assessing the potential of the existing object detectors for classifying and detecting coronary artery stenosis using angiographic imaging series. 100 patients who underwent coronary angiography at the Research Institute for Complex Issues of Cardiovascular Diseases were retrospectively enrolled in the study. To automate the medical data analysis, we examined and compared three models (SSD MobileNet V1, Faster-RCNN ResNet-50 V1, FasterRCNN NASNet) with various architecture, network complexity, and a number of weights. To compare developed deep learning models, we used the mean Average Precision (mAP) metric, training time, and inference time. Testing results show that the training/inference time is directly proportional to the model complexity. Thus, Faster-RCNN NASNet demonstrates the slowest inference time. Its mean inference time per one image made up 880 ms. In terms of accuracy, FasterRCNN ResNet-50 V1 demonstrates the highest prediction accuracy. This model has reached the mAP metric of 0.92 on the validation dataset. SSD MobileNet V1 has demonstrated the best inference time with the inference rate of 23 frames per second.</p>
      </abstract>
      <kwd-group>
        <kwd>Stenosis Detection</kwd>
        <kwd>X-ray Angiography</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Transfer Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Coronary artery disease (CAD) is the leading cause of mortality worldwide [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. CAD
is commonly caused by atherosclerotic plaques encroaching the coronary artery
luCopyright © 2020 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).
men and resulting in its narrowing or complete blockage. To date, invasive coronary
angiography is the gold standard for diagnosing coronary artery stenosis using X-ray
visualisation of a radiopaque agent. Therefore, analysis and interpretation of coronary
angiography data play an important role in the accurate diagnosis of coronary artery
stenosis. The severity of stenosis and the SYNTAX score are used for selecting either
minimally invasive extravascular surgery or invasive intervention.
      </p>
      <p>
        Despite recent advances in diagnostic tools and algorithms, capable of detecting the
location of coronary artery stenosis (82 - 95%) and classifying it (80 - 97%) [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6 ref7 ref8">2–8</xref>
        ] there
are significant limitations, necessitating further studies. Major drawbacks include poor
scalability and flexibility of preprocessing algorithms that require fine-tuning. Most
detection algorithms use the cascading principle, prone to the accumulation of errors.
Therefore, our study is aimed at developing, training, and assessing several neural
networks to determine coronary artery stenosis with the highest predicting accuracy on
original angiographic imaging series.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Source data</title>
      <p>One hundred patients who underwent coronary artery angiography using angiography
systems Coroscop (Siemens) and Innova (GE Healthcare) at the Research Institute for
Complex Issues of Cardiovascular Diseases (Kemerovo, Russia) were retrospectively
enrolled in the study. Patients with multivessel CAD were excluded from the study.
Angiographic imaging series of the radiopaque overlaid coronary arteries with stenotic
segments were selected by an interventional cardiologist. Thus, 8325 input images in
grayscale (one channel) of 512 512 pixels to 1000 1000 pixels were ultimately
included for further study. Of them, 7492 (90%) images were used for training, and 833
(10%) images were used for validation. Data were labelled using a free, open-source
version of SaaS (Software as a Service) solution { LabelBox. Typical data labelling of
the source images is shown in Fig. 1.</p>
      <p>To analyse the source dataset, we estimated the size of the stenotic region computing
the area of the bounding box. Similarly to the Common Objects in Context (COCO)
dataset, we divided objects by their area into three types: small (area &lt; 322), medium
(322 area 962), and large (area &gt; 962) objects. A total of 2509 small objects
(30%), 5704 medium objects (69%), and 111 large objects (1%) were obtained in the
input data. Considering the unbalanced distribution of classes in much training data,
we suppose that the models may perform poorer on larger objects than on small and
medium ones.</p>
      <p>To determine the stenosis location accurately, we evaluated the distribution of the
stenosis coordinates along the vessel in the input images. The coordinates of the centre
point of the bounding box around the stenotic lesion were normalised and assessed.
Based on this assessment, a distribution map of the coordinates of the stenosis centres
was generated and is shown in Fig. 2. The distribution of the coordinates highlights
two centres with relative coordinates (0.50; 0.20) and (0.27; 0.27) along the stenotic
vessel segment. The coordinates of the centres are evenly distributed without explicit
statistical outliers.</p>
      <p>Automatic coronary stenosis detection in X-ray angiography... 3
(a) Patient 1
(b) Patient 2</p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <sec id="sec-3-1">
        <title>Models description</title>
        <p>
          We applied machine learning algorithms to detect coronary artery stenosis on the CAG
imaging series. Machine learning has shown beneficial potential in computer vision
and image processing. We used SSD [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and Faster-RCNN [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] object detectors from
the Tensorflow Detection Model Zoo [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] based on such models as MobileNet [
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ],
ResNet [
          <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
          ] and NASNet [
          <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
          ]. Three models with various architectures, network
complexity, and a number of weights were selected. The lightweight SSD MobileNet
V1 SSD detector enabling real real-time data processing was chosen as the reference
model. Faster-RCNN NASNet, with over 80 million weights, was the most complex
model selected for the study. A brief description of the models is presented in Table 1.
Characteristics of neural networks, including mAP, are reported based on their training
on the COCO dataset.
When training neural network models, their base configuration is similar to that used to
train on the COCO dataset. For the unambiguous comparison of the selected models,
the total number of training steps was set to 100 equal to 100000 iterations of learning.
Regarding the loss functions, Weighted Smooth L1 loss (see equation 3 in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]) was
the localisation loss, and Weighted Focal Loss [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] was the classification loss. It should
be noted that the SSD-based model was trained using the Cosine decay with the
warmup. This technique allowed gradually decreasing the learning rate (LR) depending on
the learning step. To train the networks, we used P2 (Nvidia Tesla K80 12 Gb, 1.87
TFLOPS) and P3 instances (Nvidia Tesla V100 16 Gb, 7.8 TFLOPS) from Amazon
Web Services. Table 2 summarises the main characteristics of the model training.
        </p>
        <p>Serial changes in precision were tested on the validation set during the training
process. The mAP metric, as the metric of interest, with a predefined threshold value
for Intersection over Union equal to 0.5 (mAP@0.5) was used. Fig. 3 shows smooth
changes in the mAP on the validation set during the training process. As seen, all models
converge to a specific value of the mAP asymptotic accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Comparative analysis</title>
        <p>Automatic coronary stenosis detection in X-ray angiography... 5
Faster-RCNN
Faster-RCNN</p>
        <p>NASNet
with it. Fig. 4 and 5 report the basic metrics of the model performance (mAP, training
time, and inference time). The graphs are sorted in a way to present the models with
superior metrics first.</p>
        <p>The inference time was estimated using P3 instance (Nvidia Tesla V100 16 Gb, 7.8
TFLOPS) of Amazon Web Services. Based on the obtained results, the inference time
directly depends on the complexity of the model and the total number of its weights.
Thus, Faster-RCNN NASNet was the slowest in predictions. Its mean processing time
per one image was 880 milliseconds. The model based on the MobileNet backbone was
the fastest one with the inference time per one image of 43 milliseconds. Thus, it may
be used for predicting the location of stenosis in real-time.</p>
        <p>Our results suggest that Faster-RCNN ResNet-50 V1 is the most accurate model.
The mean Average Precision of this model on the validation set is 0.92 with the
inference time of 98 milliseconds per image ( 10 frames per second). The fastest and
relatively lightweight SSD MobileNet V1 model has the mean Average Precision of
0.70 with the inference time of 43 milliseconds per image ( 23 frames per second).
Faster-RCNN NASNet has over a 3-fold advantage in the number of weights compared</p>
        <p>Model</p>
        <p>SSD
MobileNet V1
Faster-RCNN
ResNet-50 V1
Faster-RCNN</p>
        <p>NASNet
to Faster-RCNN ResNet-50 V1. However, the accuracy of Faster-RCNN ResNet-50 V1
is 12% higher than that of the Faster-RCNN NASNet model. Therefore, Faster-RCNN
ResNet-50 V1 seems to be an optimal solution, capable of processing data quickly with
a relatively high level of accuracy.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Models testing</title>
        <p>The capabilities of the selected neural networks are presented using the data of two
patients with the referenced labeling, whose data were not used to train the models.
Fig. 6 shows the images with the red-marked area of stenosis. The models with the
best values of the loss function and mAP were used for testing. Table 4 reports the best
steps with the model optimal weights. The resultant predicting probability and location</p>
        <p>Automatic coronary stenosis detection in X-ray angiography... 7
of stenosis are shown in Fig. 7. Additionally, Intersection over Union (IoU) and Dice
Similarity Coefficient (DSC) metrics were used to assess the localization performance.</p>
        <p>The comparative study proves that the models may accurately detect the location
of stenosis. However, there were several false positives, while testing Faster-RCNN
NASNet. In both cases, this model detected the location of the false stenotic segment
in the right coronary artery and the anterior descending artery (two segments) besides
the reference with a probability of over 90% (see Fig. 7c). However SSD MobileNet
V1 erroneously predicted the absence of stenosis of the first patient. In addition, the
efficiency of the detector based on the ResNet architecture, Faster-RCNN ResNet-50
V1, should be noted. The average DSC metric on the test data for this model was 0.85.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We trained three well-known and promising detectors based on different neural
network architectures (MobileNet, ResNet-50, NASNet) to locate the single-vessel
disease. Faster-RCNN ResNet-50 V1 has an optimal accuracy-to-speed ratio. This model
demonstrates the mean Average Precision (mAP@0.5 metric) of 0.92 at 10 frames per
second. The fastest and relatively lightweight SSD MobileNet V1 model has the mean
Average Precision of 0.70 at 23 frames per second. The ML-based approach proposed
in this study is of particular interest in detecting the location of multivessel coronary
artery disease. This approach ensures accurate detection of the stenosis location and
may provide additional characteristics of the stenotic segment, such as its length,
diameter, lateral branches, etc.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>Data mining, data pre-processing and development of the ML-based approach to
detect stenosis were supported by a grant from the Russian Science Foundation, project
No. 18-75-10061 ”Research and implementation of the concept of robotic minimally
invasive prosthetics of the aortic valve”. The training of the developed models using
Amazon Web Services was funded by the Ministry of Science and Higher Education,
project No. FFSWW-2020-0014 “Development of the technology for robotic
multiparametric tomography based on big data processing and machine learning methods for
studying promising composite materials”. The selection of the primary metrics
suggesting the model performance and their analysis was supported by the grant of the Russian
Foundation of Basic Research, project No. 19-07-00351=19 “Methods and intelligent
technologies for the scientific justification of strategic solutions on digital
transformation”.
Automatic coronary stenosis detection in X-ray angiography... 9</p>
      <p>(a) SSD MobileNet V1
(b) Faster-RCNN ResNet-50 V1</p>
      <p>(c) Faster-RCNN NASNet</p>
      <p>Fig. 7. Example of new data prediction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>GBD 2017 Causes of Death Collaborators: Global, regional, and national age-sex-specific mortality for 282 causes of death in 195 countries and</article-title>
          territories, 1980
          <article-title>-2017: a systematic analysis for the Global Burden of Disease Study 2017</article-title>
          .
          <string-name>
            <surname>Lancet</surname>
          </string-name>
          (London, England)
          <volume>392</volume>
          (
          <issue>10159</issue>
          ),
          <fpage>1736</fpage>
          -
          <lpage>1788</lpage>
          (nov
          <year>2018</year>
          ). https://doi.org/10.1016/S0140-
          <volume>6736</volume>
          (
          <issue>18</issue>
          )
          <fpage>32203</fpage>
          -
          <lpage>7</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Antczak</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liberadzki</surname>
          </string-name>
          , Ł.:
          <article-title>Stenosis Detection with Deep Convolutional Neural Networks</article-title>
          .
          <source>MATEC Web of Conferences</source>
          <volume>210</volume>
          , 04001 (oct
          <year>2018</year>
          ). https://doi.org/10.1051/matecconf/201821004001
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toe</surname>
            ,
            <given-names>K.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhong</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Stenosis detection and quantification on cardiac CTCA using panoramic MIP of coronary arteries</article-title>
          .
          <source>In: 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</source>
          . pp.
          <fpage>4191</fpage>
          -
          <lpage>4194</lpage>
          . IEEE (jul
          <year>2017</year>
          ). https://doi.org/10.1109/EMBC.
          <year>2017</year>
          .8037780
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Slomka</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arsanjani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakazato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ko</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berman</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuo</surname>
            ,
            <given-names>C.C.J.</given-names>
          </string-name>
          :
          <article-title>Structured learning algorithm for detection of nonobstructive and obstructive coronary plaque lesions from computed tomography angiography</article-title>
          .
          <source>Journal of Medical Imaging</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ),
          <volume>014003</volume>
          (mar
          <year>2015</year>
          ). https://doi.org/10.1117/1.jmi.
          <volume>2</volume>
          .1.
          <fpage>014003</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Slomka</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakazato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arsanjani</surname>
          </string-name>
          , R., Cheng, V.Y.,
          <string-name>
            <surname>Min</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berman</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jay Kuo</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automated knowledge-based detection of nonobstructive and obstructive arterial lesions from coronary CT angiography</article-title>
          .
          <source>Medical Physics</source>
          <volume>40</volume>
          (
          <issue>4</issue>
          ),
          <volume>041912</volume>
          (apr
          <year>2013</year>
          ). https://doi.org/10.1118/1.4794480
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Toledano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindenbaum</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lessick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dragu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghersin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Learning to Detect Coronary Artery Stenosis from Multi-Detector CT imaging</article-title>
          .
          <source>Tech. rep.</source>
          , Technion - Israel Institute of Technology, Haifa (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. de Vos,
          <string-name>
            <given-names>B.D.</given-names>
            ,
            <surname>Wolterink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Leiner</surname>
          </string-name>
          , T., de Jong, P.A.,
          <string-name>
            <surname>Lessmann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isˇgum</surname>
          </string-name>
          , I.:
          <article-title>Direct Automatic Coronary Calcium Scoring in Cardiac and Chest CT</article-title>
          .
          <source>IEEE transactions on medical imaging 38(9)</source>
          ,
          <fpage>2127</fpage>
          -
          <lpage>2138</lpage>
          (sep
          <year>2019</year>
          ). https://doi.org/10.1109/TMI.
          <year>2019</year>
          .2899534
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zreik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van</surname>
            <given-names>Hamersvelt</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.W.</given-names>
            ,
            <surname>Wolterink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Leiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Viergever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.A.</given-names>
            ,
            <surname>Isˇgum</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.:</surname>
          </string-name>
          <article-title>A Recurrent CNN for Automatic Detection and Classification of Coronary Artery Plaque and Stenosis in Coronary CT Angiography</article-title>
          .
          <source>IEEE Transactions on Medical Imaging</source>
          <volume>38</volume>
          (
          <issue>7</issue>
          ),
          <fpage>1588</fpage>
          -
          <lpage>1598</lpage>
          (jul
          <year>2019</year>
          ). https://doi.org/10.1109/TMI.
          <year>2018</year>
          .2883807
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anguelov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reed</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>C.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          : SSD:
          <article-title>Single Shot MultiBox Detector</article-title>
          .
          <source>In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          , vol.
          <volume>9905</volume>
          LNCS, pp.
          <fpage>21</fpage>
          -
          <lpage>37</lpage>
          . Springer Verlag (dec
          <year>2016</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -46448-0 2
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <string-name>
            <surname>Faster</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          :
          <article-title>Towards Real-Time Object Detection with Region Proposal Networks</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>39</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1137</fpage>
          -
          <lpage>1149</lpage>
          (jun
          <year>2017</year>
          ). https://doi.org/10.1109/TPAMI.
          <year>2016</year>
          .2577031
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rathod</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korattikara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fathi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wojna</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guadarrama</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Speed/accuracy trade-offs for modern convolutional object detectors</article-title>
          .
          <source>In: Proceedings - 30th IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2017</year>
          . vol. 2017-January, pp.
          <fpage>3296</fpage>
          -
          <lpage>3305</lpage>
          .
          <article-title>Institute of Electrical and Electronics Engineers Inc</article-title>
          . (nov
          <year>2017</year>
          ). https://doi.org/10.1109/CVPR.
          <year>2017</year>
          .351
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalenichenko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weyand</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andreetto</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adam</surname>
          </string-name>
          , H.:
          <article-title>MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications</article-title>
          (apr
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sandler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhmoginov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.C.</given-names>
          </string-name>
          :
          <article-title>MobileNetV2: Inverted Residuals and Linear Bottlenecks</article-title>
          .
          <source>In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <fpage>4510</fpage>
          -
          <lpage>4520</lpage>
          . IEEE (jun
          <year>2018</year>
          ). https://doi.org/10.1109/CVPR.
          <year>2018</year>
          .00474
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source>
          . vol. 2016-Decem, pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . IEEE Computer Society (dec
          <year>2016</year>
          ). https://doi.org/10.1109/CVPR.
          <year>2016</year>
          .90
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Identity mappings in deep residual networks</article-title>
          .
          <source>In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          . vol.
          <volume>9908</volume>
          LNCS, pp.
          <fpage>630</fpage>
          -
          <lpage>645</lpage>
          . Springer Verlag (
          <year>2016</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -46493-0 38
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zoph</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Neural Architecture Search with Reinforcement Learning</article-title>
          .
          <source>5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings (nov</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zoph</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasudevan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shlens</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Learning Transferable Architectures for Scalable Image Recognition</article-title>
          .
          <source>In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <fpage>8697</fpage>
          -
          <lpage>8710</lpage>
          . IEEE (jun
          <year>2018</year>
          ). https://doi.org/10.1109/CVPR.
          <year>2018</year>
          .00907
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Girshick</surname>
          </string-name>
          , R.:
          <string-name>
            <surname>Fast R-CNN</surname>
          </string-name>
          .
          <source>In: 2015 IEEE International Conference on Computer Vision</source>
          (ICCV). pp.
          <fpage>1440</fpage>
          -
          <lpage>1448</lpage>
          . IEEE (dec
          <year>2015</year>
          ). https://doi.org/10.1109/ICCV.
          <year>2015</year>
          .169
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Focal Loss for Dense Object Detection</article-title>
          .
          <source>In: 2017 IEEE International Conference on Computer Vision (ICCV)</source>
          . vol.
          <volume>42</volume>
          , pp.
          <fpage>2999</fpage>
          -
          <lpage>3007</lpage>
          . IEEE (oct
          <year>2017</year>
          ). https://doi.org/10.1109/ICCV.
          <year>2017</year>
          .324
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>