<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Technology for Indoor Drone Positioning Based on CNN Detector</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>V.A. Gorbachev</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu.B. Blokhinov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.D. Nikitin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E.E. Andrienko</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>vadim.gorbachev@gosniias.ru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>yury.blokhinov@gosniias.ru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FSUE “GosNIIAS”</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moscow</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>The article presents the drone positioning technology in a multi-camera system by using the detection algorithm. Paper describes positioning system and algorithm for calculating 3d drone coordinates based on its image position, detected on images of stationary video cameras. Positioning enables automatically control the drone when precise data from satellite navigation systems are not available, for example, in closed hangars. The developed technology is used to create a complex of automatic visual control of aircraft. The ways of adaptation of neural network detection algorithm to the problem of drone detection are presented. The main attention is paid to the methods of training data preparation. It is shown that high accuracy can be achieved using synthesized images without any real data or manual labelling.</p>
      </abstract>
      <kwd-group>
        <kwd>object detection</kwd>
        <kwd>neural networks</kwd>
        <kwd>drones</kwd>
        <kwd>positioning</kwd>
        <kwd>indoor navigation</kwd>
        <kwd>multi-camera system</kwd>
        <kwd>image synthesis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Currently, due to the increase in the aircraft flow in the
airspace, the complexity of their timely and high-quality visual
inspection during the maintenance at the airport has increased
significantly. Significantly increased the total downtime of
aircraft during unscheduled inspections, caused, for example,
the impact of atmospheric electricity on the surface of the
fuselage of the aircraft in flight. External human inspection of
hard-to-reach areas of the aircraft, such as the upper fuselage or
tail, aimed to identify the effects of lightning today takes a
significant time, leading to downtime of aircraft or even flight
delays. For companies which have a fleet of more than 200
aircraft, such as Aeroflot, such an event is not uncommon:
according to the company, it occurs about 300-400 times a year,
leading to significant time and financial losses. Large
companies such as Airbus, Lufthansa, EasyJet, American
Airlines start applying drones to solve the problems of
accelerating the visual inspection of the aircraft. However,
currently, the use of drones is carried out in manual mode,
which does not allow to completely reveal the potential of the
technology. According to experts, the use of programmable
drones will significantly reduce the time of inspection of the
aircraft and, no less significantly, make the technology itself
completely digital.</p>
      <p>The article proposes an approach to the creation of
automated drone control technology based on its real time
positioning using a system of stationary cameras. This
technology is necessary to ensure the functioning of the drone
control system in enclosed spaces such as aircraft hangars. The
development of a special positioning technology is necessary,
since the signals of global satellite navigation systems (GPS,
GLONASS, etc.) may be partially or completely inaccessible in
the hangar where aircraft maintenance is carried out. At the
same time, the inertial navigation system of the drone can’t
provide sufficient accuracy throughout its flight. Due to the fact
that the flight of the drone must be carried out at a short
distance from the aircraft (no more than 1.5 meters), ensuring
the accuracy of the trajectory is a critical aspect for the safety
and applicability of the technology. Visual positioning system is
the most preferable in the described conditions, as it is able to
provide sufficient accuracy, it does not require the installation
of additional equipment on the drone, it is passive, so, it does
not emit any radio or other signals except Wi-Fi.</p>
      <p>
        During maintenance, the drone flies over the aircraft on a
programmed trajectory and makes a high resolution video of the
surface of the fuselage and wings (Fig. 1). Based on the
coordinates obtained from the visual positioning system, the
onboard drone control system monitors compliance with the
choosen trajectory. By results of the automatic analysis of the
received videos the decision on existence of damages on a
covering of aircraft is made. This technology allows complete
automating the process of visual inspection of aircraft [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The paper describes the features of creating such a
technology in terms of positioning drones through the use of
CNN-based detectors.</p>
      <p>Fig. 1. The drone flight over the aircraft during the tests.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Review of detection algorithms</title>
      <p>The proposed technology is based on an algorithm for
detecting objects in images (namely, video frames). The most
advanced detection algorithms today are algorithms based on
deep convolutional neural networks (CNN). Neural network
architectures for detection are divided into two main types:
single-stage and two-stage. In two-stage approaches, the task of
detecting objects is divided into two steps: identifying areas of
interest, then classifying the class of object in the area, and
predicting the parameters of the bounding box.</p>
      <p>
        The two-stage approach was first introduced in 2014 by
Girshik [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. His work R-CNN (Regions with CNNs) uses a
selective search method [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to detect regions of interest in input
images and uses a regional classifier based on DCN
(Deformable Convolutive Networks) to self-classify regions of
interest. Fast-RCNN [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] improves R-CNN by extracting regions
of interest from feature maps. Faster R-CNN [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is a
modification of the method of Fast R-CNN and R-CNN. The
method is based on the idea of region proposals. The key
difference between Faster R-CNN and its predecessors is that
regions are calculated not from the original image, but from the
feature map obtained from the convolutional neural network. To
do this, a module called Region Proposal Network (RPN) was
added. Obtained with the help values are passed to two parallel
fully connected branches: bounding box prediction (regression)
and classification framework. The outputs of these layers are
based on the so called anchor areas (ancor boxes) – several
frames for each position of the window, having different sizes
and aspect ratios. The regression layer for each such rectangle
produces 4 parameters that adjust the position of the bounding
rectangle, and the classification layer produces the probability
that the rectangle contains an object and the probability that the
object in the frame corresponds to each of the classes. Cascade
R-CNN [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] solves the problem of increasing the accuracy of the
bounding box detection by applying a sequence of detectors
with varying thresholds.
      </p>
      <p>In single-stage approaches, there is no stage of finding
regions of interest, the regression of bounding boxes and the
classification</p>
      <p>
        of candidates in anchor areas is performed
directly.
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] uses a small number of anchor regions (dividing the input
image with a rectangular grid) and is based on the VGG-16
neural network. YOLOv2 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] improves the performance due to
the use of a new method of bounding the regression framework
and a new neural network Darknet-19. YOLOv3 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] continues
to improve Darknet-19, offering a deeper neural network with
skip connections. Architecture YOLOv2 and YOLOv3 allow to
change the balance between accuracy and speed of detection by
varying the number of areas able to solve the problem of
detection in real time.
      </p>
      <p>
        A slightly different approach is used by CenterNet [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], a
detection algorithm based on methods for key points detection
using neural networks. It learns to predict the centers of objects
and form a feature
      </p>
      <p>
        map. The parameters of the bounding
rectangle are then regressed for the detected centers. Corner Net
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] is another detection
algorithm
based
on
key
points
prediction. Unlike the CenterNet, CornerNet detects an object
using a pair of corners of its frame.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Indoor positioning system</title>
      <p>calculated coordinates of the object are transmitted to its
onboard control system via Wi-Fi channel. The scheme of the
proposed navigation system is shown in Fig. 2.</p>
      <p>
        To calculate the three-dimensional coordinates of the object
based on its position in the images, a method is used, which is a
special case of block triangulation by the method of ligaments
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Since the camera orientation parameters are known, only
three unknown 3D coordinate values are calculated. The idea of
the method is to minimize the deviation of the projection of the
calculated three-dimensional point on the image from the real
position of the object (more precisely, the sum of squared errors
for all cameras). The projection equations are:
  = −   3(  − )+  3(   − )+ 3(  − ),
      </p>
      <p>1(  − )+  1(   − )+ 1(  − )
  = −   3(  − )+  3(   − )+ 3(  − ),
 2(  − )+  2(   − )+ 2(  − )
(1)
(2)
where (   ,</p>
      <p>,    ) are camera positions, ( ,  ,  ) is 3D object
position, (xi,yi) is its projection on image i, fi is focus distances,
 1  2  3
  = ( 1

 1
 2

 2
 3)

 3
is rotation matrix for camera i.</p>
      <p>ATA ΔX + ATB = 0,</p>
      <p>This is a well-known problem, which is solved by the
method of iterative approximations. Each increment step of the
three-dimensional coordinates ΔX is determined from the
solution of the system of equations:
where A is the matrix of partial differential of projection
equations (1),(2) by drone coordinates over all cameras (size
3*3*number of cameras in the system), B is the discrepancy
vector (size 2*number of cameras), containing deviations of
object projections from real positions on images.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Detection algorithm details</title>
      <p>
        As the detection algorithm YOLOv2 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] CNN architecture
was used. This architecture is slightly concede to YOLOv3 in
accuracy, but has a higher calculaton speed, and demonstrates
one of the best ratios of accuracy and performance, which in
our task is of key importance. Performance determines the
frequency of control signals delivered to the drone, which
directly affects the accuracy of control and maximum safe flight
speed.
      </p>
      <p>The network receives a three-channel image as input, and
outputs a tensor of size X×X×Y, where X is the number of cells
in the input image. The length of the tensor Y depends on the
number of classes detected and the number of anchor regions in
the cell. For each
anchor area, 5
basic parameters are
calculated: the coordinates of the upper left corner of the
rectangle, the width, the height, and the probability that this
rectangle contains any object. In addition, the probability of the
object belonging to each selected class is determined. The
hyperparameters of the algorithm are the number and size of the
anchor areas and the size of the input image.</p>
      <p>The image size determines the number of cells for which
the features are calculated, since the cell size is fixed and is
equal to
32x32
pixels. Therefore, it directly
affects the
performance and accuracy of the network, as the number of
cells increases the number of network filters. On the other hand,
if there are more cells, each of them contains fewer objects; the
features calculated in it correspond more accurately to each
object and allow to build a more reliable prediction. The plot in
number of anchors is set to 5, as the higher number of anchor
areas decreases performance. K-means clustering of bounding
rectangles on our training data set was used to determine anchor
sizes.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Automated data preparation</title>
      <p>Since the work uses AI detection algorithms, training data is
required to learn them. In our case, data are images with
annotation: the type of object and its coordinates (bounding
box) in the image. The CNN detectors used are very flexible but
have a very large number of parameters. In this regard, a lot of
training data is required. In order to avoid time-consuming
manual data labelling, automatic synthesis of images was used
for training and testing the algorithm.</p>
      <p>
        The data were synthesized based on the rendering of the
existing three-dimensional model of the drone (Fig. 4). Special
3D environments were not used during data endering, as their
preparation requires additional manual labor of designers.
Instead, the process was structured as follows. The model of the
drone in different angles was rendered in a 3D modeling system
on a uniform-colored background. The object in the image was
automatically cut out, and its mask was built. Then the image
and mask were subjected to random transformations: rotation,
scaling, displacement, reflection, perspective transformation,
blurring, salt/pepper noise, shift of color channel values (Fig.
5). After that, the image of the object on his mask was ovelayed
on arbitrary backgrounds. Both random images and images
from the test hangar where the subsequent testing was carried
out were used as backgrounds. In order to make such insertion
look natural and the network did not remember overlay artifacts
as informative features of the object, local smoothing of objects
with a Gaussian filter with randomized intensity was performed.
In addition, objects from the Coil-100 collection were added to
the images to increase the discriminative ability of the network
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>To prepare the test data and expand the training base
through real images, the real drone flight and video capture in
the test hangar were carried out (Fig. 6). To get rid of manual
annotation video files, the following automatic labellng
algorithm was used. Optical flow maps were calculated for each
video frame. The area with the maximum magnitude of the
optical flow was selected on the maps. Since normally there
were no other moving objects in the experimental scene, this
area was thought to correspond to a drone. Sometimes due to
the presence of foreign moving objects and shadows, as well as
inaccuracy of segmentation, such labelling contained several
errors. An experimental study was devoted to the estimate of
the influence of different types of training data on the detection
results.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Experiments</title>
      <p>The accuracy of the detection algorithms was tested in a
series of computational experiments on various training
collections. We had three main data collections: synthetic,
where images of drones were obtained by rendering 3D models,
and the backgrounds are taken arbitrarily; semi-synthetic, where
backgrounds for rendering was real images of the hangar in
which the experiment was carried out; and autolabelled real,
obtained by automated labelling drone videos (using optical
flow). Incorrectly labelled data was manually deleted. Testing
data was the part of the real data collection that was not used for
training. The obtained precision, recall and IoU for different
sets of training data are shown in table 1.</p>
      <p>According to the results of the experiments, the following
conclusions can be drawn. First, it is possible to train the
algorithm with high accuracy on fully synthetic data, which was
the purpose of the work. Secondly, the smoothing of objects
when overlaying them on the background image plays a crucial
role. Without smoothing, artifacts at the boundaries of objects
become too important feature for the neural network, and it
overfits to detect only artificial objects. Third, the use of a large
number of random backgrounds was better than the use of a
small number of real backgrounds from the test hangar. Despite
the fact that the background images on the test data were similar
to the training ones (but not the same), the network overfits, that
means it has a low generalizing ability and does not cope with
new scenes. Fourth, the inclusions of random objects
(distractor) in the training images allowed significantly improve
the accuracy of the work. Although these objects are not
labelled in the test data, the network has learned to better
distinguish drones from any other objects (see table 1).</p>
      <p>In addition, during the experiments it was found that when
training the network on the data obtained by the
abovedescribed autolabelling method, the accuracy was worse than on
synthetic data. This is due to the fact that optical flow map is
blurred, and the resulting bounding box is greater than the real
object bounding box (Fig. 6). Also, the available real data are
not sufficiently diverse.
7. Conclusion
100%
27.53%
18.25%
41.13%
45.56%
33.92%</p>
      <p>Recall
98.69%
97.71%
86.60%
92.48%
93.32%
37.91%</p>
      <p>IoU
98.65%
97.63%
86.41%
91.58%
98.64%
37.89%</p>
      <p>The paper describes the indoor drone positioning
technology based on stationary visual sensors and the algorithm
of drone detection. Given camera orientation and detection
results the 3D position is reconstructed using a special
algorithm of iterative minimization of the total reprojection
error. The ways of adaptation of the CNN-based detector to the
subject area were investigated. Both the automated process of
creating training data and hyperparameter tuning are described.
The influence of the data generation methods on the result is
studied, in particular the inclusion of distracting objects in the
data, artifacts of object overlay, the use of various background
images. Conducted experiments showed that it is possible to
train high accuracy detector exclusively on automatically
synthesized images obtained using the renderings of a
threedimensional model of the drone without any real samples.</p>
    </sec>
    <sec id="sec-7">
      <title>8. Acknowledgements</title>
      <p>This work was supported by the Russian Foundation for
Basic Research, project no. 17-08-00191 a.</p>
    </sec>
    <sec id="sec-8">
      <title>9. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yu</surname>
            .
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Blokhinov</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          <string-name>
            <surname>Gorbachev</surname>
            ,
            <given-names>A.D.</given-names>
          </string-name>
          <string-name>
            <surname>Nikitin</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          <string-name>
            <surname>Skryabin</surname>
          </string-name>
          .
          <article-title>Technology for Visual Inspection of Aircraft Surfaces using Programmable Unmanned Aerial Vehicles</article-title>
          .
          <source>Journal of Computer and Systems Sciences International. Received by the editor 28.06</source>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Donahue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Darrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Malik</surname>
          </string-name>
          .
          <article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , p.
          <fpage>580</fpage>
          -
          <lpage>587</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Uijlings</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. E. Van De Sande</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Gevers</surname>
            ,
            <given-names>A. W.</given-names>
          </string-name>
          <string-name>
            <surname>Smeulders</surname>
          </string-name>
          .
          <article-title>Selective search for object recognition</article-title>
          .
          <source>International journal of computer vision</source>
          ,
          <volume>104</volume>
          (
          <issue>2</issue>
          ):
          <fpage>154</fpage>
          -
          <lpage>171</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick. Fast</surname>
          </string-name>
          r-cnn.
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          , p.
          <fpage>1440</fpage>
          -
          <lpage>1448</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun. Faster</surname>
          </string-name>
          R-CNN:
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , p.
          <fpage>91</fpage>
          -
          <lpage>99</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cai</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Vasconcelos. Cascade</surname>
          </string-name>
          R-CNN:
          <article-title>Delving into high quality object detection</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , pages
          <fpage>6154</fpage>
          -
          <lpage>6162</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Anguelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Erhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reed</surname>
          </string-name>
          , C.- Y. Fu,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Berg</surname>
          </string-name>
          . SSD:
          <article-title>Single shot multibox detector</article-title>
          .
          <source>ECCV</source>
          , p.
          <fpage>21</fpage>
          -
          <lpage>37</lpage>
          . Springer,
          <year>2016</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.-Y.</given-names>
            <surname>Fu</surname>
          </string-name>
          , W. Liu,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ranga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tyagi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Berg</surname>
          </string-name>
          . DSSD:
          <article-title>Deconvolutional single shot detector</article-title>
          .
          <source>arXiv preprint arXiv:1701.06659</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Divvala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          .
          <article-title>You only look once: Unified, real-time object detection</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , p.
          <fpage>779</fpage>
          -
          <lpage>788</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Farhadi.</surname>
          </string-name>
          <article-title>Yolo9000: better, faster, stronger</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , p.
          <fpage>7263</fpage>
          -
          <lpage>7271</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Farhadi.</surname>
          </string-name>
          <article-title>Yolov3: An incremental improvement</article-title>
          .
          <source>arXiv preprint arXiv:1804.02767</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Krähenbühl</surname>
          </string-name>
          .
          <article-title>Object as Points</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .07850v2,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Law</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          . Cornernet:
          <article-title>Detecting objects as paired keypoints</article-title>
          .
          <source>In Proceedings of the European conference on computer vision</source>
          , p.
          <fpage>734</fpage>
          -
          <lpage>750</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.P.</given-names>
            <surname>Mikhailov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.G.</given-names>
            <surname>Chibunichev. Photogrammetry - MIIGAIK Publishing</surname>
          </string-name>
          , Moscow,
          <year>2016</year>
          , 294 p.
          <article-title>(In Russian language</article-title>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Nene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Nayar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Murase</surname>
          </string-name>
          .
          <article-title>Columbia Object Image Library (COIL-100)</article-title>
          .
          <source>Technical Report CUCS-006- 96. February</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>