<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Distributed Tracking Algorithm for Counting People in Video by Head Detection?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>nis Kuply</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Timur M</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton Konushin</string-name>
          <email>anton.konushing@graphics.cs.msu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Computational Mathematics and Cybernetics, Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NRU Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Video Analysis Technologies</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We consider the problem of people counting in video surveillance. This is one of the most popular tasks in video analysis, because this data can be used for predictive analytics and improvement of customer services, traffic control, etc. Our method is based on the object tracking in video with low framerate. We use the algorithm from [1] as a baseline and propose several modifications that improve the quality of people counting. One of the main modifications is to use a head detector instead of a body detector in the tracking pipeline. Head tracking is proved to be more robust and accurate as the heads are less susceptible to occlusions. To find the intersection of a person with a signal line, we either raise the signal lines to the level of the heads or perform a regression of bodies based on the available head detections. Our experimental evaluation has demonstrated that the modified algorithm surpasses original in both accuracy and computational efficiency, showing a lower counting error on a lower detection frequency.</p>
      </abstract>
      <kwd-group>
        <kwd>Computer Vision</kwd>
        <kwd>Video Analytics</kwd>
        <kwd>Tracking</kwd>
        <kwd>People Counting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Counting people passing through certain zones of a public infrastructure, such as
pedestrian crossings, sidewalks, squares, etc., is a practically important task. There are many
solutions to this problem, one of them is object tracking. The task of object tracking
is to create tracks for each person. A track unambiguously corresponds to a person. It
marks this particular person locations in all frames in which he or she is visible. In order
to count people a signal line is usually specified in the frame (see fig. 1). If the track
crosses the signal line, we can say with confidence that the person also crossed it.
? Publication is supported by RFBR grant 18-08-01484</p>
      <p>We propose a fully automatic people counting algorithm. The algorithm takes as
input a video stream fFigi=1 of frames captured by stationary camera and signal line that
specified by an ordered pair of points (La; Lb) on the frame. The output of the
algorithm is a set of events fEigi=1 that represented by triples of values Ei = (ki; ri; di).
The first value indicates the number of the frame where the signal line was crossed, the
second value specifies the coordinates of the bounding box and the last value indicates
the direction of the signal line intersection.</p>
      <p>
        Our solution is an extension of the algorithm described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this paper we
propose the following improvements:
– use of the head detector instead of full-body detector;
– detection matching procedure is modified in order to work with small head
bounding boxes;
– the use of body regression on the heads at different stages of tracking;
– an algorithm to automatically determine a region of interest (ROI) by signal line
position is proposed to speed up detection.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The modern methods of tracking are based on tracking-by-detection. There are a lot
of ways to detect the desired object on the frame. Three most popular methods are:
detection of body [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2,3,4</xref>
        ], detection of head [
        <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
        ] and using of key points. The first
solution is popular because there are many datasets and ready-made solutions.
      </p>
      <p>The head-tracking approach is well suited to track people in a crowd: usually video
surveillance cameras are installed above the height of the person, where heads in a
crowd can be seen better than full bodies. Heads are more resistant to overlapping than</p>
      <p>
        A Distributed Tracking Alg. for Counting People in Video by Head Det... 3
bodies. The number of ready-made solutions and data for training is less than for bodies.
There are methods that use body parts detectors for tracking [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], key points of human
pose [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], combined solutions (body and head) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and detector ensembles [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        After detection we need to bind all the detections to tracks. As in the detection task,
there are a lot of methods. The first group of algorithms for creating tracks is greedy
algorithms. In most online algorithms tracks are constructed frame by frame, each frame
creates a matrix of the cost of matching new detections and existing tracks, then the
problem of matching is solved by a greedy algorithm (searching for the maximum in
each row/column) [
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ] or by a Hungarian algorithm [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2,3,4</xref>
        ]. Sometimes MCMC
is used to bind all the detections to tracks [
        <xref ref-type="bibr" rid="ref14 ref5">5,14</xref>
        ].
      </p>
      <p>
        Recently, neural networks have been used more often in tracking. For example,
in paper [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] authors suggest using detector to obtain new detection by regression of
detection on previous frame. However, this method has disadvantages. For example,
it is able to work well only at a high frame rate and it also increases the load on the
detector.
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>The Proposed Algorithm</title>
      <sec id="sec-3-1">
        <title>Baseline</title>
        <p>
          We use the solution from [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] as the baseline, which is an extension of SORT tracking
algorithm [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We choose this algorithm as it is capable to work by detection on a sparse
set of frames, which is performed on remote servers. This significantly reduces the
amount of computational resources required for large-scale video surveillance systems
(see fig. 2). Our proposed method inherits distributed nature of baseline.
        </p>
        <p>
          Baseline works in online mode and use the Hungarian algorithm [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] to match
detections. To improve results at a low detection rate ASMS visual tracking [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] is used
to evaluate a speed of people between frames. The same approach with visual tracking
speed estimation is used in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>The baseline algorithm consists of the following steps: (1) detection; (2) evaluation
of the speed of detections using visual tracking; (3) prediction of the position of tracks
by the Kalman filter; (4) matching; (5) extrapolation of the tracks; (6) detection of signal
line crossing events.</p>
        <p>Proposed improvements to some of the steps above are described below.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Detection</title>
        <p>
          Since heads are seen better in the video and are less prone to occlusions we decided
to use the head detector based on SSD [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] approach instead of the body detector.
Another advantage of the head detector is that neural network body detectors can combine
nearby people into one bounding box, which is less frequent for heads. The detector
were trained on CrowdHuman [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] public dataset and on the dataset collected by Video
Analysis Technologies. Experimental evaluation showed AUC of 0.66 on the test part
of CrowdHuman. Usage of the head detector leads to a sub-task of restoring a bounding
box of an entire body to find intersections of the signal line, which is located on the
ground. We describe it in section 3.4.
        </p>
        <p>Datacenter</p>
        <p>. . .
During the experiments we realized that IOU metric used in the basic algorithm is not
suitable for head matching. As head bounding boxes are small enough due to the errors
in position and speed estimation bounding box pairs required to match are closely
positioned, but have no intersection. IOU equal to zero acquired in the case which leads to
track partitioning. Therefore we suggest to increase by s times the size of the bounding
boxes (while saving the center of the bounding box) before matching them using IOU
metric. This approach solves the problem described above.
We got the problem of missing signal line intersections as head tracks are located on a
human height and the signal line is located on the ground. Problem reveals itself most
on the scenes where the signal line is placed orthogonal to the camera. Therefore we
offer two solutions: rising of signal lines and body bounding box regression by head.
Rising of Signal Lines We propose to raise the signal line to the level of the human
head (see fig. 3). This solution reduces the computational cost and time of the algorithm
as it doesn’t require any extra steps to project head track to the ground plane.</p>
        <p>The signal line can be raised either manually or automatically. Automatic approach
we propose is using the following anatomical fact: the height of the head is 18 of the</p>
        <p>A Distributed Tracking Alg. for Counting People in Video by Head Det... 5
height of the entire human body. We can use a body detector to calculate average height
of human body for the scene. And raise the signal line to a height equal to 78 of the
average height of a person on the scene.</p>
        <p>This approach has a drawback: people’s growth is different, which means that
people’s heads are on different planes, while people’s legs are always on the same plane.
Therefore, the raised signal line doesn’t look so clearly since it’s not clear where the
plane is located.
Bounding Box Regression by Head To solve the drawbacks that arose in the previous
paragraph we propose to use regression of a body by a head, which is trained for the
specific scene. Our idea consists of two stages: combining the detections of heads and
bodies for the scene and training linear regression on the received data.
Combining head and body detections First of all we launch head and body detectors.
To combine heads and bodies a matrix of head and body correspondence is built on
each frame.</p>
        <p>CM k; i; j = costmatch(Bk; i; Hk; j ); i 2 [1; nbk]; j 2 [1; nhk]
(1)</p>
        <p>In eq. (1) nbk; nhk – the number of body and head detections on the k-th frame,
Bk; i; Hk; j – the bounding box of the i-th body and j-th head on the k-th frame. We
calculate matching cost as:
costmatch(Bk; i; Hk; j ) =
(iohk; i; j ; if iohk; i; j
0; else</p>
        <p>I1 ^ 1
dist(Bk; i; Hk; j )
iohk; i; j =
area (intersection(Bk; i; Hk; j ))</p>
        <p>area(Hk; j )
dist(Bk; i; Hk; j ) =
centery(Hk; j )
height(Hk; j )
top(Bk; i)</p>
        <p>In eq. (2) I1 is threshold for iohk; i; j (we are using I1 = 0:5), 1; 2 – minimum
and maximum normalized distances between the center of the head and the upper point
of the body (we are using 1 = 1 and 2 = 1). So at least half of head bounding box
area should be inside body bounding box area and vertically head center shouldn’t be
far from body top by at least one head height.</p>
        <p>Next the assignment problem is solved using the Hungarian algorithm to
maximize cost. Combined head and body detections with non-zero match cost form training
dataset for a regression (see fig. 4).</p>
        <p>Linear regression After the previous step we have the data to train a linear regression.
The training dataset consists of head and body bounding box pairs: (Bi; Hi). The
regressor predicts the following values: (Bh; Bw; shif tx), where Bh; Bw – height and
width of the body bounding box, shif tx – a shift from the center of the head to the
center of the predicted body normalized by head width:
2
(2)
(3)
(4)
(5)
(7)
(8)
Bcx = centerx(H)</p>
        <p>Bcw + shif txHw
2</p>
        <p>Bcy = Hy</p>
        <p>This approach helps us keep the signal line on the ground plane, which solves
drawbacks of the previous approach.</p>
        <p>shif tx =
centerx(B)</p>
        <p>centerx(H)</p>
        <p>Hw
Prediction is done by linear regression with quadratic members:
Bch = k0 + k1Hx + k2Hy + k3Hw + k4Hh + k5Hx2 + k6HxHy+</p>
        <p>The same way Bcw; s\hif tx are predicted, but with separate sets of coefficients. After
learning coefficients of linear regression on training dataset, we use predicted value to
restore body bounding box:
: : : + k13HhHw + k14Hh2
(6)</p>
        <p>A Distributed Tracking Alg. for Counting People in Video by Head Det... 7</p>
        <p>Now we have a choice to do speed estimation by visual tracking with head or
regressed body detections. As heads are less prone to occlusions visual tracking of them
may be more reliable. But body bounding boxes are several times larger and have more
visual information to track. We check both choices on experimental evaluation.
The entire frame is not required to find events of crossing the signal line. We can limit
the detection area to the area around the signal line – region of interest (ROI). This
solution speeds up the SSD detector as it is slower for images with higher resolution.</p>
        <p>If the signal line is horizontal then the ROI is located on the top and bottom of the
line. Otherwise the ROI is located on the left and right side of the line, as well as on the
top of the line.</p>
        <p>Let be 2 – the minimum angle between signal line and horizon, wmean; hmean
– the average width and height of body detections and (xa; ya), (xb; yb) – the
coordinates of the beginning and the end of the signal line.</p>
        <p>x1 = min(xa; xb)
swwmean</p>
        <p>1 +
x2 = max(xa; xb) + swwmean</p>
        <p>1 +
y1 = min(ya; yb)</p>
        <p>shhmeancos
y2 = max(ya; yb) + shhmeancos
sin</p>
        <p>2
sin
2
(9)
(10)
(11)
(12)</p>
        <p>Then A = (x1; y1; x2; y2) is the region of interest. sw; sh in equations 9-12 are
parameters. After visual testing we selected the following values for these variables:
sw = 2; sh = 1.
4
4.1</p>
      </sec>
      <sec id="sec-3-3">
        <title>Datasets</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Evaluation</title>
      <p>
        For an experimental evaluation of our algorithm we need datasets filmed by static
camera with body tracks markup. Video sequences should be long enough to evaluate people
counting quality. If dataset provides head tracks markup it allows us to check effect of
body regression by comparing with head tracking and raised signal lines (section 3.4).
Most of the public datasets including popular MOTChallenge dataset [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] have short
videos, only body tracks markup or filmed by moving camera. So we used 19 videos
from the collection of the Video Analysis Technologies company and the Towncentre
dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to test our algorithm. For all videos signal lines were manually drawn at
ground level as well as at head level. The table 1 provides detailed information about
each test video.
4.2
      </p>
      <sec id="sec-4-1">
        <title>Metrics</title>
        <p>
          As a quality metric we use the average error of counting the number of intersections
(events) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. The resulting events can include both true and false ones. The false events
        </p>
        <p>
          A Distributed Tracking Alg. for Counting People in Video by Head Det... 9
have no correspondences in the reference labeling. We say that an event Ei in the input
set of data matches the event Eci in the reference labeling if they correspond to the same
person crossing the signal line at the same time. We match all events as described in the
paper [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>After the events have been matched we divide videos to segments with 10 reference
events and calculate the following characteristics on them:
– GT seg is the number of reference events on the segment;
– F P seg is the number of unmatched events from the algorithm on the segment;
– F N seg is the number of unmatched events from the reference events on the
segment;
– Eseg = F P seg F Nseg is an error on the segment.</p>
        <p>GT seg</p>
        <p>i=1 ENseg , where N is the number of the
segThen final error is calculated as E = PN
ments.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Experimental Results</title>
        <p>Rising of Signal Lines At first we have tested baseline with modifications proposed in
sections 3.2, 3.3 and manually raised signal lines as described in section 3.4. The
algorithm is marked as heads-no-regression in the results table (see table 2) and parameter
s shows how many times head bounding boxes have been increased.</p>
        <p>Experiments with rising the signal line clearly shows the advantage of the head
detector over the body detector. The head detector gives a significant increase in an
accuracy reducing the error by 2 times. The algorithm without matching modification
(s = 1) has poor results at low FPS due to the previously voiced problems which shows
importance of the modification proposed in the section 3.3.</p>
        <p>Body Bounding Box Regression by Head Next we have tested body bounding box
regression by head. As mentioned in section 3.4 there are two alternative ways to apply
it. There are two algorithms we have tested (see table 3):
– heads-regression-vistrk — baseline with modifications proposed in sections 3.2,
3.3, visual tracking of regressed bodies to estimate speed and signal lines on the
ground plane (section 3.4);</p>
        <p>– heads-vistrk-regression — baseline with modifications proposed in sections 3.2,
3.3, visual tracking of heads to estimate speed and signal lines on the ground plane
(section 3.4).</p>
        <p>As you can see the second configuration gives better results. Visualization showed
us that visual tracking of heads is more reliable as they are less prone to occlusions. It
worth noting that heads-vistrk-regression performed better than heads-no-regression
(s = 2) on low detection frequency.</p>
        <p>
          Limiting the Detection Area Next we tested limiting of the detection area
(section 3.5). It allowed us to increase the speed of the algorithm almost without affecting
the counting error (see table 4). Detection area reduced by almost 60% on some of the
videos.
We have proposed the algorithm of counting people in a video, which is an extension
of the algorithm described in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Usage of the head detector and body bounding box
regression allowed to increase a counting accuracy. Detection area limiting saves
computational resources. Comparing with the baseline the proposed method is able to work
on a lower detection frequency showing lower counting error.
        </p>
        <p>
          For the future work we plan to use automatic camera calibration algorithms (see
fig. 5) [
          <xref ref-type="bibr" rid="ref20 ref21">20,21</xref>
          ]. This will allow to perform person tracking on the ground map and further
improve accuracy of people counting.
        </p>
        <p>A Distributed Tracking Alg. for Counting People in Video by Head Det... 11</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kuplyakov</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shalnov</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>A Distributed Tracking Algorithm for Counting People in Video</article-title>
          .
          <source>Programming and Computer Software</source>
          <volume>45</volume>
          (
          <issue>4</issue>
          ),
          <fpage>163</fpage>
          -
          <lpage>170</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bewley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramos</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Upcroft</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Simple Online and Realtime Tracking</article-title>
          .
          <source>CoRR abs/1602</source>
          .00763 (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Wojke</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bewley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulus</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Simple Online and Realtime Tracking with a Deep Association Metric</article-title>
          . ArXiv e-prints (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
          </string-name>
          , J.: Poi:
          <article-title>Multiple object tracking with high performance detection and appearance feature</article-title>
          .
          <source>In: European Conference on Computer Vision</source>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>42</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Benfold</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Stable Multi-Target Tracking in Real-Time Surveillance Video</article-title>
          . In: CVPR, pp.
          <fpage>3457</fpage>
          -
          <lpage>3464</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Shalnov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An improvement on an MCMC-based video tracking algorithm</article-title>
          .
          <source>Pattern Recognition and Image Analysis</source>
          <volume>25</volume>
          ,
          <fpage>532</fpage>
          -
          <lpage>540</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dehghan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oreifej</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hand</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Part-based multiple-person tracking with partial occlusion handling</article-title>
          .
          <source>In: CVPR</source>
          , pp.
          <fpage>1815</fpage>
          -
          <lpage>1821</lpage>
          . IEEE Computer Society (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Xiu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Pose Flow: Efficient Online Pose Tracking</article-title>
          . In: (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Henschel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leal-Taixe</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cremers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenhahn</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Improvements to Frank-Wolfe optimization for multi-detector multi-object tracking</article-title>
          .
          <source>CoRR abs/1705</source>
          .08314 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cobos</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abad</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A fast multi-object tracking system using an object detector ensemble</article-title>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bochinski</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Senst</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sikora</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Extending IOU Based Multi-Object Tracking by Visual Information</article-title>
          .
          <source>In: 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bochinski</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselein</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sikora</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>High-Speed tracking-by-detection without using image information</article-title>
          .
          <source>In: Advanced Video and Signal Based Surveillance (AVSS)</source>
          ,
          <year>2017</year>
          14th IEEE International Conference on, pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>H.W.:</given-names>
          </string-name>
          <article-title>The Hungarian method for the assignment problem</article-title>
          .
          <source>Naval Research Logistics (NRL) 2</source>
          (
          <issue>1-2</issue>
          ),
          <fpage>83</fpage>
          -
          <lpage>97</lpage>
          (
          <year>1955</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kuplyakov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shalnov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Markov chain Monte Carlo based video tracking algorithm</article-title>
          .
          <source>Programming and Computer Software</source>
          <volume>43</volume>
          (
          <issue>4</issue>
          ),
          <fpage>224</fpage>
          -
          <lpage>229</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Bergmann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meinhardt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leal-Taixe</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Tracking Without Bells and Whistles</article-title>
          .
          <source>In: The IEEE International Conference on Computer Vision</source>
          (ICCV), (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Vojir</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noskova</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
          </string-name>
          , J.:
          <article-title>Robust scale-adaptive mean-shift for tracking</article-title>
          .
          <source>In: Scandinavian Conference on Image Analysis</source>
          , pp.
          <fpage>652</fpage>
          -
          <lpage>663</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anguelov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reed</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , C.-Y.,
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          :
          <source>SSD: Single Shot MultiBox Detector. Lecture Notes in Computer Science</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>CrowdHuman: A Benchmark for Detecting Human in a Crowd</article-title>
          . arXiv preprint arXiv:
          <year>1805</year>
          .
          <volume>00123</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Dendorfer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rezatofighi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cremers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leal-Taixe</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>MOT20: A benchmark for multi object tracking in crowded scenes</article-title>
          . arXiv:
          <year>2003</year>
          .09003[cs] (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Shalnov</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.S.:</given-names>
          </string-name>
          <article-title>Convolutional neural network for camera pose estimation from object detections</article-title>
          . ISPRS - International
          <source>Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</source>
          <volume>42</volume>
          (
          <fpage>2</fpage>
          -
          <lpage>W4</lpage>
          ),
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Shalnov</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gringauz</surname>
            ,
            <given-names>A.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>Estimation of the people position in the world coordinate system for video surveillance</article-title>
          .
          <source>Programming and Computer Software</source>
          <volume>42</volume>
          (
          <issue>6</issue>
          ),
          <fpage>361</fpage>
          -
          <lpage>366</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>