<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Adversarial Learning for Visual Tracking Research Idea</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>l Di N</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Milan</institution>
          ,
          <addr-line>Milan MI 20122</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>101</fpage>
      <lpage>106</lpage>
      <abstract>
        <p>The doctoral research activity1 mainly focuses on methodologies in the field of computer vision. In particular, the work is focused on designing, developing and validating novel approaches, also based on deep learning methodologies, for visual tracking. Visual tracking in video sequences has always been a main topic in computer vision and interesting results have been obtained by approaches based on Support Vector Machine, Siamese Networks and Discrete Correlation Filters. However, these techniques are limited due to the low discriminative ability of the used features for object detection. In his research activities, Emanuel Di Nardo proposes a novel approach, based on Generative Adversarial Networks for feature extraction or regression. In particular, using Generative Adversarial Networks we are able to characterize the elements to be traced in the scene and make them easier to recognize.</p>
      </abstract>
      <kwd-group>
        <kwd>Deep Learning</kwd>
        <kwd>Adversarial Learning</kwd>
        <kwd>Feature Extraction</kwd>
        <kwd>Visual Tracking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Visual tracking in video sequences has always been a topic that arouses the
attention of the scientific community. It consists in detecting and following an
object that moves in a scene. The object will inevitably undergo modifications
during its movement (Fig. 1). It happens both because of the displacement itself
being free of constraints and because the scene itself, which could present
obstacles between the camera and the object and still due to problems caused by the
acquisition conditions as in the case of non-ideal lighting. Therefore it is needed
to use techniques that are defined as robust with a fair compromise of accuracy.</p>
      <p>
        There are many challenges for this kind of task and one is VOT (Visual Object
Tracking) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Usually, in this context the following three parameters are taken
into account:
1. Accuracy. Mean Overlap between the target and the ground truth
2. Robustness. How many times the target is lost
3. Expected Average Overlap (EAO). Mean of accuracy over multiple
video sequences with the same visual properties. It combines accuracy and
robustness
Copyright c 2019 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
      </p>
      <p>In the VOT challenge, it is possible to identify two kinds of tracking called
short-term and long-term. In the former, an object is always visible in the scene
and it is possible to detect it in each frame. In the latter, the target could not be
present in the scene for long time due to a total occlusion because it came out
of the scene. In this case, the tracker can not report any position for the object,
but it can provide a confidence score that the object is not present.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Most of the works in visual tracking are compared on various challenges such as
VOT [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and MOT (Multiple Object Tracking) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. On the one hand, a strong
evolution of techniques is based on the template matching of the whole object
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. On the other hand, other approaches tend to take into account
the movement and to estimate the possible position in which the target is
located, relying for example on the optical flow [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Other methodologies
called part-based tend to scan the areas close to the initial target by estimating
which points have the greatest probability that there is a target or a part of
it [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Nowadays, most trackers use approaches based on artificial Neural
Networks (NN) at various stages of the tracking process. Some use them to have
meaningful features that can be representative of the object [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In recent
years, moreover, techniques based on Siamese networks have emerged, which see
two parallel networks that work together to estimate the location of the object
in the scene [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Other techniques use filters that allow, through a
domain transformation, to discriminate the probable position in a robust and a
highly-efficient way [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Some methodologies mix all these approaches
together to be more and more precise [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The approaches based on Deep
Learning [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] use a pre-trained Neural Network on known datasets
for classifying objects in the images. A recent technique is VITAL [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. It uses
1 Ph.D. supervisor: Angelo Ciaramella (University of Naples Parthenope);
cosupervisor: Fabio Narducci (University of Naples Parthenope).
a Generative Adversarial Networks to generate a mask that represents the most
relevant features in the image, based on the input target. Another recent
approach uses compressive sensing for trajectory tracking [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] in order to reduce
the image complexity and be able to know where the target is moving on.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Research Idea</title>
      <p>
        The main objective of the research activity is the introduction of a novel approach
based on Generative Adversarial Networks (GANs) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for visual tracking. In a
first scenario a GAN is used for tracking. Usually, the adversarial networks are
used to generate samples that are as close as possible to the real ones. This is
possible thanks to the property they have to learn the distribution of the data
they want to generate. This type of approach leads to the generation of a latent
space for the representation of the data. Here, the idea is to use this property
for extracting representative samples (i.e., features). Some studies reported this
approach combined with autoencoders [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. In particular, the generator
network encodes the vector representation extracted form the latent space. In a
second scenario GANs could be used directly to perform a regression operation
[
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] to define where an object is found, or at least, as a support to the
localization through probable positions. Differently from what has already been
done [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], in the proposed approach GAN should be able to extract a meaningful
representation of the object with a concrete reconstruction of the target
localization in the image instead of a simple dropout mask that suggests what are
the areas that are more sensible to the input. Furthermore, we want the ability
of domain adaptation of GANs to be discriminating for this activity without
recurring, as it happens in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] to pre-training on the task of tracking.
Another aspect that should not be underestimated is that of scene and the
object regularizing. In this context, GANs could be used to remove alterations in
images to make tracking easier or even generating objects in positions different
from those known to better estimate how an object can be changed over time.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Planning</title>
      <p>Phase One A generative network is built to perform and study the
segmentation of the image trying to localize a target object in an image. The investigation
aims to generate an image that is related to the ground truth used in the
discriminator network. As the first point, the activity tries to relate an image and a
target that is visible in it. The first experiments are conducted using the model
in Fig. 2 a generative model with two inputs (the image and the target) that
are encoded individually. In the end, they are concatenated on the feature
dimension to achieve an association between them. Further methodologies are in
development to enforce the relationship between the image and the target. In
this context, it has to face some problems. The first one is the multi-domain
property that the network tries to approximate because it is not trained on a set
of objects that belong to the same category, but on a large variety of them. On
the other hand, the segmentation purpose should help to normalize this behavior
because the objective function is calibrated to work on a less complex solution.
Another problem is related to the segmentation quality. It is possible that the
result is not accurate with a degeneration to a kind of output that can be more
similar to a heat-map.</p>
      <p>Phase Two The segmentation obtained from the first step can help to
understand what is the discriminatory effect of the learned latent space. It can be used
trying to work only on the encoding of the input without bringing it to the
generative output in a pure autoencoder fashion. The main challenge is on the usage
of only encoded features because space on which it is mapped could be lost some
important properties if partially described. Another important investigation can
be done on the discriminator side of the GAN. Usually, it is used only in the
generative learning step and not in the operative phase, but it learns to extract
characteristic features of real data. It is an important property that could be
used in a verification step to avoid low-quality output or in a distraction-aware
manner.</p>
      <p>
        Future proposals In addition to GAN based plan, other research paths can
be invested in future investigation. It involves analyzing the techniques based on
dictionary learning, compressive sensing [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and ODE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that appear to be a
valid alternative to the classical operations found in all tracking algorithms.
      </p>
      <p>The shown planning would achieve a strong representation of a generic object
and, consequently, learning also the localization in the space. It can be used as
tracking-by-detection solution in a short-term challenge. It gives the possibility
to join the VOT challenge to evaluate the research work using the benchmark
tools provided in the competition and validate the effectiveness of the developed
model.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>Aim of this study is the introduction of innovative methodologies for visual
tracking in the field of computer vision. In particular, it could be noted that
new elements that are currently used in different fields with excellent results
could give a strong boost to the current research status.</p>
      <p>
        These should be validated comparing with proposed state-of-the-art
methodologies to understand how much they characterize or not. From here, it is possible
to analyze how to use the whole adversarial model as a detector of objects
without going through other methods. This study could also lead to using different
characterization techniques from convolutional networks. In particular, in [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]
the proposed approach deviates from the classic CNN and that appears to be a
good alternative to prevent the number of parameters from exploding.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ian</surname>
          </string-name>
          , et al.
          <article-title>"Generative adversarial nets</article-title>
          .
          <source>" Advances in neural information processing systems</source>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Candes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Romberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <article-title>"Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information,"</article-title>
          <source>in IEEE Transactions on Information Theory</source>
          , vol.
          <volume>52</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>489</fpage>
          -
          <lpage>509</lpage>
          , Feb.
          <year>2006</year>
          . doi:
          <volume>10</volume>
          .1109/TIT.
          <year>2005</year>
          .862083
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ricky</surname>
            <given-names>T. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            and
            <given-names>Yulia</given-names>
          </string-name>
          <string-name>
            <surname>Rubanova</surname>
            and
            <given-names>Jesse</given-names>
          </string-name>
          <string-name>
            <surname>Bettencourt</surname>
            and
            <given-names>David</given-names>
          </string-name>
          <string-name>
            <surname>Duvenaud</surname>
          </string-name>
          .
          <source>Neural Ordinary Differential Equations</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cortes</surname>
            , Corinna, and
            <given-names>Vladimir</given-names>
          </string-name>
          <string-name>
            <surname>Vapnik</surname>
          </string-name>
          .
          <article-title>"Support-vector networks</article-title>
          .
          <source>" Machine learning 20.3</source>
          (
          <year>1995</year>
          ):
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bertinetto</surname>
          </string-name>
          ,
          <string-name>
            <surname>Luca</surname>
          </string-name>
          , et al.
          <article-title>"Fully-convolutional siamese networks for object tracking</article-title>
          .
          <source>" European conference on computer vision</source>
          . Springer, Cham,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lukezic</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alan</surname>
          </string-name>
          , et al.
          <article-title>"Discriminative correlation filter with channel and spatial reliability</article-title>
          .
          <source>" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kristan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Matej</surname>
          </string-name>
          , et al.
          <article-title>"The sixth visual object tracking vot2018 challenge results</article-title>
          .
          <source>" Proceedings of the European Conference on Computer Vision (ECCV)</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dendorfer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Patrick</surname>
          </string-name>
          , et al.
          <article-title>"CVPR19 Tracking and Detection Challenge: How crowded can it get?." arXiv preprint arXiv:</article-title>
          <year>1906</year>
          .
          <volume>04567</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>An adaptive combination of multiple features for robust tracking in real scene</article-title>
          .
          <source>In: 2013 IEEE International Conference on Computer Vision Workshops (ICCVW)</source>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>136</lpage>
          ,
          <year>December 2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xing</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Graph embedding based semi-supervised discriminative tracker</article-title>
          .
          <source>In: 2013 IEEE International Conference on Computer Vision Workshops (ICCVW)</source>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>152</lpage>
          ,
          <year>December 2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Maresca</surname>
            , Mario Edoardo, and
            <given-names>Alfredo</given-names>
          </string-name>
          <string-name>
            <surname>Petrosino</surname>
          </string-name>
          .
          <article-title>"Matrioska: A multi-level approach to fast tracking by learning</article-title>
          .
          <source>" International Conference on Image Analysis and Processing</source>
          . Springer, Berlin, Heidelberg,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Wendel</surname>
            , Andreas,
            <given-names>Sabine</given-names>
          </string-name>
          <string-name>
            <surname>Sternig</surname>
            , and
            <given-names>Martin</given-names>
          </string-name>
          <string-name>
            <surname>Godec</surname>
          </string-name>
          .
          <article-title>"Robustifying the flock of trackers." 16th Computer Vision Winter Workshop</article-title>
          . Citeseer.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Maresca</surname>
            , Mario Edoardo, and
            <given-names>Alfredo</given-names>
          </string-name>
          <string-name>
            <surname>Petrosino</surname>
          </string-name>
          .
          <article-title>"Clustering local motion estimates for robust and efficient object tracking</article-title>
          .
          <source>" European Conference on Computer Vision</source>
          . Springer, Cham,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lukežič</surname>
            , Alan, Luka Čehovin Zajc, and
            <given-names>Matej</given-names>
          </string-name>
          <string-name>
            <surname>Kristan</surname>
          </string-name>
          .
          <article-title>"Deformable parts correlation filters for robust visual tracking."</article-title>
          <source>IEEE transactions on cybernetics 48</source>
          .6 (
          <year>2017</year>
          ):
          <fpage>1849</fpage>
          -
          <lpage>1861</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Battistone</surname>
            , Francesco,
            <given-names>Alfredo</given-names>
          </string-name>
          <string-name>
            <surname>Petrosino</surname>
            , and
            <given-names>Vincenzo</given-names>
          </string-name>
          <string-name>
            <surname>Santopietro</surname>
          </string-name>
          .
          <article-title>"Watch out: Embedded video tracking with BST for unmanned aerial vehicles</article-title>
          .
          <source>" Journal of Signal Processing Systems 90.6</source>
          (
          <year>2018</year>
          ):
          <fpage>891</fpage>
          -
          <lpage>900</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Chong</surname>
          </string-name>
          , et al.
          <article-title>"Learning spatial-aware regressions for visual tracking</article-title>
          .
          <source>" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Bo</surname>
          </string-name>
          , et al.
          <article-title>"High performance visual tracking with siamese region proposal network."</article-title>
          <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Li</surname>
            , Yuhong,
            <given-names>and Xiaofan</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          .
          <article-title>"SiamVGG: Visual Tracking using Deeper Siamese Networks." arXiv preprint arXiv:</article-title>
          <year>1902</year>
          .
          <volume>02804</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kiani</surname>
            <given-names>Galoogahi</given-names>
          </string-name>
          , Hamed,
          <string-name>
            <given-names>Ashton</given-names>
            <surname>Fagg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Lucey</surname>
          </string-name>
          .
          <article-title>"Learning backgroundaware correlation filters for visual tracking</article-title>
          .
          <source>" Proceedings of the IEEE International Conference on Computer Vision</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Lukezic</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alan</surname>
          </string-name>
          , et al.
          <article-title>"Discriminative correlation filter with channel and spatial reliability</article-title>
          .
          <source>" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
          </string-name>
          , et al.
          <article-title>"Learning spatial-temporal regularized correlation filters for visual tracking</article-title>
          .
          <source>" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Valmadre</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jack</surname>
          </string-name>
          , et al.
          <article-title>"End-to-end representation learning for correlation filter based tracking</article-title>
          .
          <source>" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Yun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sangdoo</surname>
          </string-name>
          , et al.
          <article-title>"Action-decision networks for visual tracking with deep reinforcement learning</article-title>
          .
          <source>" Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Song</surname>
          </string-name>
          , Yibing et al. “
          <article-title>VITAL: VIsual Tracking via Adversarial Learning</article-title>
          .”
          <source>2018 IEEE/CVF Conference on Computer Vision</source>
          and Pattern
          <string-name>
            <surname>Recognition</surname>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>: n. pag</article-title>
          . Crossref. Web.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Nam</surname>
          </string-name>
          , Hyeonseob, and Bohyung Han.
          <article-title>“Learning Multi-Domain Convolutional Neural Networks for Visual Tracking</article-title>
          .”
          <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2016</year>
          )
          <article-title>: n. pag</article-title>
          . Crossref. Web.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Kracunov</surname>
            , Marijana,
            <given-names>Milica</given-names>
          </string-name>
          <string-name>
            <surname>Bastrica</surname>
          </string-name>
          , and Jovana Tesovic. “
          <article-title>Object Tracking in Video Signals Using Compressive Sensing</article-title>
          .”
          <source>2019 8th Mediterranean Conference on Embedded Computing (MECO)</source>
          (
          <year>2019</year>
          )
          <article-title>: n. pag</article-title>
          . Crossref. Web.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Pinho</surname>
            , Eduardo, and
            <given-names>Carlos</given-names>
          </string-name>
          <string-name>
            <surname>Costa</surname>
          </string-name>
          .
          <article-title>"Feature Learning with Adversarial Networks for Concept Detection in Medical Images: UA</article-title>
          . PT Bioinformatics at ImageCLEF
          <year>2018</year>
          .
          <article-title>"</article-title>
          <source>CLEF (Working Notes)</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Sohn</surname>
            , Kihyuk,
            <given-names>Honglak</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>and Xinchen</given-names>
          </string-name>
          <string-name>
            <surname>Yan</surname>
          </string-name>
          .
          <article-title>"Learning structured output representation using deep conditional generative models</article-title>
          .
          <source>" Advances in neural information processing systems</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Olmschenk</surname>
            , Greg,
            <given-names>Zhigang</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
            , and
            <given-names>Hao</given-names>
          </string-name>
          <string-name>
            <surname>Tang</surname>
          </string-name>
          . “
          <article-title>Generalizing Semi-Supervised Generative Adversarial Networks to Regression Using Feature Contrasting</article-title>
          .”
          <source>Computer Vision and Image Understanding</source>
          <volume>186</volume>
          (
          <year>2019</year>
          ):
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . Crossref. Web.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <given-names>Karan</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Kirchmeyer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pranjul</given-names>
            <surname>Yadav</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. Sathiya</given-names>
            <surname>Keerthi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Gallinari</surname>
          </string-name>
          .
          <source>Regression with Conditional GAN</source>
          .
          <year>2019</year>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Ullah</surname>
            , Ihsan, and
            <given-names>Alfredo</given-names>
          </string-name>
          <string-name>
            <surname>Petrosino</surname>
          </string-name>
          .
          <article-title>"A strict pyramidal deep neural network for action recognition</article-title>
          .
          <source>" International Conference on Image Analysis and Processing</source>
          . Springer, Cham,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>