<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The neural network image captioning model based on adversarial training</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>K P Korshunova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The Branch of National Research University "Moscow Power Engineering Institute" in Smolensk</institution>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>438</fpage>
      <lpage>444</lpage>
      <abstract>
        <p>The paper represents the model for image captioning based on deep neural networks and adversarial training process. The model consists of a convolutional network as image encoder, a recurrent network as natural language generator and another convolutional network as an adversarial discriminator. The structure of the model, the training algorithm, some experimental results and evaluation using popular metrics are proposed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Nowadays complex artificial intelligence tasks that require processing of combination of visual and
linguistic information has received increasing attention from both the computer vision and natural
language processing communities. These tasks are called multimodal. They are challenging because of
requiring accurate computational visual recognition, comprehensive world knowledge, and natural
language generation. In addition to computer vision and natural language processing problems there
are some problems related to the combination of the fields. One of the most challenging tasks is
automatic Image Captioning [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] known from 1990s [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Image Captioning Task</title>
      <p>Automatic Image Captioning systems generate one or more descriptive sentences in natural language
given a sample image.</p>
      <p>
        The task is the intersection of two data analysis fields: pattern recognition and natural language
processing. In addition to visual objects, attributes and relations recognizing it requires further
describing them as a natural language text [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The task of generating image descriptions can be understood as translation from one representation
(visual features) to another (text features). In this aspect it is similar to machine translation task that is
to transform data representation written in one language/modality (an input image I) into its
representation in the target language/modality (a target sequence of words C) by maximizing the
likelihood p(C|I) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>Automatic Image Captioning systems include two subsystems: “encoder” and “decoder”. An
“encoder” reads the source data (raw pixels of the given image) and transforms it into a rich
fixedlength vector representation, which in turn is used as the initial hidden state of a “decoder” that
generates the target descriptive sentence in natural language.</p>
      <p>
        The most successful Image Captioning approaches are based on deep neural networks:
Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN). General Image
Captioning approach (Figure 1): convolutional neural network (first pre-trained for an image
classification task) is used as an image “encoder”, then the last hidden layer is used as an input to the
RNN decoder that generates sentences [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The neural network image captioning model based on adversarial training</title>
      <p>
        Generative Adversarial Nets (GANs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) that implement adversarial training have been used to
produce samples of photorealistic images, to model patterns of motion in video, to reconstruct 3D
models of objects from images, to improve astronomical images, etc. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>
        However in this paper we propose image captioning approach based on the Sequence Generative
Adversarial Nets (Sequence GANs [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]).
      </p>
      <p>GANs represent a combination of two neural network: one network (generative model G) generates
candidates and the other (discriminative model D) evaluates them. Typically, the generator G learns to
map from a latent space to a particular data distribution of interest, while the discriminator D
discriminates between instances from the true data distribution and candidates produced by the
generator. This is the implementation of adversarial training: the generative model’s training objective
is to increase the error rate of the discriminative model (i.e., "fool" the discriminator network by
producing novel synthesised instances that appear to have come from the true data distribution).
3.1. The structure of the model
The general structure of the proposed neural network model is represented in the Figure 2.</p>
      <p>The model consists of:
1) convolutional neural network that is used as an image “encoder”;
2) recurrent network that produces natural language descriptions;
3) another convolutional neural network that is used as the discriminator during adversarial
training process.
as image encoder, recurrent network as natural language generator and convolutional network as a
adversarial discriminator.</p>
      <p>
        VGG16 model [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is used for image encoding (CNN), LSTM (Long-Short Term Memory [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ])
recurrent network is used for generating text descriptions (G). We choose the convolutional network
as the discriminator (D) as this kind of deep networks have recently been shown of great effectiveness
in text (token sequence) classification [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>3.2. Training algorithm</title>
        <p>The training process of the proposed model consists of the following steps:</p>
        <p>Step 1. Initialization and pre-training:
1.1. Pre-train CNN and G;
1.2. Generate negative samples using CNN and G;
1.3. Pre-train D;
Step 2. Training (N epochs):
2.1. Train G for g epochs;
2.2. Generate negative samples using CNN and G;
2.3. Train D for d epochs.</p>
        <p>
          We use the reinforcement learning (RL) modification [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] to train the proposed model. The
generative model is treated as an agent of RL. In the case of adversarial training the discriminative net
D learns to distinguish whether a given data instance is real or not, and the generative net G learns to
confuse D by generating high quality data.
        </p>
        <p>The discriminator provides the adversariness of the training process. The CNN and the generator G
work during production of the model: raw pixels of the given image are read and transformed into a
rich fixed-length vector representation by the encoder CNN, then generator G generates the target
descriptive sentence in natural language from this representation.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Experiments results</title>
        <p>
          We have performed some experiments on the challenging public available Microsoft COCO Caption
dataset [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It includes images from Microsoft Common Objects in COntext (COCO) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] database.
All data are divided into training set and validation set. We use 32,000 images and 160,000
corresponding text descriptions (five per image) as training set and 40,000 pairs “image-sentence” as
validation set.
        </p>
        <p>Several sample descriptions provided by the model after 75 training epochs are represented in the
Figure 3.</p>
        <p>In many cases descriptions made by the proposed model can describe the content of the depicted
scenes (despite grammatical and semantic inaccuracies). However there are some gross mistakes.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Evaluation</title>
        <p>
          Although it is sometimes not clear whether a description should be deemed successful or not given an
image, prior art has proposed several evaluation metrics [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. These metrics are
based on evaluating the similarity of two sentences (candidate caption and reference caption). We use
popular metrics BLEU-1, BLEU-2, BLEU-3, BLEU-4 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], ROUGE-L [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], CIDEr [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>We compare the proposed neural network model based on adversarial training to an CNN+RNN
baseline.</p>
        <p>The image captioning performance of the proposed (GAN) and known (CNN+RNN) models are
represented in the Table 1 and Figures 4-5.</p>
        <p>Table 1 and Figures 4-5 show that the proposed image captioning model based on adversarial
training outperforms the compared baseline (CNN+RNN) in various metrics. The best improvement is
achieved for 100 training epochs. Obviously, the performance of the proposed model depends on the
detailed model structure and training strategy. Choosing the attributes of the model structure (number
of layers, etc.) and values of the training process parameters (number of training and pre-training
epochs) is the problem for further research.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In this paper, we proposed a neural network image captioning model based on adversarial training.
The model combines a convolutional neural net for image processing and Sequence Generative
Adversarial Net for generating text descriptions. Some experimental work to measure the effectiveness
of the model has been performed on the challenging Microsoft COCO Caption dataset. It shows that
the proposed model could provide better automatic Image Captioning compared to known baseline of
CNN and RNN.</p>
      <p>Acknowledgments
This work was supported by the Russian Foundation for Basic Research (Grant No. 18-07-00928)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Borisov</surname>
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korshunova</surname>
            <given-names>K. P.</given-names>
          </string-name>
          <string-name>
            <surname>Direct</surname>
          </string-name>
          and
          <article-title>Reverse Image Captioning problem definition. Postanovka priamoi i obratnoi zadachi poiska i generirovaniia tekstovykh opisanii po izobrazheniiam</article-title>
          . Energetika, informatika, innovatsii
          <article-title>- 2017 (elektroenergetika, elektrotekhnika i teploenergetika, matematicheskoe modelirovanie i informatsionnye tekhnologii v proizvodstve). [Power engineering</article-title>
          , computer science, innovations - 2017
          <source>. Proceedings of the VII international scientific conference]. Smolensk</source>
          ,
          <year>2017</year>
          , pp
          <fpage>228</fpage>
          -
          <lpage>230</lpage>
          (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Korshunova</surname>
            <given-names>K. P.</given-names>
          </string-name>
          <string-name>
            <surname>Automatic Image</surname>
          </string-name>
          <article-title>Captioning: Tasks and Methods</article-title>
          .
          <source>Systems of Control, Communication and Security</source>
          ,
          <year>2018</year>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>77</lpage>
          . Available at: http://sccs.intelgr.com/archive/2018-01/
          <fpage>02</fpage>
          -Korshunova.pdf (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Abella</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kender</surname>
            <given-names>J. R.</given-names>
          </string-name>
          ,
          <source>Starren J. Description Generation of Abnormal Densities found in Radiographs // Proc. Symp</source>
          . Computer Applications in Medical Care,
          <source>Journal of the American Medical Informatics Association</source>
          .
          <year>1995</year>
          . pp
          <fpage>542</fpage>
          -
          <lpage>546</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Anderson</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernando</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gould</surname>
            <given-names>S. SPICE</given-names>
          </string-name>
          :
          <article-title>Semantic propositional image caption evaluation</article-title>
          <source>// Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <fpage>9909</fpage>
          <lpage>LNCS</lpage>
          .
          <year>2016</year>
          . pp
          <fpage>382</fpage>
          -
          <lpage>398</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Chen</surname>
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick</surname>
            <given-names>C. L.</given-names>
          </string-name>
          <article-title>Mind's eye: A recurrent visual representation for image caption generation //</article-title>
          <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2015</year>
          . pp
          <fpage>2422</fpage>
          -
          <lpage>2431</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Chen</surname>
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            <given-names>T. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vedantam</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollár</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick C. L. Microsoft COCO</surname>
          </string-name>
          <article-title>Captions: Data Collection and Evaluation Server</article-title>
          . arXiv.org,
          <year>2015</year>
          . Available at: https://arxiv.org/abs/1504.00325 (
          <issue>accessed</issue>
          : 01
          <source>February</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Denkowski</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavie</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Meteor</surname>
          </string-name>
          <article-title>Universal: Language Specific Translation Evaluation for Any Target Language //</article-title>
          <source>Proceedings of the Ninth Workshop on Statistical Machine Translation</source>
          .
          <year>2014</year>
          . pp
          <fpage>376</fpage>
          -
          <lpage>380</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Gerber</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagel</surname>
            <given-names>N. H.</given-names>
          </string-name>
          <article-title>Knowledge representation for the generation of quantified natural language descriptions of vehicle traffic in image sequences //</article-title>
          <source>Proceedings of the International Conference on Image Processing</source>
          .
          <year>1996</year>
          . pp
          <fpage>805</fpage>
          -
          <lpage>808</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Goodfellow</surname>
            <given-names>I. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouget-Abadie</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirza</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warde-Farley</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ozair</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y</given-names>
          </string-name>
          . Generative Adversarial Networks // Proceedings of NIPS.
          <year>2014</year>
          . pp
          <fpage>2672</fpage>
          -
          <lpage>2680</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Gu</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Stack-Captioning</surname>
          </string-name>
          :
          <article-title>Coarse-to-Fine Learning for Image Captioning // Association for the</article-title>
          <source>Advancement of Artificial Intelligence</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hochreiter</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urgen Schmidhuber J. Long</surname>
            <given-names>Short-Term</given-names>
          </string-name>
          <string-name>
            <surname>Memory</surname>
          </string-name>
          .
          <source>Neural Computation</source>
          ,
          <year>1997</year>
          , Vol.
          <volume>9</volume>
          (
          <issue>8</issue>
          ), pp
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Karpathy</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei L</surname>
          </string-name>
          .
          <article-title>Deep visual-semantic alignments for generating image descriptions //</article-title>
          <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Convolutional Neural Networks for Sentence Classification /</article-title>
          / EMNLP.
          <year>2014</year>
          . pp
          <fpage>1746</fpage>
          -
          <lpage>1751</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Lantao</surname>
            <given-names>Yu</given-names>
          </string-name>
          , Weinan Zhang, JunWang, Y. Y. SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient // JAMA Internal Medicine,
          <volume>177</volume>
          (
          <issue>3</issue>
          ).
          <year>2017</year>
          . pp
          <fpage>326</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Lin</surname>
            <given-names>C. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Rouge</surname>
          </string-name>
          :
          <article-title>A package for automatic evaluation of summaries //</article-title>
          <source>Proceedings of the Workshop on Text Summarization Branches out (WAS</source>
          <year>2004</year>
          ).
          <year>2004</year>
          . Vol.
          <volume>1</volume>
          . pp
          <fpage>25</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Lin</surname>
            <given-names>T. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maire</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perona</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramanan</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollár</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick C. L. Microsoft</surname>
            <given-names>COCO</given-names>
          </string-name>
          :
          <article-title>Common objects in context</article-title>
          .
          <source>European conference on computer vision</source>
          ,
          <year>2014</year>
          ,
          <fpage>pp740</fpage>
          -
          <lpage>755</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Papineni</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roukos</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            <given-names>W.</given-names>
          </string-name>
          <article-title>BLEU: a method for automatic evaluation of machine translation</article-title>
          .
          <source>Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2002</year>
          , pp
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Robertson</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>Understanding inverse document frequency: on theoretical arguments for IDF /</article-title>
          / Journal of Documentation.
          <year>2004</year>
          . №
          <volume>60</volume>
          (
          <issue>5</issue>
          ). pp
          <fpage>503</fpage>
          -
          <lpage>520</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Simonyan</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Very deep convolutional networks for large-scale image recognition // arXiv</article-title>
          .org.
          <year>2014</year>
          . - URL: https://arxiv.org/abs/1409.1556 (
          <issue>accessed</issue>
          : 01
          <source>February</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Sutton</surname>
            <given-names>R. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barto</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Reinforcement learning: an introduction</article-title>
          . University College London, Computer Science Department, Reinforcement Learning Lectures,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Vedantam</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick</surname>
            <given-names>C. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parikh</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>CIDEr: Consensus-based image description evaluation //</article-title>
          <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2015</year>
          . pp
          <fpage>4566</fpage>
          -
          <lpage>4575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Vinyals</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toshev</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Show</surname>
          </string-name>
          and
          <string-name>
            <surname>Tell: A Neural Image Caption Generator</surname>
          </string-name>
          // Conference on Computer Vision and Pattern Recognition.
          <year>2015</year>
          . pp
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23] Generative Adversarial Network // Wikipedia, the free encyclopedia.
          <year>2018</year>
          . URL: https://en.wikipedia.org/wiki/Generative_adversarial_
          <source>network (accessed: 01 May</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Xu</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiros</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemel</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Show</surname>
          </string-name>
          ,
          <source>Attend and Tell: Neural Image Caption Generation with Visual Attention // International Conference on Machine Learning</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>