<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigation of the Batch Size Influence on the Quality of Text Generation by the SeqGAN Neural Network</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nikolay Krivosheev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ksenia Vik</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yulia Ivanova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Spitsyn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Tomsk Polytechnic University</institution>
          ,
          <addr-line>Lenin Avenue, 30, Tomsk, 634050</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tomsk State University of Architecture and Building</institution>
          ,
          <addr-line>Solyanaya square, 2, Tomsk, 634003</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One of the problems of text generation using the LSTM neural network is a decrease in the quality of generation with an increase in the length of the generated text. There are various solutions to improve the quality of text generation based on generative adversarial neural networks. This work uses preliminary training of the LSTM neural network based on the MLE approach and further training based on the SeqGAN neural network. Based on the presented results, we can conclude that the SeqGAN-based approach allows to increase the quality of text generation according to the NLL and BLEU metrics. The study of the influence of the batch size, in the process of competitive training of the SeqGAN neural network, on the quality of text generation has been carried out. It is shown that with an increase in the batch size, in the process of adversarial learning, the quality of LSTM neural network training increases. In this work, the Monte Carlo algorithm is not used in the training process of the SeqGAN neural network. For training and testing algorithms, image captions from the COCO Image Captions data sample are used. The quality of text generation based on the NLL and BLEU metrics has been assessed. Examples of the results of generating texts with an assessment of the quality of examples according to the BLEU metric are given.</p>
      </abstract>
      <kwd-group>
        <kwd>1 SeqGAN</kwd>
        <kwd>text generation</kwd>
        <kwd>adversarial learning</kwd>
        <kwd>reinforcement learning</kwd>
        <kwd>batch size</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Algorithms for automatic text generation are in demand and are widely used in various tasks, such
as: automatic generation of random texts, automatic translation of texts, text summarization, generation
of captions for images.</p>
      <p>
        One of the problems of text generation using the LSTM [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] neural network is a decrease in the
quality of generation with an increase in the length of the generated text. There are various solutions to
improve the quality of text generation based on generative adversarial neural networks. In this work,
we use preliminary training of the LSTM neural network based on the MLE [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] approach and further
training based on the SeqGAN [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] neural network. The training process based on the SeqGAN neural
network does not use the Monte Carlo algorithm proposed in paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The study of the influence of
the batch size on the quality of LSTM neural network training in the process of adversarial training
based on the SeqGAN neural network was carried out. It is shown that with an increase in the batch
size, in the process of adversarial learning, the quality of LSTM neural network training increases. Also,
an increase in the batch size leads to an increase in the training time of the algorithm.
      </p>
      <p>
        For training and testing algorithms, captions to images from the COCO Image Captions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] data
sample are used. The proposed sample contains captions in English. In this work, word-by-word text
generation is used. The maximum sentence length is 20 words; most sentences in this set contain about
10 words. COCO Image Captions sample data is taken from analogue [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The data is posted on the
website [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Thus, the preprocessing of the data selection coincides with that implemented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The testing process uses a pretrained neural network based on the MLE approach. In the MLE-based
pre-learning process, the batch size is 1,000 examples and does not change. Next, we tested the LSTM
neural network trained on the basis of the SeqGAN approach using various batch sizes. Tested on a
batch size of 40, 400 and 4000 samples.</p>
      <p>
        In this work, the software implementation of the investigated approaches in the Python language
was carried out. The software implementation is presented on the site [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and is based on the work [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Description of the used data sample</title>
      <p>In this work, we used the COCO Image Captions (Common Objects in Context) dataset. This set
consists of 330000 images, of which 220000 samples are marked. All images are accompanied by
annotations stored in json format. This set has various types of annotation, such as:
 object detection;
 keypoint detection;
 stuff segmentation;
 panoptic segmentation;
 denepose;
 image captioning.</p>
      <p>
        To form a textual data sample, captions to images in English from the COCO Image Captions sample
are used. The maximum length for an example is 20 words. Most of the sample texts contain about 10
words. The data sample was taken from the source [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], so its preprocessing corresponds to that presented
in article [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Words mentioned less than 10 times have been removed from the sample, examples
containing these words have also been removed. The number of words used in the dictionary is 4837
words. This dataset contains 80000 examples in the training set and 5000 examples in the test set.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Description of the influence of batch size on the SeqGAN learning process</title>
      <p>
        In this work, we use the SeqGAN neural network proposed in paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our work does not use the
Monte Carlo algorithm proposed in work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. One of the hyperparameters affecting the training process
of the SeqGAN neural network is the batch size. This parameter affects the number of examples
generated by the neural network by the generator and further evaluated by the neural network by the
discriminator. An example of text evaluation by a neural network by a discriminator is shown in the
image Figure 1.
      </p>
      <p>From the presented image, you can see that words used several times in different sentences receive
a better and more average grade. In the process of generating text, the neural network generator can
generate a good beginning of a sentence, but the end of a sentence may be of poor quality. As a result,
the discriminator neural network will give a poor estimate of the entire generated sequence. Thus,
increasing the number of examples in batch size allows for a better estimate of the generator neural
network.</p>
      <p>
        In paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], it is proposed to use the Monte Carlo algorithm to assess the quality of the generated
samples. This approach allows you to increase the quality of the text evaluation, as well as the use of
an increased batch size.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Assessment of the quality of text generation according to the BLEU metric</title>
      <p>
        In this work, the text quality is assessed based on the BLEU [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] metric implemented in the nltk
library.
      </p>
      <p>Before evaluating sample texts, the placeholder word is removed. This word is used to increase the
length of the text to a given length during training. This operation allows you to get a better assessment
of texts, since the matches of placeholder words are not taken into account when evaluating samples.</p>
      <p>
        In this work, when evaluating samples, the smoothing method is used. This smoothing method is
described in paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The proposed modification allows you to get a better assessment of the quality
of texts.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Test results of the considered approaches</title>
      <p>
        The neural network SeqGAN was developed and tested on the task of word-by-word generation of
short texts [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Implementation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and testing of the SeqGAN neural network were carried out on a
data sample with captions to images from the COCO Image Captions sample [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>To assess the quality of text generation according to the BLEU metric, samples from the training
sample (written by people) were tested on a test sample. The results obtained are used for subsequent
comparison with the generation results. The results of testing 500 random examples from the training
sample are presented in Table 1.</p>
      <p>
        The quality of the generation of texts generated by the LSTM neural network trained on the basis of
the MLE approach was assessed. The quality was assessed based on the BLEU [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and negative
loglikelihood (NLL) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] metrics. The test results are presented in Table 2:
Table 2
MLE-trained LSTM test results
      </p>
      <p>The generator neural network in SeqGAN is an LSTM trained using MLE. Table 3 shows the results
of additional LSTM training using the SeqGAN neural network. Lot sizes of 40, 400 and 4000 samples
were chosen. This choice is based on the formation of a larger difference in the size of batches, 10 and
100 times.</p>
      <p>
        Examples of texts generated using a neural network trained by SeqGAN are presented in Table 4.
For each example, an estimate is given according to the BLEU metric [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Based on the graphs and the results presented in the tables above, we can conclude that the batch
size affects the final quality of the neural network training according to the NLL metric. It should also
be emphasized that in this work, in the process of training the SeqGAN neural network, the Monte
Carlo-based algorithm used in the original work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is not used. The Monte Carlo algorithm can improve
the quality of training a neural network, but is not considered in this paper.
      </p>
      <p>Using the SeqGAN neural network improves the quality of text generation using the BLEU and NLL
metrics. The improvement in text quality can be seen by comparing the learning outcomes based on
MLE and SeqGAN. Increasing the batch size reduces the error in the NLL metric. The disadvantage of
this approach is the increase in the training time of the algorithm.</p>
      <p>When using the batch size parameter equal to 40, the text quality according to the BLEU metric is
better than when using the batch size equal to 400 and 4000. An increase in the quality of the text may
be associated with a decrease in its diversity, as indicated by a large error in the NLL metric. This can
also be evidenced by the higher accuracy of the discriminator, equal to 86%, while when using packets
400 and 4000 it is 81%. It should be noted that the quality of text generation by the neural network
trained on the basis of the MLE and SeqGAN approach is inferior to the examples from the training
sample according to the BLEU metric.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>
        As part of this work, the MLE and SeqGAN approaches were implemented and tested. Based on the
presented results, we can conclude that the SeqGAN-based approach allows to increase the quality of
text generation according to the NLL and BLEU metrics. In this work, the Monte Carlo algorithm
proposed in paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] was not used.
      </p>
      <p>The LSTM neural network was trained based on the SeqGAN approach using different batch sizes.
Based on the results obtained, it was concluded that the increased batch size makes it possible to
increase the training quality of the LSTM neural network.</p>
      <p>It should be noted that the quality of text generation by the neural network trained on the basis of
the MLE and SeqGAN approach is inferior to the examples from the training sample according to the
BLEU metric.</p>
      <p>
        In this work, the software implementation of the investigated approaches in the Python language
was carried out. The software implementation is presented on the site [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and is based on the work [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgements</title>
      <p>This research was supported by Tomsk Polytechnic University Competitiveness Enhancement
Program.</p>
    </sec>
    <sec id="sec-8">
      <title>8. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Long</given-names>
            <surname>Short-Term</surname>
          </string-name>
          <string-name>
            <surname>Memory</surname>
          </string-name>
          ,
          <source>Neural Computation</source>
          <volume>9</volume>
          (
          <issue>8</issue>
          ) (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          . doi:
          <volume>10</volume>
          .1162/neco.
          <year>1997</year>
          .
          <volume>9</volume>
          .8.1735.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.S.</given-names>
            <surname>Cramer</surname>
          </string-name>
          , Econometric Applications of Maximum Likelihood Methods, Cambridge University Press,
          <year>1986</year>
          . doi:
          <volume>10</volume>
          .1017/CBO9780511572050.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Yu,</surname>
          </string-name>
          <article-title>SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient</article-title>
          ,
          <source>in: AAAI'17: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>2852</fpage>
          -
          <lpage>2858</lpage>
          . arXiv:
          <volume>1609</volume>
          .
          <fpage>05473</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          , T.-
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Vedantam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.L.</given-names>
            <surname>Zitnick</surname>
          </string-name>
          , (
          <year>2015</year>
          )
          <article-title>Microsoft COCO Captions: Data Collection and Evaluation Server</article-title>
          . arXiv:
          <volume>1504</volume>
          .
          <fpage>00325</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wang</surname>
          </string-name>
          , (
          <year>2018</year>
          )
          <article-title>Long Text Generation via Adversarial Training with Leaked Information, in: The Thirty-</article-title>
          <source>Two AAAI Conference on Artificial Intelligence</source>
          . vol.
          <volume>32</volume>
          . no.
          <issue>1</issue>
          . pp.
          <fpage>5141</fpage>
          -
          <lpage>5148</lpage>
          . arXiv:
          <volume>1709</volume>
          .
          <fpage>08624</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Sampling</given-names>
            <surname>Image</surname>
          </string-name>
          <string-name>
            <surname>COCO</surname>
          </string-name>
          ,
          <year>2020</year>
          . URL: https://github.com/CR-Gjx/LeakGAN /tree/master/Image%20COCO/save
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[7] SeqGAN neural network implementation</source>
          ,
          <year>2020</year>
          . URL: https://github.com /NikolayKrivosheev/Generation-of-short-texts-SeqGAN
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>SeqGAN</surname>
          </string-name>
          ,
          <year>2020</year>
          . URL: https://github.com/suragnair/seqGAN
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W.-J. Zhu,
          <article-title>BLEU: a Method for Automatic Evaluation of Machine Translation</article-title>
          ,
          <source>in: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . doi:
          <volume>10</volume>
          .3115/1073083.1073135.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cherry</surname>
          </string-name>
          ,
          <article-title>A Systematic Comparison of Smoothing Techniques for Sentence-Level BLEU</article-title>
          ,
          <source>in: Proceedings of the Ninth Workshop on Statistical Machine Translation</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>362</fpage>
          -
          <lpage>367</lpage>
          . doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>W14</fpage>
          -3346.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nikolenko</surname>
          </string-name>
          ,
          <article-title>Deep learning</article-title>
          .
          <source>Immersion in the world of neural net-works, Piter</source>
          ,
          <year>2018</year>
          .
          <article-title>(In Russian)</article-title>
          .
          <source>ISBN: 978-5-496-02536-2.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>