<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Synthesizing Retro Game Screenshot Datasets for Sprite Detection</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Chanha Kim, Jaden Kim, Joseph C. Osborn Formal Analysis of Interactive Media Lab Pomona College 185</institution>
          <addr-line>East 6th Street Claremont, California 91711</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Scenes in 2D videogames generally consist of a static terrain and a set of dynamic sprites which move around freely. AI systems that aim to understand game rules (for design support or automated gameplay) must be able to distinguish moving elements from the background. To this end, we re-purposed an object detection model from deep learning literature, developing along the way YOLO Artificial Retro-game Data Synthesizer, or YARDS, which efficiently produces semirealistic, retro-game sprite detection datasets without manual labeling. Provided with sprites, background images, and a set of parameters, the package uses sprite frequency spaces to create synthetic gameplay images along with their corresponding labels.</p>
      </abstract>
      <kwd-group>
        <kwd>Figure 1</kwd>
        <kwd>Real (left) vs</kwd>
        <kwd>Synthetic (right) Screenshot from Super Mario Bros</kwd>
        <kwd>on NES</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Many videogames employ a visual language which presents
a static terrain (with background and foreground elements)
juxtaposed with dynamic, animated, and freely moving
sprites. Sprites are often game characters or objects of
interest: potential threats, powerups, or the player’s
character. Knowledge about these sprites (e.g. their type,
location, or speed) is vital for AI systems meant to understand
games, especially in systems such as automated game
design learning
        <xref ref-type="bibr" rid="ref25 ref37 ref8">(Osborn, Summerville, and Mateas 2017)</xref>
        and
learning-based level generation
        <xref ref-type="bibr" rid="ref12 ref18 ref27 ref27 ref35">(Guzdial and Riedl 2016;
Summerville et al. 2016a)</xref>
        as well as general videogame
playing and automated approaches to accessibility.
      </p>
      <p>
        Besides manual image labeling, current methods for
obtaining sprite segmentations generally involve game-specific
image processing or deep instrumentation of game
emulators, as in CHARDA
        <xref ref-type="bibr" rid="ref25 ref37 ref8">(Summerville, Osborn, and Mateas
2017)</xref>
        . These techniques can be difficult to generalize and
may be expensive, slow, or potentially fragile (e.g. the
locations of hardware sprites in memory do not correspond
exactly to the positions of human-legible sprites). In this work,
we instead generate synthetic images starting from readily
accessible spritesheets and sprite-free background images.
We can then use the synthetic datasets generated from these
resources to train sprite detection models for their
corresponding games, which should generalize better than
classical computer vision techniques or emulator instrumentation.
      </p>
      <p>While our approach still requires that users collect
spritesheets and background images, it eliminates the need
for manually labeling images or writing computer vision
code. We have packaged this generator in a Python
package called YOLO Artificial Retro-game Data Synthesizer
(YARDS). YARDS is a command-line tool which, given
sprites and background images, generates synthetic
training images that mimic patterns found in real game
screenshots. YARDS pre-formats the synthetic data for integration
with YOLO and can generate 1,000 labeled images in 10
seconds—a task which took the authors 10 long hours!</p>
      <p>In this paper, we present our two key contributions. First,
we apply recent progress made in deep learning and
computer vision to aid in producing semi-realistic datasets
suitable for training sprite detection models. Second, we
introduce a software package that can rapidly generate large
datasets. In the following sections, we will discuss our
approach for synthesizing retro-game screenshots,
demonstrate how YARDS works, and evaluate several models
trained on synthetic data against those trained on manually
labeled images.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>The intersection of computer vision and games is a growing
area of research that can benefit both designers and
players. In 2018, Luo et al. demonstrated how transfer learning
could improve the task of extracting player experiences
directly from gameplay videos. In the same year, Zhang et al.
introduced the problem of content-based retrieval of game
moments and presented a prototype search engine for
retrieving such moments based on user-provided game
screenshots (2018). Our research contributes to this area by
applying synthetic data generation techniques to the task of
training sprite detection models and by introducing a software
package that supports end-users with synthesizing their own
datasets.</p>
      <p>
        In our research, we use YOLO (You Only Look Once) to
detect sprites in four retro games: Super Mario Bros. (NES),
Earthbound (SNES), Super Street Fighter II (SNES), and
Super Mario World (SNES). YOLO is a well-known
object detection model
        <xref ref-type="bibr" rid="ref29">(Redmon et al. 2015)</xref>
        which has
seen several iterations
        <xref ref-type="bibr" rid="ref12 ref27 ref33 ref4">(Redmon and Farhadi 2016; 2018;
Bochkovskiy, Wang, and Liao 2020)</xref>
        . The original model’s
architecture is inspired by GoogLeNet
        <xref ref-type="bibr" rid="ref39">(Szegedy et al. 2014)</xref>
        and has 24 convolutional layers followed by 2 fully
connected layers (see Fig. 3). This end-to-end model allows for
efficient training and detection of objects in both images and
videos. The specific YOLO version that we are using is
Ultralytic’s YOLOv5
        <xref ref-type="bibr" rid="ref42">(Ultralytics 2020)</xref>
        , a recent
implementation of YOLO in PyTorch.
      </p>
      <p>
        The use of synthetic data to train models is a
wellknown practice in computer vision. Motivations for
generating synthetic datasets include the high cost of manually
labelling real images
        <xref ref-type="bibr" rid="ref30">(Roig et al. 2020)</xref>
        , privacy concerns
when using real user data
        <xref ref-type="bibr" rid="ref28 ref33 ref41">(Triastcyn and Faltings 2018;
Shaked and Rokach 2020)</xref>
        , and the shortage of real
training examples for rare cases
        <xref ref-type="bibr" rid="ref2">(Beery et al. 2019)</xref>
        . Other
arguments for using synthetic data include the lack of extensive
datasets in niche application domains
        <xref ref-type="bibr" rid="ref43">(Wong et al. 2019)</xref>
        and the under-representation of the full target distribution
in small datasets
        <xref ref-type="bibr" rid="ref18">(Lateh et al. 2017)</xref>
        . For such reasons,
researchers have suggested numerous approaches to
generating synthetic data over the past decade
        <xref ref-type="bibr" rid="ref23 ref32 ref33">(Nikolenko 2019;
Seib, Lange, and Wirtz 2020)</xref>
        .
      </p>
      <p>
        One common approach is to first generate a synthetic
image and then stylize the image to be more realistic
        <xref ref-type="bibr" rid="ref10 ref43 ref8 ref9">(Dwibedi,
Misra, and Hebert 2017; Georgakis et al. 2017; Wong et al.
2019)</xref>
        . For example, Wang et al. 2019 generated
photorealistic synthetic images using a virtual 3D object-environment
reconstruction method and style transfer techniques.
        <xref ref-type="bibr" rid="ref14">Hinterstoisser et al. 2019</xref>
        used 3D CAD models and pose
curricula to generate foreground-background compositions, then
made those compositions photorealistic via rendering
techniques.
      </p>
      <p>
        A significant issue arising from synthetic datasets is the
synthetic-to-real domain gap
        <xref ref-type="bibr" rid="ref40 ref44 ref45">(Tremblay et al. 2018; Yun et
al. 2019b)</xref>
        . This gap occurs when the synthetic images used
to train a model are not representative of the target image
distribution. In terms of model performance, research shows
that models trained with both real and synthetic images
achieve the best performance, followed by models trained
with purely real images
        <xref ref-type="bibr" rid="ref10 ref26 ref31 ref44 ref45 ref8 ref8 ref9">(Rozantsev, Lepetit, and Fua 2015;
Dwibedi, Misra, and Hebert 2017; Georgakis et al. 2017;
Rajpura, Bojinov, and Hegde 2017; Yun et al. 2019a)</xref>
        . These
studies also demonstrate that training with purely synthetic
images seems to detract from model performance. However,
comparable performance is achievable in image
segmentation tasks
        <xref ref-type="bibr" rid="ref7">(Di Cicco et al. 2017)</xref>
        and fine-tuning models
trained on synthetic data with additional real images can
yield better performance than mixed training
        <xref ref-type="bibr" rid="ref24">(Nowruzi et
al. 2019)</xref>
        .
      </p>
      <p>Our synthetic data generation approach is most similar to
the one presented by Dwibedi et al. 2017. Since our games
of interest use a pixelated art style, we can skip the step of
increasing the realism of the generated images; it is enough
to simply paste the sprites onto the backgrounds in a
reasonable distribution. We therefore focus on developing an
efficient technique for pasting sprites onto background
images based on sprite frequency distributions observed in real
gameplay images.</p>
    </sec>
    <sec id="sec-3">
      <title>Synthetic Data Generation Approach</title>
      <p>Our approach involves two steps: first, collect the sprites
and background images for a given game; and second, paste
sprites onto background images using their frequency
distributions. Both sprites and background images are easily
obtainable by extracting the data from an emulator,
borrowing from archives compiled by fans, or utilizing approaches
proposed by researchers like Summerville et al. 2016b.
Additional methods for obtaining sprites and background
images include scripting some game-specific image extraction
code or providing the assets directly if the user is the one
developing the game.</p>
      <p>One reason why our approach is so effective in the
videogame domain is because we are working with much
smaller image spaces (sets of possible images) than the
real-world image spaces typically used in computer vision
tasks. While the immense complexity of real-world images
can depend on virtually anything, from lighting conditions
to object textures, our focus on low-resolution, retro-game
screenshots allows us to synthesize realistic screenshots just
by pasting sprites onto background images.</p>
      <sec id="sec-3-1">
        <title>Sprite Frequency Spaces</title>
        <p>Even though we are working with low-resolution images,
synthesizing images that roughly mimic those seen in real
gameplay is a nontrivial task. We generate synthetic images
by providing the sprite frequency space for each class of
sprites to be detected. We define the sprite frequency space
for a given class as the discrete probability distribution over
the frequency of appearances for that class on a given
gameplay image. That is, it is a function mapping the numbers of
sprites in a class to the probabilities that those numbers of
sprites actually appear on a screenshot.</p>
        <p>These sprite frequency spaces can either be defined by the
user (as in our reported results) or approximated by
feeding pre-labeled images into YARDS. At runtime, YARDS
will approximate the sprite-frequency spaces by counting
the frequencies of the desired sprite classes in the image
labels. Using these sprite frequency spaces, we can determine
the number of times each sprite class should appear on each
generated image. For example, in Super Mario Bros. (NES),
there should almost always be one player on the screen,
while there may be any number of enemy sprites from zero
to 16, each number appearing at a different rate.</p>
        <p>For each output image, we choose the number of times
each class of sprites should appear by sampling the class’s
corresponding sprite frequency space. We thereby
incorporate the given or estimated frequencies with which our
target sprite classes appear in authentic gameplay images. This
allows the YOLO model to train on data that roughly
mimics our target distribution and avoid the previously-discussed
domain gap issues.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Edge Handling with Transparency Quadrants</title>
        <p>Given that sprites often have transparent pixels, we must
ensure that sprites are at least partially visible in screenshots no
matter their shape. For example, an L-shaped sprite placed in
the lower left corner of the screen might have no visible
pixels (i.e., opaque pixels) in the screenshot (see Fig. 4).
Training a model would be very difficult if our training set asserts
that background pixels are in fact part of a character. To
handle these edge cases, we propose an approach that relies on
transparency quadrants for determining whether enough of
the sprite is visible in the screenshot.</p>
        <p>For a given sprite, we first determine its transparency
quadrants by using the locations of the first and last
visible (non-transparent) pixels along each of its border axes.</p>
        <p>The upper-left quadrant (Fig. 5) is determined by the first
visible pixel in the first row of visible pixels and the first
visible pixel in the first column of visible pixels.
Corresponding horizontal and vertical lines are drawn from each
pixel, and their intersection defines the quadrant. Similarly,
the bottom-right quadrant (Fig. 5) is determined by the last
visible pixel in the last row of visible pixels and the last
visible pixel in the last column of visible pixels. The
upperright and bottom-left quadrants are formed analogously. In
general, the quadrant is defined by the intersection of the
horizontal and vertical lines drawn from each relevant pixel.</p>
        <p>Once we have the sprite’s transparency quadrants, we
determine where the sprite is clipped by the boundaries of the
screenshot. For example, if it is partially off the left side of
game title
num images
train size
mix size
label all classes
labeled classes
max sprites per class
transform sprites
clip sprites1
classification scheme2
The game’s title, which is prepended to each image’s filename to avoid naming conflicts.
The total number of images to generate.</p>
        <p>The proportion of total images which should be included in the train set.</p>
        <p>The proportion of total real images which should be included in the train set.
Determines whether all classes should be labeled or if only specific classes should be.
Determines which classes to label if label all classes is false. Useful for focusing
attention on a single sprite and introducing noise in the form of other sprites or random
images.</p>
        <p>The maximum number of sprites per class which can appear in any given image. If set to
-1, no cap will be set. Provides a means for limiting noise. Useful primarily when setting
classification scheme to random, as it allows for more control of the distribution.
Another means for introducing noise. If set to true, transforms sprites by rotating a multiple
of ninety degrees, mirroring, or scaling to twice their original size. The reason for the set
scaling is because pixel art gets distorted by any non-double scaling.</p>
        <p>Determines whether to keep all sprites entirely on screen or to allow some sprite clipping.</p>
        <p>Determines the classification scheme by which to place sprites.
the screenshot, we check the right-most quadrants (Fig. 5).</p>
        <p>Then, if the larger width of the two quadrants (i.e., the
transparency space) is greater than the width of the sprite that is
visible after clipping (i.e., the clipping space), not enough
of the sprite is within the boundaries of the screenshot. For
instance, assume the transparency space is greater than or
equal to the clipping space for a sprite being clipped off the
left of an image. In this case, we average the x-positions
of the inner vertical edges of all four transparency
quadrants. Then, we crop the sprite to be from the resulting
average x-position to the sprite’s rightmost border and paste
the cropped sprite into the screenshot with the sprite’s left
border aligned with the screenshot’s left border. Because we
use all four transparency edges to determine how to crop
the sprite, we know that enough of the sprite’s useful
information will appear on the generated image. We perform an
analogous procedure for each screen boundary that the sprite
overlaps.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Using YARDS</title>
      <p>Integrating YARDS into object detection projects is simple.
The development pipeline with YARDS involves three main
steps: preprocessing, synthetic data generation, and model
training. In preprocessing, we gather sprite and background
images for a given game and define the corresponding folder
locations in the configuration file. During synthetic data
generation, we define the parameters in the rest of the
configuration file and run YARDS via command line to generate the
images. The command-line package for YARDS takes up to
two parameters. The configuration parameter (--config
or -c) defines the path to the configuration file, and the
visualize parameter (--visualize or -v) tells the package
to draw bounding boxes for a sample of images. Finally, in
model training, we train the model and validate it in a
conventional machine learning pipeline.</p>
      <sec id="sec-4-1">
        <title>Configuration Parameters</title>
        <p>YARDS has multiple configuration parameters that the user
must define prior to using the package. Table 1 summarizes
what each parameter does, and the complete documentation
is available in the project’s source code repository.</p>
        <p>clip_sprites1 and classification_scheme2
are the parameters that control our synthetic data
generation approach. clip_sprites1 determines whether or
not the synthetic screenshots should have clipped sprites.
classification_scheme2 accepts one of four
keywords that define different methods for characterizing the
sprite frequency space: distribution, mimic-real,
random, and discrete. The distribution method
takes a set number of predefined classes such as player,
enemy, or item and corresponding sprite frequency spaces
for each class, represented by an array. For instance,
player: [0.20, 0.40, 0.40] means that for the
player class, zero sprites should appear twenty percent
of the time, one sprite should appear forty percent of the
time, and two sprites should appear forty percent of the time.
The mimic-real method analyzes a set of pre-labeled
images to approximate the sprite distribution in a dataset
and takes as input an array of class numbers, which
correspond to the class numbers in the image labels. It then uses
the approximated distributions to generate the images. The
random method samples each class with a uniform
distribution, given the maximum number of sprites for each class.
The discrete method takes inspiration from games like
Street Fighter II where each screen has a constant number
of sprites, and it takes a constant number of sprites to
display on each screenshot.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Tests and Results</title>
      <p>To evaluate our dataset synthesizer we compare model
performance for different datasets from a single game, train
binary classifiers to detect real versus synthetic data, and
demonstrate the generalizability of our approach.</p>
      <sec id="sec-5-1">
        <title>Training on Synthetic, Real, and Mixed Datasets</title>
        <p>
          To compare the performance of models trained on synthetic
data to those trained on real data, we trained nine YOLO
models on various datasets for Super Mario Bros.
(screenshot dimension: 256 192). Each YOLOv5 model
          <xref ref-type="bibr" rid="ref42">(Ultralytics 2020)</xref>
          was trained for 200 epochs with batch-size 32, and
the best weights for each model (i.e. the weights that yielded
the best model performance in training) were validated on
250 real images. We used mAP@0.5, a mean average
precision metric for measuring the performance of object
detection models, as our single-valued evaluation metric.
        </p>
        <p>Table 2 confirms the results shown by researchers in other
image recognition domains, suggesting that training mixed
datasets of real and synthetic images yields the best model
performance—although models trained on 3,750 and 10,000
synthetic images show that training with large purely
synthetic datasets can also work well. The model trained on
750 synthetic images also improved significantly after
finetuning with 750 real images for 41 epochs. We initialized
fine-tuning to train the model’s weights for 200 epochs, but
the package fast-forwarded the fine-tuning process to the
last 41 epochs. Based on these results, we recommend
using YARDS to generate either very large synthetic datasets
or smaller supplementary synthetic datasets to boost
manually labeled images.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Binary Classification of Synthetic vs. Real Data</title>
        <p>
          To test whether computer vision models could distinguish
between real and synthetic data, we trained LeNet-5
          <xref ref-type="bibr" rid="ref19">(Lecun
et al. 1998)</xref>
          , AlexNet
          <xref ref-type="bibr" rid="ref16">(Krizhevsky, Sutskever, and Hinton
2012)</xref>
          , and ResNet-50
          <xref ref-type="bibr" rid="ref13">(He et al. 2015)</xref>
          —three classic CNN
architectures of increasing complexity—to classify real
versus synthetic images. Each architecture was modified to take
in inputs of 256 256 3, configured with binary
crossentropy loss and the Adam optimizer, and trained for 50
epochs with batch-size 64. LeNet-5 was modified to use
max-pooling and ReLU activation. We trained and tested
the classifiers on 4000-image datasets composed of real and
synthetic training images from Super Mario Bros.
        </p>
        <p>Contrary to our original hypothesis that each model would
achieve roughly a 50% accuracy, Table 3 suggests that
architectures with many weights (e.g. AlexNet and ResNet-50)
are able to distinguish between real and synthetic images,
while smaller ones like LeNet-5 are not. We therefore need
to develop a training strategy to account for the discrepancy
between synthetic and real data; in the future, we may be
able to leverage these binary classifiers to guide further
improvements to our synthetic data generation approach.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Generalization</title>
        <p>We trained two very large datasets composed of synthetic
data for three separate games: Super Mario World, Super
Street Fighter II, and Earthbound (screenshot dimension:
256 224). We generated 20,000 images for each game,
combining them into a total dataset of 60,000 images. We
split the dataset at a 0.8 train-test ratio and generated two
variants: one with clipping and one without. The results can
be seen in Table 4. Model performance dropped for games
outside of those the model was trained on, which may be
due to the lack of sprites representing the larger sphere of
NES/SNES games. Given a wider variety of sprites and
games, however, we believe in the possibility of training a
general detection model for most NES/SNES games.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>Although models trained on purely synthetic datasets do not
perform as well as those trained on purely real datasets, our
results suggest that training models with mixed
syntheticand-real datasets can increase overall sprite detection
performance. Furthermore, synthesizing artificial data is much
more efficient than collecting real gameplay screenshots and
accelerates object-detection model development, ultimately
allowing for better-performing models. Additionally, large
synthetic datasets may outweigh the advantages of using real
images and make sprite detection tools more accessible.</p>
      <p>Training videogame object detectors on mixed datasets
can be useful for many applications. For example, it may
accelerate research in automated game design learning and
help relax the requirement of deep visibility into the
inner workings of emulated game hardware for
distinguishing game sprites from the level geometry. We also envision
methods for game developers to improve the accessibility
of their games by verifying that a trained model recognizes
sprites and their labels in a way which is consistent with
a designer’s intention (e.g. that enemies “read” as enemies,
that a character is not easy to misinterpret as background
texture, etc.). Such a model could help predict whether
future players would be able to make the same assumptions
and easily identify the playable parts of the game.</p>
      <p>
        Synthetic data generation could also be useful as a feature
extraction tool for reinforcement learning or other general
game-playing agents. By training models that can accurately
identify features of sprites belonging to classes like helpful,
harmful, item, enemy, etc.
        <xref ref-type="bibr" rid="ref3">(perhaps borrowed from an
affordance grammar like that of Bentley and Osborn 2019)</xref>
        , sprite
detection models may help reinforcement learning agents
train faster and generalize more effectively.
      </p>
      <p>
        That being said, there are numerous ways to improve
our approach. First, we suggest using model visualization
techniques (e.g. class activation maps, occlusion
sensitivity, gradient ascent) to visualize what the models see when
trained with synthetic versus real data. Second, an addition
that could greatly improve our current approach would be to
define spatial curves in addition to sprite frequency spaces.
Using the spatial curves to paste sprites into regions where
they would appear in real gameplay images can serve as
a way of increasing the realism of the synthetic images.
Third, trying out existing approaches, such as domain
randomization
        <xref ref-type="bibr" rid="ref21 ref33 ref40 ref6">(Liu, Liu, and Luo 2020; Borrego et al. 2018;
Tremblay et al. 2018)</xref>
        , increasing the accuracy of images
in relation to natural data
        <xref ref-type="bibr" rid="ref21 ref33">(Liu, Liu, and Luo 2020)</xref>
        ,
using generative models or GANS
        <xref ref-type="bibr" rid="ref1 ref21 ref28 ref3 ref33 ref41">(Goodfellow et al. 2014;
Liu, Liu, and Luo 2020; Bailo, Ham, and Shin 2019;
Triastcyn and Faltings 2018)</xref>
        , and procedural content
generation
        <xref ref-type="bibr" rid="ref23">(Nikolenko 2019)</xref>
        , may provide further insight into how
to refine our approach.
      </p>
      <p>Our YARDS implementation can also benefit from
additional features. First, incorporating multiprocessing would
greatly increase the speed of synthetic data generation. Our
current package generates 80 images per second on a
single core with no GPU acceleration for Super Mario Bros.,
and parallelizing this task would increase the package’s
efficiency. Second, adding basic image rendering and filtering
functions such as blurring or pixelating sprites may be
useful for videogames that do not use the pixelated style and
resolution common to the four games we examined in this
work. Third, color filtering functions may help the object
detection models learn the sprites’ essential features and avoid
overfitting to their color patterns. Fourth, we would like to
see added support for games in a wider variety of gameplay
styles and genres. Finally, adding text detection functions
may help with including basic user-interface elements.</p>
      <p>In summary, this paper has introduced an application of
existing synthetic data generation research to the problem
of sprite detection and a software package that enables an
end-user to rapidly generate large, synthetic training
images based on sprite frequency spaces and edge-handling.
An open-source and working prototype of YARDS is
available at https://github.com/faimSD/yards. We hope that our
paper and software package will inspire further research in
sprite detection and in computer vision and games.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bailo</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ham</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Shin</surname>
            ,
            <given-names>Y. M.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Red blood cell image generation for data augmentation using conditional generative adversarial networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Beery</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Morris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Piavis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Kapoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Meister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Joshi</surname>
          </string-name>
          , N.; and Perona,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Synthetic examples improve generalization for rare classes</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bentley</surname>
            ,
            <given-names>G. R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Osborn</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>The videogame affordances corpus</article-title>
          .
          <source>In 2019 Experimental AI in Games Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bochkovskiy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            , C.-Y.; and Liao, H.-
            <given-names>Y. M.</given-names>
          </string-name>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>Yolov4: Optimal speed and accuracy of object detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Borrego</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dehban</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Figueiredo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Moreno,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Bernardino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ; and
            <surname>Santos-Victor</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Applying domain randomization to synthetic data for object category detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Di</given-names>
            <surname>Cicco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Potena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Grisetti</surname>
          </string-name>
          , G.; and
          <string-name>
            <surname>Pretto</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>Automatic model based dataset generation for fast and accurate crop and weeds detection</article-title>
          .
          <source>2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Dwibedi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Misra</surname>
            ,
            <given-names>I.;</given-names>
          </string-name>
          and Hebert,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Cut, paste and learn: Surprisingly easy synthesis for instance detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Georgakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mousavian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Kosecka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Synthesizing training data for object detection in indoor scenes</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          2014.
          <article-title>Generative adversarial networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Guzdial</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Riedl</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Toward game level generation from gameplay videos</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>CoRR abs/1512</source>
          .03385.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hinterstoisser</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Pauly</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Heibel</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ; Marek,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and Bokeloh,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>An annotation saved is an annotation earned: Using fully synthetic training for object instance detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Koylu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Deep neural networks and kernel density estimation for detecting human activity patterns from geo-tagged images: A case study of birdwatching on flickr</article-title>
          .
          <source>ISPRS International Journal of GeoInformation</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.;</given-names>
          </string-name>
          and Hinton,
          <string-name>
            <surname>G. E.</surname>
          </string-name>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          . In
          <string-name>
            <surname>Pereira</surname>
          </string-name>
          , F.;
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>C. J. C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <article-title>and Weinberger, K. Q</article-title>
          ., eds.,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>25</volume>
          . Curran Associates, Inc.
          <fpage>1097</fpage>
          -
          <lpage>1105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Lateh</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Muda</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yusof</surname>
            ,
            <given-names>Z. I. M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Muda</surname>
            ,
            <given-names>N. A.</given-names>
          </string-name>
          ; and Azmi,
          <string-name>
            <surname>M. S.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Handling a small dataset problem in prediction model by employ artificial data generation approach: A review</article-title>
          .
          <source>Journal of Physics: Conference Series</source>
          <volume>892</volume>
          :
          <fpage>012016</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Lecun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and Haffner,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <article-title>Gradient-based learning applied to document recognition</article-title>
          .
          <source>In Proceedings of the IEEE</source>
          ,
          <fpage>2278</fpage>
          -
          <lpage>2324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ; Liu, J.; and
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Can synthetic data improve object detection results for remote sensing images?</article-title>
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Guzdial</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Liao</surname>
            , N.; and Riedl,
            <given-names>M.</given-names>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>Player experience extraction from gameplay video</article-title>
          . CoRR abs/
          <year>1809</year>
          .06201.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Nikolenko</surname>
            ,
            <given-names>S. I.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Synthetic data for deep learning</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Nowruzi</surname>
            ,
            <given-names>F. E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kapoor</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kolhatkar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hassanat</surname>
            ,
            <given-names>F. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Laganiere</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Rebut</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>How much real data do we actually need: Analyzing object detection performance using synthetic and real data</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Osborn</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Summerville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and Mateas,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Automated game design learning</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Rajpura</surname>
            ,
            <given-names>P. S.</given-names>
          </string-name>
          ; Bojinov, H.; and Hegde,
          <string-name>
            <surname>R. S.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Object detection using deep cnns trained on synthetic images.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Redmon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Farhadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Yolo9000: Better, faster, stronger</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Redmon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Farhadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Yolov3: An incremental improvement</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Redmon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Divvala,
          <string-name>
            <given-names>S.</given-names>
            ; Girshick, R.; and
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Roig</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Varas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Masuda</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ; Riveiro,
          <string-name>
            <surname>J. C.</surname>
          </string-name>
          ; and BouBalust,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>Unsupervised multi-label dataset generation from web data</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Rozantsev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lepetit</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ; and Fua,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>On rendering synthetic images for training an object detector</article-title>
          .
          <source>Computer Vision and Image Understanding</source>
          <volume>137</volume>
          :
          <fpage>24</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Seib</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lange</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Wirtz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Mixing real and synthetic data to enhance neural network training - a review of current approaches</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Shaked</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rokach</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Privgen: Preserving privacy of sequences through data generation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          2016a.
          <article-title>Learning player tailored content from observation: Platformer level generation from video traces using lstms</article-title>
          .
          <source>In AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.</source>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Summerville</surname>
            ,
            <given-names>A. J.</given-names>
          </string-name>
          ; Snodgrass,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Mateas,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and Ontano´n,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2016b</year>
          .
          <article-title>The vglc: The video game level corpus</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>arXiv preprint arXiv:1606</source>
          .
          <fpage>07487</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Summerville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Osborn</surname>
            , J.; and Mateas,
            <given-names>M.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Charda: Causal hybrid automata recovery via dynamic analysis</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <source>arXiv preprint arXiv:1707</source>
          .
          <fpage>03336</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ; Jia,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Sermanet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Anguelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Erhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ; and
            <surname>Rabinovich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Going deeper with convolutions</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>Tremblay</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Prakash</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Acuna</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Brophy,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Jampani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ;
            <surname>Anil</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          ; To,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Cameracci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Boochoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Birchfield</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Training deep networks with synthetic data: Bridging the reality gap by domain randomization</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <surname>Triastcyn</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Faltings</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Generating artificial data for private deep learning</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <surname>Ultralytics</surname>
          </string-name>
          .
          <year>2020</year>
          . Yolov5.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>M. Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kunii</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Baylis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ong</surname>
            ,
            <given-names>W. H.</given-names>
          </string-name>
          ; Kroupa,
          <string-name>
            <given-names>P.</given-names>
            ; and
            <surname>Koller</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Synthetic dataset generation for object-to-model deep learning in industrial applications</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>Yun</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Eldin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Huyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Chow</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2019a</year>
          .
          <article-title>Small target detection for search and rescue operations using distributed deep learning and synthetic data generation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <string-name>
            <surname>Yun</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kim</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2019b</year>
          .
          <article-title>Balancing domain gap for object instance detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Holtz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <source>In Proceedings of the 13th International Conference on the Foundations of Digital Games</source>
          , FDG '
          <fpage>18</fpage>
          . New York, NY, USA: Association for Computing Machinery.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>