<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Image Foreground Extraction and Its Application to Neural Style Transfer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Plekhanov Russian University of Economics</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stremyanny lane</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moscow</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia v.v.kitov@yandex.ru</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Leninskie gory, 1, GSP-1, Moscow, 119991</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Foreground extraction plays important role in different computer vision applications: photo enhancement, image classification and understanding, style transfer improvement and others. New images dataset with annotation into foreground/background is proposed. Several recent neural segmentation models are trained on this dataset to extract foreground automatically and their performance is compared. The benefits of automatic foreground extraction are demonstrated on style transfer task - a popular technique for automatic rendering of photo (or content image) in the style defined by the style image, for example - the painting of a famous artist.</p>
      </abstract>
      <kwd-group>
        <kwd>foreground extraction</kwd>
        <kwd>background removal</kwd>
        <kwd>image segmentation</kwd>
        <kwd>image generation</kwd>
        <kwd>style transfer</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Foreground extraction plays important role in computer vision applications, such
as photo editing, photo enhancement, image classification, image and video
understanding, surveillance systems, style transfer. Common method to extract foreground
uses GraphCut algorithm [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] but requires human interaction to select part of
foreground area and limit the foreground into a bounding box. Some articles, such as [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
propose an automatic foreground extraction algorithms, which try to automate human
interaction in GraphCut by automatic extraction of salient regions on the image. But
such approaches require large training sets to work accurately. We propose a new
image dataset with labeled foreground and background objects, which can be used to
train and finetune automatic foreground extraction models. We propose to use
segmentation algorithms for this purpose. A segmentation algorithm takes image as input
and produces image of the same shape where each pixel is assigned to particular
object class. We propose to use binary classification with two classes – foreground and
background. Performance of two recent segmentation models is compared.
      </p>
      <p>Finally we demonstrate the benefit of automatic foreground extraction in style
transfer application. Image style transfer is a popular task of rendering input photo
(called content image) in arbitrary style, represented by style image, as shown on
fig.1. It may be applied in making creative and memorable advertisements, to improve
design of sites, groups and community pages in social networks, interiors, etc. It may
be used to apply effects in movies, cartoons, music clips and virtual reality systems
such as computer games. Various online services provide this service, such as
alterdraw.com, depart.io, ostagram.me, as well as desktop applications, such as Deep
Art Effects and mobile applications, such as prisma, artiso, vinci to mention just a
few. Adobe Photoshop has announced inclusion of this technique in their photo filters
in the 2021 version.</p>
      <p>
        Early approaches [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] used algorithms with human engineered features targeting
to impose particular styles. In 2016 Gatys et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposed an algorithm of imposing
arbitrary style taken from user defined style image on arbitrary content image by
using representations of images that could be obtained with deep convolutional
networks. Gatys et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] extended this framework in many ways. In particular weighted
stylization loss was proposed to apply masks and mix different styles together.
      </p>
      <p>
        Style transfer produces changes to the original content image which can make
important parts of it, such as human faces, figures or gestures, unrecognizable.
Schekalev et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposed to solve this problem by using weighted approach from
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to control spatial strength of stylization in different regions of content image – it
was proposed to decrease the strength of stylization for important objects (to better
preserve their structure) and to increase stylization strength on the rest of the image
(to better impose the style). Important objects were selected by regular patches,
superpixels [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and segmentation results, which showed the best result. This paper
extends their work.
      </p>
      <p>A new image dataset proposed in this work consists of over 6000 images with
automatically extracted and manually verified foreground/background mask. This
dataset allows to train new segmentation algorithms trained specifically to extract
foreground objects new on images. Two neural segmentation models are trained on the
dataset and their accuracies compared. The better model is applied to extended style
transfer algorithm where foreground objects are stylized less and background objects
are stylized more. Qualitative analysis of different resulting images shows that this
extension improves the quality of style transfer, allowing to increase recognizability
of foreground objects and by imposing style more vividly on background objects,
which is especially important for portrait stylization and advertisements where central
object on the image (a person or an advertised good) needs to stand out on the image.</p>
      <p>Proposed images dataset with labeled foreground and the neural algorithm, trained
to extract foreground automatically on new images, may be helpful not only in style
transfer applications but in other tasks, such as photo editing (to remove background),
photo enhancement (by blurring the background), image classification and
understanding (by removing unimportant background objects from consideration) which
may provide benefits in design, marketing, image &amp; video search and
recommendations as well as in automatic surveillance systems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Proposed New Images Dataset With Labeled</title>
    </sec>
    <sec id="sec-3">
      <title>Foreground</title>
      <p>
        To improve the quality of style transfer and for general photo enhancement,
including background removal and background blurring it is very important to extract the
foreground objects on the input image. This can be done with modern image
segmentation architectures, but such architectures require a training dataset to learn their
parameters. Such dataset was proposed in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], however, it consists only of 715 images,
which is not enough to train a modern deep neural network model for accurate image
segmentation. Some objects, labeled as foreground do not actually represent
foreground according to our more strict criteria, such as on figure 2. Moreover, it does not
contain images with missing foreground, which appear quite frequently in practice.
      </p>
      <p>
        Fig. 2. Examples of images from [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] having foreground objects, that are not
considered foreground according to our more strict criteria
      </p>
      <p>
        We propose a new image dataset with labeled foreground objects:
github.com/victorkitov/foreground_dataset. It has 6073 images, 1057 of which do not
contain foreground objects; the average image fraction, occupied by foreground is
0.26, the standard deviation of this fraction is 0.16. To form this dataset,
postprocessed subset of images were used from the following datasets: MC COCO
datasets [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], INRIA [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Clothing Co-Parsing [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], SUN RGB-D [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Additionally
320 images without foreground were taken from publicly available images.
      </p>
      <p>Foreground was labeled using the fact that most often the foreground includes
objects that:
• Occupy a certain share images (do not fill it entirely and are not too small);
• Located approximately in its central part, not on the edges;
• The distance to them is significantly less than to surrounding pixels;
• Belong to the class that has small spatial dimensions (thus such classes as road,
sky, sea, forest, etc. are excluded).</p>
      <p>For each dataset, formal criteria were determined for highlighting the foreground.
Then all images pre-selected according to formal criteria were manually scanned for
compliance of the selected objects with the notion of foreground.
2.1</p>
      <sec id="sec-3-1">
        <title>SUN RGB-D Dataset Processing</title>
        <p>
          SUN RGB-D [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] contains 10335 images with semantic labeling of objects and
corresponding depth maps (a depth map is a grayscale image with the same size, each
pixel value corresponds to the distance of object, located at that pixel, from the
camera). Objects were ordered by their proximity to the camera, most distant objects were
excluded, as well as objects from the following excluded classes: wall, floor, door,
window, picture, blinds, desk, curtain, mirror, clothes, ceiling, paper, whiteboard and
toilet. Also objects were excluded that occupied less than 5% of total image area or
less than 40% of area occupied by all objects of their class. Finally objects not
intersecting with central part of the image (a rectangle with width and height equal to 80%
of width and height of the original image) were also excluded. All other objects were
combined to represent the foreground of the image.
2.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Microsoft COCO Dataset Preprocessing</title>
        <p>
          MC COCO object detection 2018 validation dataset [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] consists of 5000 images
with segmented objects. For each object, information about its area and minimal
bounding box, containing the whole object, is supplied. Objects having area below
8% of total image area were not considered as well as objects, whose center did not
belong to central region of the image (a rectangle with width and height equal to 80%
of width and height of the original image). Foreground was represented by largest
object augmented by smaller objects having intersecting bounding box with the
largest object.
2.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>INRIA and Clothing Co-Parsing Datasets Preprocessing</title>
        <p>
          On INRIA images [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] people and cars were separately annotated (420 images of
people, 311 with cars). Our annotation was obtained from the original one by
combining the masks of objects located in the center of the image and occupying more than
40% of its area.
2.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Summary statistics</title>
        <p>Our combined images dataset with annotated foreground consists of 6073 images,
1057 of which do not contain foreground objects. All segmentation results were
manually scanned for compliance with the notion of foreground.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Original set</title>
        <p>Stanford
Background
SUN
RGB-D</p>
      </sec>
      <sec id="sec-3-6">
        <title>Total</title>
      </sec>
      <sec id="sec-3-7">
        <title>Included</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Automatic Foreground Segmentation</title>
      <p>
        To apply foreground segmentation automatically two segmentation models are
trained: LW RefineNet [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and Fast-SCNN [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. LW RefineNet stands for Light
Weight RefineNet and is a more efficient implementation of RefineNet model [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>Fast-SCNN uses two path architecture – image is encoded and passed through two
paths, the outputs of both paths are summed and the result is passed through a
decoder. The first path contains a convolution and the second path has multiple bottleneck
convolution layers as well as spatial pooling and upsampling. The second path
extracts high level low resolution features whereas the first path contains low lever high
resolution features.</p>
      <p>LW RefineNet encodes the image using pretrained convolution blocks of
ResNet50 classification model and applies splitting of image representations into multiple
streams with different resolutions and levels of feature abstraction which are later
joined by bilinear upscaling and summation.</p>
      <p>
        We used python realizations of LW RefineNet [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and Fast-SCNN [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] in pytorch
framework. Our foreground dataset was divided into train (4530 images), validation
(1043 images) and test sets (500 images). Images were rescaled to equal size and we
applied stochastic gradient descent algorithm with learning rate until cross
entropy loss stopped decaying on the validation set (210 epochs). To compensate class
imbalance, foreground class was accounted for with weight 0.7 and background class
– with weight 0.3.
      </p>
      <p>Performance comparison for both models is shown on table 1.</p>
      <p>LW RefineNet is more accurate than Fast-SCNN. This may be attributed to better
structure: it uses pretrained convolution blocks from ResNet-50 classification model
and combines features of diverse resolution and diverse levels of abstraction to
generate better result. Qualitative analysis shows that LW RefineNet model selects
foreground objects in most cases, while Fast-SCNN model frequently may also select
some of the background objects. LW RefineNet mask has smoother edges. On images
without foreground, LW RefineNet works better than Fast-SCNN.</p>
    </sec>
    <sec id="sec-5">
      <title>Foreground Extraction In Style Transfer</title>
      <p>
        We use style transfer method of Gatys et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] modified to preserve foreground
objects. We apply modification proposed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], namely we use spatially weighted
multiplier of content loss. This multiplier is initialized using predicted foreground by
foreground segmentation model, trained on our image dataset with labeled
foreground. Higher values of the multiplier are set for foreground area and lower values –
for background area. Such weighting allows to preserve foreground recognizability by
stylizing it less and impose vivid style on the background. Stylization results for
standard and proposed method (with foreground preservation) are shown on figure 4.
For illustrative purposes stylization strength is set to zero (no stylization) for
foreground objects.
      </p>
      <p>It can be seen that proposed image foreground dataset is sufficient to train an
accurate foreground extraction model, which in turn can be used for style transfer
improvement: important foreground objects are stylized less and are preserved more,
whereas style is applied vividly to the background.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>A new images dataset with labeled foreground objects was proposed, which may
be used for training automatic foreground extraction algorithms for wide range of
purposes, including: photo editing (automatic background removal), photo
enhancement (automatic background blurring), better image compression (with better
preservation of important objects on the foreground), automatic image captioning and scene
understanding improvement, surveillance systems (tracking of foreground objects)
and more. Two recent segmentation models were trained on the dataset and their
accuracy compared – LW RefineNet and Fast-SCNN. The former has better quality
which may be attributed to more advanced structure, utilizing ResNet-50 encoder
with skip-connections, and combinations of multiple features with different levels of
abstraction.</p>
      <p>We demonstrated the benefit of automatic foreground extraction for improving
neural style transfer. By applying spatially weighted style transfer it becomes possible
to improve stylization result by decreasing stylization strength of foreground objects
(allowing to preserve them better) and increasing stylization strength of background
(allowing to transfer style more vividly). Such improved approach has applications in
advertisement generation, design, virtual reality and entertainment industry in general.
6</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>This work proposed a new images dataset with labeled foreground objects, together
with methodology of foreground extraction and discussion of statistical properties of
the obtained dataset. Two recent automatic segmentation models were trained on this
dataset and their quality compared. Such models have many perspective applications
in various computer vision tasks. In particular it was shown how to improve image
style transfer using such models by applying style weaker to the foreground and
stronger – to the background of the image, which may have applications in design,
marketing, virtual reality, entertainment and other industries.</p>
      <p>Fig. 4. Comparison of standard and foreground aware style transfer.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Rother</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolmogorov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>"GrabCut" interactive foreground extraction using iterated graph cuts</article-title>
          .
          <source>ACM transactions on graphics 23(3)</source>
          ,
          <fpage>309</fpage>
          -
          <lpage>314</lpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Automatic foreground extraction for images and videos</article-title>
          .
          <source>In: 2010 IEEE International Conference on Image Processing</source>
          . pp.
          <fpage>2993</fpage>
          -
          <lpage>2996</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gooch</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gooch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Non-photorealistic rendering</article-title>
          . CRC Press, USA (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Strothotte</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlechtweg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Non-photorealistic computer graphics: modeling, rendering, and animation</article-title>
          . Morgan Kaufmann, USA (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gatys</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ecker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bethge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Image style transfer using convolutional neural networks</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <fpage>2414</fpage>
          -
          <lpage>2423</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gatys</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexander</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matthias</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aaron</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eli</surname>
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Controlling perceptual factors in neural style transfer</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>3985</fpage>
          -
          <lpage>3993</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Schekalev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kitov</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Style Transfer with Adaptation to the Central Objects of the Scene</article-title>
          .
          <source>In: International Conference on Neuroinformatics</source>
          <year>2019</year>
          , pp.
          <fpage>342</fpage>
          -
          <lpage>350</lpage>
          . Springer, Cham (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Superpixels introduction, https://medium.com/@darshita1405,
          <source>last accessed</source>
          <year>2020</year>
          /10/30.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gould</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fulton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Koller</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Decomposing a scene into geometric and semantically consistent regions</article-title>
          .
          <source>In: 2009 IEEE 12th international conference on computer vision</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perona</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramanan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Microsoft coco: Common objects in context</article-title>
          .
          <source>In: European conference on computer vision</source>
          . pp.
          <fpage>740</fpage>
          -
          <lpage>755</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Marszalek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Accurate object localization with shape masks</article-title>
          .
          <source>In: 2007 IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Clothing co-parsing by joint image segmentation and labeling</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <fpage>3182</fpage>
          -
          <lpage>3189</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lichtenberg</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Sun rgb-d: A rgb-d scene understanding benchmark suite</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <fpage>567</fpage>
          -
          <lpage>576</lpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Nekrasov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Light-weight refinenet for real-time semantic segmentation</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>03272</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Poudel</surname>
            ,
            <given-names>R. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liwicki</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cipolla</surname>
          </string-name>
          , R. Fast-SCNN:
          <article-title>Fast semantic segmentation network</article-title>
          . arXiv preprint arXiv:
          <year>1902</year>
          .
          <volume>04502</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Refinenet: Multi-path refinement networks for highresolution semantic segmentation</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , pp.
          <fpage>1925</fpage>
          -
          <lpage>1934</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>LW</surname>
          </string-name>
          <article-title>RefineNet realization</article-title>
          , https://github.com/DrSleep/lightweight-refinenet,
          <source>last accessed</source>
          <year>2020</year>
          /10/30.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <article-title>Fast-SCNN realization</article-title>
          , https://github.com/Tramac/Fast-SCNN-pytorch,
          <source>last accessed</source>
          <year>2020</year>
          /10/30.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>