<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Text Scene Detection with Transfer Learning in Price Detection Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladimir Fomenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Botov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julius Klenin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Chelyabinsk State University</institution>
          ,
          <addr-line>Chelyabinsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses the use of the transfer learning method in a text scene detection task. The transfer learning is an e ective method in image analysis tasks, in particular, classi cation and object detection. Nevertheless, the application of this approach in combination with various methods for detecting text scenes almost is not described. The experiment is conducted to transfer knowledge about text detection to the detection for price tags using Fully Convolutional Network. COCO-Text, based on the MS COCO dataset, ImageNet dataset and dataset made up of a large number of images of the prices of various stores obtained during monitoring are taken as base datasets for the transfer learning. The target dataset is compiled from photographs of price monitoring with price tags and prices marked on them. The results of the experiment show how the application of transfer learning a ects the training speed of a FCN in the task of detection price tags and prices for each of the base datasets.</p>
      </abstract>
      <kwd-group>
        <kwd>Text Scene Detection</kwd>
        <kwd>Transfer Learning</kwd>
        <kwd>Convolutional Neural Networks</kwd>
        <kwd>Semantic Segmentation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Monitoring of prices allows for e ective pricing based on the prices of
competitors. At the moment, many retail chains conduct the price monitoring process in
the following way: low-quali ed personnel visits competitors' shops, photographs
the goods and price tags of the monitored goods, and writes down information
about the goods and prices to the database, and then other employees verify the
correctness of the information entering into the database. Every day more than
one million photographs are received for monitoring, and therefore this process
does not allow to quickly analyze the prices of competitors and conduct pricing
based on this information.</p>
      <p>
        The task of price tag recognition aims to signi cantly speed up the process
of monitoring the prices of competitors by reducing the number of photos being
checked. The whole task is performed in two stages: the localization of the price
tag and the price and the recognition of the name of the product and its price
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Localization can be performed by methods based on object detection and
semantic segmentation. Recognition of the name of the products can be carried
out by means of the classi cation of sentences or words or methods of OCR
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Price recognition can only be performed by OCR methods. There are also
end-to-end recognition methods [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], but they require signi cant computational
resources.
      </p>
      <p>The collection and labelling of the dataset for the detection of price tags and
prices is quite complex and expensive, and therefore this article proposes the use
of the method of transfer learning to solve the problem of nding price tags.</p>
      <p>
        In the transfer learning, two approaches can be distinguished: the transfer
learning from a similar task and the transfer learning from a similar domain
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We can assume that the task of nding the price tag is similar to the task
of nding the text and dataset can be used to search for text, such as
COCOText [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Considering that retail chains collect more than a million photographs
every day with labeling of product names, we can take a model trained to classify
goods by photo and use it as a base model for nding price tags.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>
        There are two approaches to nding text in an image. The rst approach involves
methods based on region proposals. State-of-the-art region proposals methods
such as R-CNN in its pure form are not suitable for text searching because the
design of the anchor box is not suitable for a large aspect ratio of text strings
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. There are methods that solve this problem by using long anchor boxes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
or the Region Proposal Network [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. At the moment, such methods are working
well only with horizontal text.
      </p>
      <p>
        The second approach includes methods based on segmentation, such as
TextBlock FCN [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which is a modi cation Fully Convolutional Networks [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. These
methods allow you to do per-pixel prediction of nding text in the image and
do not have problems with the text of irregular shape, but they require
timeconsuming process of separation to get the result on the words.
      </p>
      <p>
        Moreover, hybrid methods are emerging that combat the shortcomings of
both approaches [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In recent years, the transfer learning method has been applied to most image
analysis tasks, such as the classi cation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and the objects detection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
transfer learning makes it possible to accelerate the process of training computer
vision models at times. There are two approaches to the transfer learning: based
on a similar problem and based on a similar domain. Many deep learning methods
for imaging tasks are used as a base architecture methods that show high results
in the ImageNet competition [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <p>
        For the experiment, the Fully Convolutional Network method was chosen. This
method solves the problem of semantic segmentation. The key feature of this
method is that the architectures implementing this method do not contain any
fully connected layers (see Fig. 1) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        At the input, such a neural network receives an image and an output
image is created with the number of channels equal to the number of predicted
classes, where each channel is a binary mask showing where the object of the
corresponding class is on the image. In the original article [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the best results
were shown by an architecture based on the architecture of VGG16 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which
showed the best results in the ImageNet contest in 2014.
      </p>
      <p>To convert VGG16 architecture into FCN, the following operations were
performed: fully connected layers, atten layer and last max pooling layer have been
removed; convolutional layer with a kernel size of 1x1 , the number of lters equal
to 128 and the relu activation function, upsampling layer increases the output
of the last convolutional layer 16 times and convolution layer with a kernel size
of 3x3, a number of lters equal to 2, and a sigmoid activation function have
been added. The Adam optimizer was used with the parameters: learning rate
is 10 3; 1 is 0.9, 2 is 0.999, is 10 7:</p>
      <p>
        In this article, in addition to VGG16, the use of the MobileNet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
architecture is proposed. To convert MobileNet architecture into FCN, the following
operations were performed: a fully connected layer and a global average pooling
layer have been removed; were then added convolutional layer with a kernel size
of 1x1 , the number of lters equal to 128 and the relu activation function,
upsampling layer increases the output of the last convolutional layer 32 times and
convolution layer with a kernel size of 3x3, a number of lters equal to 2, and
a sigmoid activation function have been added. The Adam optimizer was used
with the parameters: learning rate is 10 4; 1 is 0.9, 2 is 0.999, is 10 7:
      </p>
      <p>The transfer learning was carried out as follows: the weights of the target
network were initialized by the weights of the networks trained on other tasks,
then the target network was trained on the target dataset.</p>
      <p>The training was terminated by the early stopping method: the training
stopped when there was no improvement in the metric on the validation set for
ve epochs.</p>
      <p>The following data augmentations were used: image rotation from -10 to 10
degrees, shift from 0 to 0.1 in any direction, zoom from 0.9 to 1.1.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Dataset</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Base datasets</title>
        <p>The transfer learning was conducted by the following datasets.</p>
        <p>1. ImageNet. Dataset used in the annual competition for pattern recognition.
At the moment it contains 14,197,122 images containing 21,841 categories.</p>
        <p>2. COCO-Text. On this dataset, the problem of text search was solved. It
is made up of pictures included in the MS COCO dataset of photos containing
text. It contains 63,686 images with 145,859 text instances.</p>
        <p>3. Dataset collected from the price monitoring pictures. This dataset was
used to solve the problem of classi cation of products from the photo. It contains
about 500,000 photos of goods divided into 150 classes.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Target dataset</title>
        <p>The target dataset consists of 10,642 price monitoring photographs of more than
10 stores, on which there is one of six types of goods and a price tag. For each
photo, the areas in which the price tag and price are located are labeled. In the
photo, there may be more than one price tag and more than one price on the
price tag and in such cases a price tag that corresponds to the product in the
photo was chosen, and the price was selected based on the expert's opinion (see
Fig. 2).
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>5.1</p>
      <sec id="sec-5-1">
        <title>Metric</title>
        <p>The Jaccard coe cient metric, also known as Intersection over Union, was
chosen, which is calculated as the intersection of the predicted and true areas of the
object's location by the image divided by their union.
As shown in Table 1, using the transfer learning, it was possible to increase the
speed of training and it was not possible to improve the quality of models. The
greatest increase in the speed of training was due to the transfer learning based
on one subject area. The transfer learning based on the same task did not give
a gain in speed in comparison with the transfer learning based on the ImageNet
dataset.
Based on the results of the experiment, it can be concluded that the use of
the transfer learning method makes it possible to make learning process several
times faster. The best approach was the approach of transferring learning from
the tasks of same subject area. In our experiment, one epoch of training took
5 minutes and we managed to reduce the training of models from 3.5 hours to
45 minutes. Acceleration of training models will allow us to more e ectively test
models with di erent parameters.</p>
        <p>In this paper, the application of the transfer learning to the Fully
Convolutional Network method, which is based on segmentation, was considered. In
the future, it is planned to explore the e ectiveness of the transfer of learning
for methods based on region proposals. In addition, in the future, the domain's
dataset will grow to more than a million images divided into more than 1000
classes, which, perhaps, will further accelerate the training process for nding
price tags and prices.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aytar</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Transfer Learning for Object Category Detection</article-title>
          . http://www.robots. ox.ac.uk/~vgg/publications/2014/Aytar14a/ (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Deep Direct Regression for Multi-Oriented Scene Text Detection</article-title>
          .
          <source>arXiv preprint arXiv:1703.08289</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalenichenko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weyand</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andreetto</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adam</surname>
          </string-name>
          , H.:
          <article-title>MobileNets: E cient Convolutional Neural Networks for Mobile Vision Applications</article-title>
          . aeXiv preprint arXiv:
          <volume>1704</volume>
          .04861 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Huh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Efros</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>What makes ImageNet good for transfer learning?</article-title>
          <source>arXiv preprint arXiv:1608.08614</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jaderberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Reading Text in the Wild with Convolutional Neural Networks</article-title>
          .
          <source>arXiv preprint arXiv:1412</source>
          .
          <year>1842</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>TextBoxes: A Fast Text Detector with a Single Deep Neural Network</article-title>
          .
          <source>arXiv preprint arXiv:1611.06779</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shelhamer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Fully Convolutional Networks for Semantic Segmentation</article-title>
          .
          <source>arXiv preprint arXiv:1605.06211</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Oquab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laptev</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sivic</surname>
          </string-name>
          , J.:
          <article-title>Learning and Transferring Mid-Level Image Representations using Convolutional Neural Networks</article-title>
          .
          <source>CVPR'14 Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>1717</fpage>
          -
          <lpage>1724</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Very Deep Convolutional Networks for Large-Scale Image Recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tian</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qiao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Detecting Text in Natural Image with Connectionist Text Proposal Network</article-title>
          .
          <source>arXiv preprint arXiv:1609.03605</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Veit</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matera</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images</article-title>
          .
          <source>arXiv preprint arXiv:1601.07140</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Wojna</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gorban</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ibarz</surname>
          </string-name>
          , J.:
          <source>Attentionbased Extraction of Structured Information from Street View Imagery. arXiv preprint arXiv:1704.03549</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Xing</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
          </string-name>
          , Y.:
          <article-title>ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene</article-title>
          .
          <source>arXiv preprint arXiv:1711.11249</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          <article-title>: Multi-Oriented Text Detection with Fully Convolutional Networks</article-title>
          .
          <source>arXiv preprint arXiv:1604.04018</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Scene text detection and recognition: recent advances and future trends</article-title>
          .
          <source>Frontiers of Computer Scienc February</source>
          <year>2016</year>
          , Volume
          <volume>10</volume>
          , Issue 1, pp.
          <fpage>19</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>