<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Image Analysis in Technical Documentation (Discussion Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fabio Carrara</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Franca Debole</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Gennaro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Amato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manufac-</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Information Science and Technologies (ISTI), Italian National Research Council (CNR)</institution>
          ,
          <addr-line>Via G. Moruzzi 1, 56124 Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the era of Big Data, manufacturing companies are overwhelmed by a lot of disorganized information: the large amount of digital content that is increasingly available in the manufacturing process makes the retrieval of accurate information a critical issue. In this context, and thanks also to the Industry 4.0 campaign, the Italian manufacturing industries have made a lot of e ort to ameliorate their knowledge management system using the most recent technologies, like big data analysis and machine learning methods. This paper presents the on-going work done within the ADA project, with special emphasis on the speci c image analysis work carried out to extract information from images contained in the so di erent document of the manufacturing companies, partners of the project.</p>
      </abstract>
      <kwd-group>
        <kwd>Image Analysis turing companies</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Big Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Manufacturing companies, which produce complex products and manage large
plants, generate a consistent ow of data and information throughout the
company processes, from acquisition to production and maintenance of the products
themselves. In this amount of data, a signi cant part consists of texts, graphics,
and images obtained as a transposition of the know-how of human personnel.</p>
      <p>Collecting and retrieving all this data and information quickly and easily is
vital for speeding up internal business activities. For example, during the design
phases of a new product, it is useful to be able to identify, in past projects,
speci cations, data and information contained in lessons learned, in risk analysis, to
carry out more reliable and innovative design activities. In case of plants
maintenance, it is invaluable for operators to have immediate access to the information
necessary to carry out their work quickly and e ectively.</p>
      <p>Copyright c 2019 for the individual papers by the papers authors. Copying
permitted for private and academic purposes. This volume is published and copyrighted by
its editors. SEBD 2019, June 16-19, 2019, Castiglione della Pescaia, Italy.</p>
      <p>The needs of the manufacturing companies described above, however, clash
with the complexity of knowledge management, due both to the large
quantity and to the heterogeneity of the data and documents to be processed: the
technological tools currently available on the market are not able to rise this
challenge and e ectively meet these needs. In this context, therefore, the main
target of the ADA project is to design and develop a platform based on big data
analytics systems that allows for the acquisition, organization, and automatic
retrieval of information from technical texts and images in the di erent phases
of acquisition, design &amp; development, testing, installation and maintenance of
products.</p>
      <p>In this paper, we illustrate the work carried out in the ADA project
focusing on the image content retrieval part: the images contained in the corporate
documents constitute a relevant source of information that could be relevant in
the manufacturing work ow. On this context, we developed speci c techniques
for the extraction, classi cation, recognition, and tagging of images within
technical documentation. The architecture and the methodologies of our work are
presented on Section 3 and Section 4, respectively.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Image recognition techniques have been extensively studied in the last decade
in the eld of computer vision and the recovery of multimedia information.
Deep learning techniques, such as those based on Convolutional Neural
Networks (CNNs), represent today the state of the art for the most varied computer
vision activities such as image classi cation, image recovery and object
recognition [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Furthermore, the use of intermediate layer activation as a high-level
descriptor (feature) of visual image content has become very popular and has
proven to be e ective as demonstrated by many scienti c papers [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
Convolutional neural networks exploit the computing power provided by the GPU-based
architectures, in order to learn from huge collections of multimedia information
(e.g. images). One of the limitations of this approach is that many collections
of images available were created for academic purposes (e.g. ImageNet [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) and
can not be used e ectively for applications such as those discussed in the ADA
project.
      </p>
      <p>
        In the project, the development of tools to search for graphic symbols
belonging to technical schemes within the technical documentation is of particular
importance. To this end, some works [
        <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
        ] that try to tackle the problem by
using CNNs seem to be promising. However, as Elyan et al. claim in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the
application of CNNs to detect and localize symbols in drawings is still a
challenging task. This is probably due to the complexity of the problem and also to
the lack of su cient annotated examples or publicly available data sets.
      </p>
      <p>
        Elyan et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presented a semi-automatic and heuristic-based approach to
localise symbols within engineering drawings, and then applied a CNN to classify
the detected symbols. Similarly, Quan et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] used an AlexNet to classify point
symbols in color topographic maps.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Image Analysis Component</title>
      <p>The technical documents of manufacturing companies often contain a large
number of heterogeneous images such as graphics, wiring diagrams, mechanical
drawings, etc. While on the one hand the images are almost always accompanied by
text such as captions, descriptions, and labels, being able to recognize their
content without using text is an increasingly requested feature: images represent a
rich source of information.</p>
      <sec id="sec-3-1">
        <title>Image Module</title>
      </sec>
      <sec id="sec-3-2">
        <title>Searcher</title>
        <p>Provides image search capabilities via category matching, object presence, and visual similarity
Image Dataset</p>
      </sec>
      <sec id="sec-3-3">
        <title>FExtractor</title>
      </sec>
      <sec id="sec-3-4">
        <title>Classifier</title>
      </sec>
      <sec id="sec-3-5">
        <title>Recognizer</title>
      </sec>
      <sec id="sec-3-6">
        <title>Indexer</title>
      </sec>
      <sec id="sec-3-7">
        <title>Image Index</title>
        <p>The Image Analysis Module (Fig. 1) aims to enable companies to extract
information from images contained within their technical documentation. This
information is necessary to perform a search and classi cation of images based
on visual content. The module in particular deals with the identi cation and
extraction from the images the content, classifying them, recognizing some graphic
elements, such as symbols, within them and looking for similar ones.</p>
        <p>The Image Analysis Module is composed by the following sub-components:
FExtractor Module. This component is intended to extract from one image
one or more visual descriptors (features) that allow one to perform searches
based on visual similarity using an image as a query.</p>
        <p>Classi er Module. The classi er component deals with the classi cation of
the images in the various possible types in the eld of technical and patent
documentation (e.g. technical drawing in perspective, sections, electronic circuit,
ow chart, etc.). The tool relies on modern Machine Learning techniques based
on Deep Learning. The automatic classi cation is based on training a neural
network on a number of examples for each of the types of images to be classi ed.
The accuracy of the tool is strictly correlated to the number of training examples
provided.</p>
        <p>Recognizer Module. This component allows the automatic recognition of
objects within the images classi ed by the classi cation tool: once the images are
classi ed through the classi cation tool, the objects are automatically recognized
within the image. For each type of image, the tool will have a number of
examples related to the objects to be identi ed. Another result of this component
is to automatically associate one or more tags with an image: the tags can be
further enriched or corrected by analyzing the text in the document that refers
to the image itself.</p>
        <p>Indexer Module. This component deals with the appropriate indexing of both
the visual features of the images and the context features related to the image:
the image is decomposed into a visual part and into a symbolic part. For each
image in fact, we will have both the global visual features necessary to perform
visual similarity queries, and local features to perform queries able to detect and
localize speci c symbols contained in the image.</p>
        <p>Searcher Module. This module deals with sorting the various types of search
supported (see the details below).</p>
        <p>The search component allows the user to perform di erent types of searches
using the index created by the Indexer module:
1. Textual Search on the text correlated or extracted from the images.
2. Similarity Search on images: search by similarity of the basic elements
(symbols) in an image archive using an example as a query.
3. Search for the categories to which an image belongs.</p>
        <p>Furthermore, the Searcher Module is able to automatically handle external
queries, i.e. using images that are not present in the index, as well as
internal queries, i.e. using images already present in the database as queries. For
the external queries, the Searcher Module uses FExtractor Module to extract
features from the query image.</p>
        <p>In the next section, we will describe the speci c methodologies exploited
for the realization of the four main mentioned modules: FExtractor, Classi er,
Recognizer, and Indexer.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <p>Due to their astonishing e ectiveness on perceptual tasks, we resort to
stateof-the-art Deep Learning techniques based on Convolutional Neural Networks</p>
      <sec id="sec-4-1">
        <title>VNCO VNCO ...</title>
        <p>VCNO VCNO</p>
        <p>CNN</p>
        <p>ResNet-101</p>
      </sec>
      <sec id="sec-4-2">
        <title>VNCO VNCO ...</title>
        <p>VNCO VNCO</p>
        <p>CNN
Features Extractor</p>
        <p>Recognizer</p>
        <p>R
O
S
S
E
R
G
E
R
0.2
-1.5
5.4</p>
        <p>…
1.0
0.0
RegL(iaoRsn-taMLlAaPCyoe)orling DeIms-c8ar.gi3petor</p>
      </sec>
      <sec id="sec-4-3">
        <title>VNCO VNCO ...</title>
        <p>VCNO VCNO</p>
        <p>CNN
Features Extractor
n
o
itt
a
u
m
r
e
P
p
e
e
D
(CNN) to implement the visual analyses conducted by the aforementioned
modules.</p>
        <p>In the following, we will describe in details the methods chosen to implement
the main functions of the Image Analysis Module, and how they are distributed
among its sub-components. As depicted on Figure 2, for each image on the
dataset:
{ we extract the features (FExtractor) as descriptors of the image on the index
(Indexer); the extracted features are also made available to the Recognizer
and Classi er modules;
{ for speci c categories, we use ad-hoc techniques for the object recognition
(Recognizer) and the classi cation of the image (Classi er); we rely on
simpler techniques based on previously extracted features otherwise;
{ all the information deriving from FExtractor, Recognizer and Classi er are
memorized in the Image Record on a speci c index (Indexer).
4.1</p>
        <p>
          Visual Similarity Search
For the implementation of the visual similarity search, we employ
state-of-theart global image descriptors extracted from CNN speci cally tailored for image
retrieval. Speci cally, we select the Region Maximum Activation of Convolution
(R-MAC) descriptors [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] as a compact and expressive image descriptors, which
enables our system to support instance-level visual object retrieval on both
natural photos and schematics/synthetic images.
        </p>
        <p>Given an image, the FExtractor module rst feeds it to a pre-trained
convolutional network to extract the output of the last convolutional layer, which
is composed by multiple feature maps having two spatial dimensions. Then, we
compute the R-MAC descriptor by pooling and aggregating di erent parts (or
regions) of those feature maps. The feature maps are max-pooled over several
regions on their spatial dimensions, and the obtained vectors are aggregated by
summation and the result l2-normalized.</p>
        <p>
          The choice of the pre-trained CNN for the extraction of the feature maps
is essential: instead of choosing a network trained on generic object recognition
tasks (such as ImageNet), we select the model developed by [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], a ResNet-101
convolutional network trained speci cally for the task of same-object retrieval.
With this con guration, we obtain 2048-dimensional image descriptors that can
be compared with the cosine similarity to search for visual matches.
        </p>
        <p>As already mentioned on Section 3, part of the work done for the ADA
project is to make use also of the text correlated or extracted from the image. In
the following paragraph, we will explain how we realized an appropriate index
supporting both the visual and textual features of the images in an e cient way.
4.2</p>
        <p>
          Indexing of Image Descriptors
While textual information, such as labels and tags, can be stored in a relational
database, we need a similarity search index structure to store high-dimensional
image features and e ciently compute the cosine-similarity to perform retrieval.
Despite open-source indexes for e cient similarity search exist [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], they still come
with caveats and, in general, are not mature and well-supported as the scalable
disk-based textual indexes such as Elastichsearch.
        </p>
        <p>
          Seeking for simplicity, we propose to adopt Elasticsearch in the Indexer and
Searcher modules as database for all the information extracted by the visual
analysis modules, and adapt our similarity search needs to the full-text search
engine provided by the software. Speci cally, the adoption of two techniques,
Deep Permutations [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and Surrogate Text Representation [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ], permit us to
represent our image descriptors as text, index them using a full-text search
engine based on inverted indexes, and perform similarity queries without the
need of a image-speci c similarity index.
        </p>
        <p>
          The Deep Permutations technique is based on the fact that in high-level
deep-learned features, each dimension represents an abstract visual concept, and
its value speci es the importance of that concept. Similar features tend to have
the same relative importance among visual concepts, and thus, we can
approximate a oat vector of features by sorting its dimensions in descending order and
keeping the sorting permutation (an integer vector). By truncating the
permutation to the top elements only, Amato et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] demonstrated that this coarse
approximation is su cient to obtain state-of-the-art trade-o s between e ciency
and e ectiveness, permitting large-scale searches with query time in the order
of seconds.
        </p>
        <p>The Surrogate Text Representation technique aims at encoding a given
integer vector in the term frequency values of a full-text search engine by
generating an appropriate textual string. Used in combination with Deep Permutations,
it permits us to store sparse truncated permutation in the inverted index for
textual data and leverage the vector document model of full-text search engines to
perform cosine similarity searches.
4.3</p>
        <p>
          Image Categorization and Object Detection
For certain documentation, the Classi er module needs to assign each image to
one or more categories de ned by the system users. Thus, we have to implement
a multi-class multi-label image classi er. For the categories with a considerable
number of training examples, a CNN can be trained to implement the image
classi er. Instead of de ning and training a new model from scratch, we propose
to leverage the transfer learning practice, exploiting the knowledge of popular
CNNs that have been successfully trained in large-scale generic object
recognition tasks. Speci cally, we plan to use the ResNet-50 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] model pre-trained on
ImageNet1k, replacing the last classi cation layer to match our number of classes
and to perform ne-tuning on the available training samples. When training
samples are scarce, we resort to a simpler kNN classi er based on visual similarity:
we reuse the same features extracted by the FExtractor module and the cosine
similarity matching function in the implementation of the kNN classi er.
        </p>
        <p>
          For speci c categories of images, the Recognizer module needs to detect
the presence of particular objects of interest. If the location of the object is not
required, R-MAC features matching can be used to reveal object presence. Being
an aggregation of local region descriptors, matching the R-MAC descriptor of
a sub-region of an image against the whole image descriptor will yield a high
score. Thus, we can perform object detection via instance-level similarity search
reusing the same modules. Instead, if more ne-grained information about
detected objects is required, such as the speci c localization in the image, and
enough training samples are available, we resort to more complex object
detection techniques based on Deep Learning, such as region-based CNNs [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] or
single-stage detectors [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], that can be ne-tuned on the speci c object category
to detect.
        </p>
        <p>Both the information extracted by the Classi er and Recognizer modules
are stored by the Indexer module in Elasticsearch as additional elds of the
image record, and accordingly searchable through the Searcher module.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper brie y introduces the on-going work on image analysis in the ADA
project, whose main aim is to support the innovation of the production process
of the manufacturing companies. We focused on the description of the
methodologies carried out to extract relevant information from images contained in the
di erent document provided by the project partners.
This work was partially funded by \Automatic Data and documents Analysis
to enhance human-based processes" (ADA), CUP CIPE D55F17000290009. We
gratefully acknowledge the support of NVIDIA Corporation with the donation
of a Tesla K40 GPU used and a Jetson TX2 board used for this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amato</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolettieri</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrara</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gennaro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Large-scale image retrieval with elasticsearch</article-title>
          .
          <source>In: The 41st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval</source>
          . pp.
          <volume>925</volume>
          {
          <fpage>928</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Amato</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrara</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gennaro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>E cient indexing of regional maximum activations of convolutions using full-text search engines</article-title>
          .
          <source>In: Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval</source>
          . pp.
          <volume>420</volume>
          {
          <fpage>423</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Amato</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gennaro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vadicamo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Deep permutations: Deep convolutional neural networks and permutation-based indexing</article-title>
          .
          <source>In: International Conference on Similarity Search and Applications</source>
          . pp.
          <volume>93</volume>
          {
          <fpage>106</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Donahue</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            <given-names>man</given-names>
          </string-name>
          , J., Zhang, N.,
          <string-name>
            <surname>Tzeng</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Decaf: A deep convolutional activation feature for generic visual recognition</article-title>
          .
          <source>arXiv preprint arXiv:1310.1531</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Elyan</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jayne</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Symbols classi cation in engineering drawings</article-title>
          .
          <source>In: 2018 International Joint Conference on Neural Networks (IJCNN)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gordo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almazan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Revaud</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larlus</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>End-to-end learning of deep visual representations for image retrieval</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>124</volume>
          (
          <issue>2</issue>
          ),
          <volume>237</volume>
          {
          <fpage>254</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1512.03385</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Douze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jegou</surname>
          </string-name>
          , H.:
          <article-title>Billion-scale similarity search with gpus</article-title>
          .
          <source>arXiv preprint arXiv:1702.08734</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.E.:
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>1097</volume>
          {
          <issue>1105</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Quan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miao</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>A combinatorial solution to point symbol recognition</article-title>
          .
          <source>Sensors</source>
          <volume>18</volume>
          (
          <issue>10</issue>
          ),
          <volume>3403</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Razavian</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azizpour</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sullivan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carlsson</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>CNN features o -the-shelf: an astounding baseline for recognition</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition Workshops (CVPRW)</source>
          ,
          <source>2014 IEEE Conference on</source>
          . pp.
          <volume>512</volume>
          {
          <fpage>519</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Redmon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farhadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Yolo9000: better, faster, stronger</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>7263</volume>
          {
          <issue>7271</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <string-name>
            <surname>Faster</surname>
          </string-name>
          r-cnn:
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>91</volume>
          {
          <issue>99</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tolias</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sicre</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jegou</surname>
          </string-name>
          , H.:
          <article-title>Particular object retrieval with integral maxpooling of cnn activations</article-title>
          .
          <source>arXiv preprint arXiv:1511.05879</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>