<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HTML Atomic UI Elements Extraction from Hand-Drawn Website Images using Mask-RCNN and novel Multi-Pass Inference Technique</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Prasang Gupta</string-name>
          <email>prasang.gupta@pwc.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Swayambodha Mohapatra</string-name>
          <email>swayambodha.mohapatra@pwc.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>PricewaterhouseCoopers US Advisory</institution>
          ,
          <addr-line>Mumbai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Website UI Design is an integral part of the world, but it is not trivial as there are a huge array of challenges that need to be conquered. A quintessential step of a website design process is to sketch the UI wireframe on paper and translating it into code later on. In an attempt to automate this process, advanced AI algorithms are explored in this study. The final approach comprises of image processing, followed by UI feature identification and localisation using Mask-RCNN and ultimately a novel Multi-Pass inference technique to boost the viability of the model. On the test dataset, the method resulted in an mAP or Mean Average Precision (IoU &gt; 0.5) value of 64.12 Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CLEF 2020, 22-25 September 2020, Thessaloniki, Greece.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>HTML</kwd>
        <kwd>UI</kwd>
        <kwd>Image Processing</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>OpenCV</kwd>
        <kwd>Mask-RCNN</kwd>
        <kwd>Multi-Pass Inference</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        With the world going increasingly global and starting to work virtually, websites
are more important than ever to expand business and reach out to customers.
However, website design requires a very specific set of skills. There are 2 major
ways of building a website. The first method is using visual website building
tools like Wix [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Constant Contact [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Squarespace [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], etc. and the second is
building by programming using languages like HTML, PHP, CSS, JavaScript,
etc. The downside of both of these approaches is that they have a very steep
learning curve.
      </p>
      <p>
        The ImageCLEF 2020 DrawnUI Task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] from ImageCLEF 2020 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is
formulated to reduce this dependency on the tools and flatten the learning curve
by enabling people to create websites using hand-drawn pictures of website
interfaces on whiteboard or a piece of paper. This would give a chance to people
having no knowledge of the aforementioned tools and languages to create
websites easily and quickly.
      </p>
      <p>The first step towards making this possible is to come up with a model that
correctly identifies the type and the position of various atomic user interface
(UI) elements in the wireframe drawing. This information can be leveraged to
generate a website layout using various heuristics. The next step to this problem
would be to convert this detected layout to code. In this study, we are focusing
on the first part of the problem.</p>
      <p>Diving into the details of the implementation, we will discuss about the
Dataset used for training the model in Section 2 and will cover the Methodology
used in Section 3. Further, we will discuss the Results in Section 4 and present
the Conclusions and any scope for Future Work in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Data Set</title>
      <p>
        ImageCLEF 2020 DrawnUI task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] was focused on extracting the atomic UI
elements from a hand-drawn image of a website. The dataset provided as part
of the challenge contained about 3000 hand-drawn images inspired from mobile
application screenshots and actual web pages containing about 1000 diferent
templates.
      </p>
      <p>The dataset was divided into two parts, the development set which contained
2363 annotated images and a test set containing 587 images which were not
annotated and were strictly for testing purposes. Each image in the development
set contained information in the form of a bounding box and a label for each UI
element present in that image. There were 21 diferent classes of labels present
as listed in Table 1.</p>
      <p>The images in the dataset were of varying sizes. All of them were RGB images
in JPEG format. The annotations were provided separately in a CSV file format.
The distribution of the width and the height of the images can be seen in Figure
1. Due to the varying sizes of the images, they need to be resized to a fixed size
that will be covered later in the modelling section.</p>
      <p>There were several challenges within the dataset. The first challenge was
that there were several repeated images in the development set which would not
Fig. 3: The image above has been taken at a very steep angle. This converts the
straight horizontal lines into diagonals and rectangles into parallelograms.
Fig. 4: The images shown above represent the overlap of diferent classes on the
”image” class. There were 2 types of overlap. The image on the left shows the
”under the image” overlap, while the image on the right shows the ”over the
image” overlap. Both of these are taken from the development dataset.
contribute much towards the training of the AI model. These images were not
exactly similar, but difered only on the basis of the background colour / tint of
the image. One of such examples of this can be seen in Figure 2.</p>
      <p>The second challenge was that there were some images with a very steep
capture angle in the dataset as shown in Figure 3. This had to be taken care of
as the apparent shapes of the UI elements would change drastically when the
images are captured at an angle.
paragradprohpdocwhnecrkabdoioxbuttonratingtoggtleextadraetaesptiecpkpeerrinputslidervideolabel table listheaderbuttonimalgineebrecaokntainer litnekxtinput</p>
      <sec id="sec-2-1">
        <title>Label Class</title>
        <p>paragradprohpdocwhnecrkabdoioxbuttonratingtoggtleextadraetaesptiecpkpeerrinputslidervideolabel table lishteaderbuttonimaligneebrecaokntainer litnekxtinput</p>
      </sec>
      <sec id="sec-2-2">
        <title>Label Class</title>
        <p>The third problem pertained to the ”image” class of the dataset. The image
class was defined as a rectangle with both the diagonals drawn. However, there
were several files containing multiple objects overlapping with an image class.
There were 2 kinds of overlap, the object over the image, where the diagonal
was ”hidden” or behind the image where the wireframe of the image would run
over these classes. This can be seen in Figure 4.</p>
        <p>Apart from these challenges, the number of labels were also skewed in the
dataset as some labels had plenty of representation, while some labels were
present quite rarely. The distribution for the labels can be seen in Figure 5. It
can be seen from the plot that the labels like ”button”, ”paragraph” and ”image”
are very commonly present while other labels like ”textarea”, ”stepperinput” and
”rating” are very sparse.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>Data Pre-Processing</title>
        <p>
          We have explored several pre-processing techniques to improve the viability of
our model. The majority of techniques are based on modifications of the data
using OpenCV [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] on C++. This was chosen to reduce the time taken to perform
operations and transformations on the images. Some of the techniques used are
described in this section.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Removing duplicate images</title>
        <p>By visual inspection of the dataset as mentioned in Data Set Section, we
found that there were repeated images (as can be seen in Figure 2) and thus,
these images would not contribute much to the training of the model. If the
Fig. 7: The image above is the result of the superposition of two candidate
images for similarity check. It can be seen that both of the them are exactly
similar but a little shifted. To gauge this, the distance between the images can
be calculated using the sum of squares of distances between the centre of the
corresponding label bounding boxes of both the images. The red line in the
image depicts the distance and the black dot is the center of the bounding box
for the corresponding labels in both the images
number is huge, then there may be a possibility that our model would start
overfitting to these similar images. Hence, to quantify the number of images
that are repeated in the dataset, we had to come up with some algorithm for
detecting the same. We will be testing the algorithms on the same set of images
as shown in Figure 2.</p>
        <p>One of the methods commonly used for checking if the images are equal or
not is the OpenCV subtract method. This method performs a pixel-by-pixel
subtraction of the images and returs an image. If the returned image is completely
black, then the starting images are same. We tried employing this method to
our images, but the results were not good as can be seen in Figure 6. This can
be owed to the fact that our similar images are not ”exactly” same as there are
some camera angle changes involved which change the orientation of the image
and hence, a pixel-by-pixel subtraction did not yield the best results here.</p>
        <p>The second approach used for finding duplicate images was an algorithm
based on finding the smallest distance between two given images. The algorithm
included making a list of size 21 (the total number of unique classes present in
the dataset) and then populating it for all the images with the number of the
classes of each type they have. This was iterated upon and all the image pairs
having the same class vector were found. The cartesian distance between these
selected pairs was calculated to verify if the images are actually similar or not
and if the distance was found to be lesser than a threshold, it was classified
as a repeated image pair. The way of calculating the distance between the two</p>
        <p>Fig. 8: The image shows the output of the DLIB model trained on just 15
instances. In this very limited learning, it has identified the general structure of
the image class with great accuracy.
images can be seen in Figure 7. Employing this algorithm, it was found that
there were only 1306 unique images in the development dataset out of the total
2363 images. Rest 1057 images were copies of the images already present.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Extracting individual elements from the image</title>
        <p>As the underlying shape of each of the label is same and only the localisation
of the labels vary across images, we tried extracting the individual elements from
the image. These extractions can then later be used for training purposes. Also,
these can be used for increasing the number of the classes whose frequency is
less in the dataset. This would allow the model to learn the features of the lesser
frequent label types as well.</p>
        <p>
          The approach selected for this was using a DLIB [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] model to capture the
general features of the class labels. The DLIB model was chosen as it could
learn the basic structure of the class with very few learning data, as would be
the case with the classes having very low representation. A sample DLIB model
output for the image class can be seen in Figure 8. The DLIB models were able
to identify most of the label classes, but were unable to segregate the image
wherever there was overlapping present.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Converting images to Grayscale</title>
        <p>The images provided in the dataset were all 3 channel RGB images. 3 channels
might be helpful in problems where the information carried by the colour is
needed to be learned by the model and should be used as a feature. But, in our
case, colour doesn’t matter as we have to detect the features only on the basis of
shape. Hence, to prevent throwing of the learning of the model by introducing
Fig. 9: The image on the left is the original image while the one on the right is
after the grayscale conversion.</p>
        <p>Fig. 10: The image on the left is the converted grayscale image while the one on
the right is a more refined sharpened grayscale image.
colour, all the images were converted to grayscale using OpenCV. A sample
grayscale conversion can be seen in Figure 9.</p>
        <p>Another factor that is helpful in grayscale images is that the number of
channels are reduced to 1. Hence, the efective size of the image reduces which
leads to a speed up in computation. To increase the visibility of the labels further,
the grayscale images were later sharpened using OpenCV. This can be seen in
Figure 10.</p>
        <p>Fig. 11: The image on the left is the grayscale image while the one on the right
is the result of a simple thresholding conversion.</p>
        <p>Fig. 12: The image on the left is the grayscale image while the one on the right
is after applying Otsu’s binarization algorithm on it</p>
      </sec>
      <sec id="sec-3-5">
        <title>Converting images to Black &amp; White</title>
        <p>There were several algorithms applied to convert the image from a grayscale
image to a Black and White Image. This was carried out to further reduce the
efect of the background elements on the model prediction, as grayscale also
carries information regarding the shade of the image or background.</p>
        <p>The first approach used to convert the grayscale images to black and white
was simple binary thresholding. But, the limitation of the model was that every
image had a diferent optimum threshold and there was no way to find it
beforehand. Hence, there was a lot of loss of information by this conversion as can be
seen in Figure 11.</p>
        <p>
          The second approach used is Otsu’s binarization algorithm [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The Otsu’s
algorithm finds a threshold for the image automatically based on its histogram
distribution. The image generated using this is shown in Figure 12. Formally,
Otsu’s algorithm tries to find a threshold value (t) which minimizes the weighted
within-class variance given by the following relation.
        </p>
        <p>σw2(t) = q1(t)σ12(t) + q2(t)σ22(t)
(1)
µ1(t) = ∑t iP (i) &amp; µ2(t) =
i=1 q1(t)</p>
        <p>I
∑ iP (i)
i=t+1 q2(t)
i=1</p>
        <p>The final approach used is an adaptive approach where a single threshold is
not applied globally to the dataset. The threshold value in this case is a
Gaussianweighted sum of the neighbourhood values minus a constant. The results using
this were good, but contained a lot of noise as can be seen in Figure 14. However,
this was by far the best conversion of grayscale to binary black and white. Hence,
this was selected as the final model and the noise was tackled by finding and
removing the small connected components of the image using C++. The final
image can be seen in Figure 15.</p>
        <p>Fig. 16: General Architecture of the Mask RCNN Model. Reproduced from</p>
        <p>
          Mask-RCNN [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Fig. 17: Output generated by Mask RCNN Model (Run 1) on one of the images
belonging to the test split of the dataset.
3.2</p>
      </sec>
      <sec id="sec-3-6">
        <title>Methods Implemented</title>
        <p>
          The images were first all transformed into single channel black and white images
and then resized into 1024*1024*1, as required by Mask RCNN Architecture [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
The general architecture of the Mask RCNN model can be found in Figure 16.
The dataset was split in the ratio 80:20 with the larger ratio corresponding to the
training set and the smaller one corresponding to the validation set. Since there
were a few classes that had a small set of images corresponding to them, care was
taken to ensure that such images were present in the same ratio while splitting
the dataset. The models were trained on a virtual Ubuntu server equipped with
a 16GB NVIDIA Tesla V100 GPU Accelerator [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] hosted on PwC’s proprietary
cloud platform, Workbench.
        </p>
        <p>
          Since this challenge mainly involved detection of small website UI elements,
Mask RCNN was chosen because it performs better than other models in
object detection. Mask RCNN model generates bounding boxes and segmentation
masks for each instance of an object in the image, and is based on Feature
Pyramid Network (FPN) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and a ResNet-101 backbone [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          We implemented Transfer Learning by using a pre-trained Mask RCNN
model trained on COCO Dataset [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Even though the images contained in
the COCO Dataset are not very similar to our dataset, we used it to ensure
that our model extracts the high-level features in all images. The ’heads’ layer
of the model was then trained for 200 epochs at a Learning Rate of 10−3. The
convergence of the model can be seen in Figure 18.
        </p>
        <p>The output from this model, obtained on an image from the test split of the
dataset can be seen in Figure 17. The model performed really well in recognising
all the major UI elements in the image. But, the problem was that it was not
able to detect smaller UI elements in the image.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Run 2 : Mask RCNN Model with novel Multi-Pass Inference Technique</title>
        <p>After evaluation, even though the overall precision score of Run 1 was high,
the overall recall score was not as high. This meant that the model was not able
to detect the smaller UI elements on the image. To improve the previous run,
we implemented a novel Multi-Pass Inference Technique.</p>
        <p>The novel Multi-Pass Inference Technique involves getting the predictions on
the input image and then filling the corresponding bounding box regions with
the background colour (white in this case). The edited image is then passed
Fig. 19: The image on the top left is obtained after passing the image through
the model once (Generally used single pass inference). The bounding boxes in
this image were filled with white (except the classes where overlapping is
happening) and then passed through the model again. The image on the top
right is the output of the 2nd pass. Both of the above images are combined
together on the basis of confidence scores and IoU overlaps to form the final
prediction image on the bottom. The blue bounding boxes in the final image
are from Pass 1 while the red boxes are from Pass 2.
again through the model to essentially ’force’ the model to make predictions
on the missed out elements. The new predictions are appended to the earlier
predictions to get the final results for a particular image. This technique can be
visualised in Figure 19. It can be observed from the figure that there are several
UI elements that have been missed in Part 1, but predicted in Part 2 making the
ifnal output contain most of the UI elements present on the page. The number
of times the edited image is passed can be varied according to the problem in
hand.</p>
        <p>Run 2: Training Loss &amp; Val Loss by Epoch</p>
        <p>For this particular run, the ’heads’ layer of the model was trained for 100
epochs and the following Learning Scheduler was used to ensure the model
converges quickly - Learning Rate of 10−2 for the first 25 epochs, Learning Rate
of 10−3 for the next 25 epochs, Learning Rate of 10−4 for the next 25 epochs,
Learning Rate of 10−5 for the last 25 epochs. The convergence of the model can
be seen in Figure 20.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Run 3 : Modified Version of Run-2</title>
        <p>This run was implemented to improve the results obtained from the previous
run. The same Multi-Pass Inference technique was implemented with a slight
modification to ensure that only the bounding boxes with the highest
confidence scores after the second pass were added to the final results of each image.
Fig. 22: This image shows the intermediate output generated by Mask RCNN</p>
        <p>Model with Multi-Pass Inference Technique (Run 3) on one of the images
belonging to the test split of the dataset after the first pass. A few smaller UI
elements are missed out by the model.</p>
        <p>Fig. 23: This image shows the final output generated by Mask RCNN Model
with Multi-Pass Inference Technique (Run 3) on one of the images belonging to
the test split of the dataset after the second pass. Most of the missed elements
in the first pass are captured in the second pass.</p>
        <p>This was done to ensure that the stray elements detected after the white space
replacement step are not added to the final results.</p>
        <p>For this particular run, the ’heads’ layer of the model was trained for 125
epochs and the following Learning Scheduler was used to ensure the model
converges quickly - Learning Rate of 10−2 for the first 25 epochs, Learning Rate of
5 ∗ 10−3 for the next 25 epochs, Learning Rate of 10−3 for the next 25 epochs,
Learning Rate of 2 ∗ 10−4 for the next 25 epochs, Learning Rate of 10−4 for the
last 25 epochs. The convergence of the model can be seen in Figure 21.</p>
        <p>The intermediate output and final output generated from this model, on an
image from the test split of the dataset can be seen in Figure 22 and 23. There
was a visible improvement in recognising smaller UI elements on the image which
was also reflected in better Mean Average Precision (mAP) scores as listed in
the Table 2.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The predictions on the test set images were collated in a csv file. For each image
on the test set, the bounding boxes corresponding to each instance of a detected
class and the confidence scores were submitted. The Mean Average Precision
(mAP) scores obtained across the three runs can be found as listed in Table 2.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>Throughout the challenge, we experimented with several processing techniques
to get the data in the best shape to be trained. We selected Mask R-CNN as
our baseline model as it is known to perform well on Object Detection
problems and this problem statement was not much diferent. We also came up
with a novel technique, Multi-Pass Inference, which improved the mAP score
drastically, hence gaining us the 3rd spot on the leaderboard of the DrawnUI
challenge.</p>
      <p>Due to the lack of time, we could not tinker around much with the models
as the training takes up a lot of time, being computationally expensive. In the
future, we can explore other models as the baseline model which have better
performance over Mask R-CNN. One such example of an improved model would
be EfficientDet, which is known to perform much better, but is deadly slow
to train. Also, there is a lot of scope in expanding the viability of the novel
Multi-Pass Inference technique and study the afect of number of passes with
performance. There is also scope for experimenting with attention mechanism
to focus on those parts of the image which are actually important.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Constant contact. https://www.constantcontact.com/website/ (
          <year>July 2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Squarespace. https://www.squarespace.
          <source>com (July</source>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Wix. https://www.wix.
          <source>com (July</source>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bradski</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The OpenCV Library</article-title>
          .
          <source>Dr. Dobb's Journal of Software Tools</source>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fichou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berari</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ştefan</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of ImageCLEFdrawnUI 2020: The Detection and Recognition of Hand Drawn Website UIs Task</article-title>
          .
          <source>In: CLEF2020 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org &lt;http://ceur-ws.
          <source>org&gt;</source>
          , Thessaloniki,
          <source>Greece (September</source>
          <volume>22</volume>
          -25
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woods</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eddins</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          :
          <article-title>Digital image processing using MATLAB. Pearson Education India (</article-title>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkioxari</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollár</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
          </string-name>
          , R.:
          <string-name>
            <surname>Mask</surname>
          </string-name>
          r-cnn.
          <source>In: 2017 IEEE International Conference on Computer Vision</source>
          (ICCV). pp.
          <fpage>2980</fpage>
          -
          <lpage>2988</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition (</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          :
          <article-title>Matplotlib: A 2d graphics environment</article-title>
          .
          <source>Computing in Science &amp; Engineering</source>
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <fpage>90</fpage>
          -
          <lpage>95</lpage>
          (
          <year>2007</year>
          ). https://doi.org/10.1109/
          <string-name>
            <surname>MCSE</surname>
          </string-name>
          .
          <year>2007</year>
          .55
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Péteri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datla</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DemnerFushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlovski</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liauchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Herrera</surname>
            ,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ninh</surname>
            ,
            <given-names>V.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , l Halvorsen,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Lux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Dang-Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.T.</given-names>
            ,
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Fichou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Berari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Brie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Dogariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ştefan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.D.</given-names>
            ,
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.G.</surname>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2020: Multimedia retrieval in lifelogging, medical, nature, and internet applications</article-title>
          .
          <source>In: Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the 11th International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), vol.
          <volume>12260</volume>
          .
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Thessaloniki,
          <source>Greece (September</source>
          <volume>22</volume>
          - 25
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>King</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          :
          <article-title>Dlib-ml: A machine learning toolkit</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>10</volume>
          ,
          <fpage>1755</fpage>
          -
          <lpage>1758</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollár</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hariharan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Feature pyramid networks for object detection</article-title>
          .
          <source>In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          . pp.
          <fpage>936</fpage>
          -
          <lpage>944</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bourdev</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perona</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramanan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollár</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Microsoft coco: Common objects in context (</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lindholm</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nickolls</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oberman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montrym</surname>
          </string-name>
          , J.:
          <article-title>Nvidia tesla: A unified graphics and computing architecture</article-title>
          .
          <source>IEEE Micro</source>
          <volume>28</volume>
          (
          <issue>2</issue>
          ),
          <fpage>39</fpage>
          -
          <lpage>55</lpage>
          (
          <year>Mar 2008</year>
          ). https://doi.org/10.1109/MM.
          <year>2008</year>
          .
          <volume>31</volume>
          , https://doi.org/10.1109/MM.
          <year>2008</year>
          .31
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>