<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>The physical
internet, Open Engineering</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1515/eng-2016-0073</article-id>
      <title-group>
        <article-title>Using Camera-Drones and Artificial Intelligence to Automate Warehouse Inventory</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>René Kessler</string-name>
          <email>rene.kessler@uol.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Melching</string-name>
          <email>christian.melching@uol.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralph Goehrs</string-name>
          <email>ralph.goehrs@abat.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Marx Gómez</string-name>
          <email>jorge.marx.gomez@uol.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>In A. Martin, K. Hinkelmann</institution>
          ,
          <addr-line>H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.)</addr-line>
          ,
          <institution>Proceedings of the AAAI 2021 Spring Symposium on Combining Machine Learning and Knowledge Engineering (AAAI-MAKE 2021) - Stanford University</institution>
          ,
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Oldenburg</institution>
          ,
          <addr-line>Ammerländer Heerstr. 114-118, Oldenburg, 26129</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>abat AG</institution>
          ,
          <addr-line>An der Reeperbahn 10, Bremen, 28127</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>6</volume>
      <issue>2016</issue>
      <fpage>33</fpage>
      <lpage>50</lpage>
      <abstract>
        <p>Inventory is a very important, but also a very time-consuming manual process in warehouse logistics. This paper presents an approach to automate manual inventory using a camera-drone and various AI procedures. Thereby, sensor technology, such as RFID, is avoided, and only the visual representation of the products and goods is used. We developed a custom dataset that was used for the training of an object detection model to extract and count all relevant objects based on an image of the warehouse. Furthermore, we can show that diferent pre-processing steps and especially image augmentation methods can significantly influence the performance of such models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;inventory</kwd>
        <kwd>logistics</kwd>
        <kwd>artificial intelligence</kwd>
        <kwd>object detection</kwd>
        <kwd>drones</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Motivation and problem statement</title>
      <p>
        Logistics plays a crucial role in today’s global economy - every company depends on reliable
and intact logistics processes to create value. At the same time, logistics is subject to very
high margin pressure. On the one hand, improved service is expected from logistics service
providers, but on the other hand, customers do not want to to pay extra for it [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. If the
margin is to be maintained or even increased, this can only be achieved through savings in
internal processes. For many companies, digital transformation and the use of so-called smart
technologies can be the solution for optimizing processes and procedures [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3, 4, 5</xref>
        ], since many
logistics processes still involve a lot of manual efort, including inventory, which is essential
for every company [6]. In practice and science, this is often referred to as Industry 4.0. Four
main design principles apply to use cases in this area: Networking, information transparency,
technical assistance and decentralized decisions [7]. The approach pursued in this work can
also be classified in these principles. Through the combined use of camera drones and methods
for processing the data, such as AI, human abilities can be imitated: The processing of visual
signals. The result is that previously manual activities, such as the inventory process in a
warehouse, could be automated.
      </p>
      <p>The goal of the inventory is to record the type and number of internal goods and to quantify
this inventory with an exact value [6]. The actual process of inventory can difer not only in
company-specific factors but also in the fact that two types of inventories are common. Thus,
a distinction can be made between physical inventory, i.e. counting physical goods, and the
non-physical inventory, where, for example, financial goods or bank balances are recorded. In
this paper, further focus will be on physical inventory. According to the German Commercial
Code (HGB), German companies are obliged to carry out an inventory atleast once a year1
(paragraph 240 German Commercial Code). While this cycle may be suficient for companies
with little movement of goods, it makes sense to keep shorter cycles especially for companies
in the retail sector. However, the inventory can also be understood as an instrument of quality
management concerning transparency in order to be able to monitor the business goals and
their achievement. With the help of an inventory, deficits, faulty processes, non-optimal flows,
or even theft can be detected. During an inventory, there is always a large amount of personnel
efort involved. Depending on the inventory’s size and scope, several people are often
exclusively occupied with the manual counting of the goods. Inventory not only causes personnel
costs, but also disrupts operational processes. Therefore, companies find themselves in a
dichotomy between transparency and costs, which is why they often work with samples whose
ifndings can then be extrapolated to the entire inventory. This business conflict can be resolved
by automating the manual localization and identification of products and goods during
inventory, resulting in greater transparency at lower costs [8]. Existing approaches are based on the
use of sensor technology and often only consider a very specific sub-area (e.g. industries). In
this paper, a generic approach is followed, which exclusively involves the visual representation
of the products and is based on deep learning methods. To also automate the acquisition of the
images and thus to evaluate a vehicle for the operationalization of the approach, a drone is also
used, which has already proven itself in comparable applications [9, 4, 10, 11]. Combined with
the current problems in inventory processes described above and the great potential through
automation of these process steps, this leads to the research question:</p>
      <p>Which AI procedures are suitable to recognize products on images in order to count
them for an inventory?</p>
      <p>To solve the described problem, a data-driven approach was followed, which was oriented
towards the established CRISP-DM model [12, 13]. During implementation, the company
cooperated intensively with a North German beverage distributor, which has several storage
locations and a high volume of trade in the B2B sector. Initially, domain knowledge about the
storage locations and the inventory process was collected in several workshops and
discussions. Also, we were granted access to the diferent warehouses in order to record the custom
dataset (3). Finally, the results of the experiments were presented to the practice partner, and
practical implications were derived.</p>
      <p>The paper is structured as follows: In the next section, Related Work (2), an overview of
the current research work is given. Section three, Dataset (3), is dedicated to the structure of
the specially compiled dataset and its annotations. The fourth section, Experiment and Results
1https://www.gesetze-im-internet.de/hgb/__240.html, accessed on 20.11.2020
(4), represents the core of the work and describes the methods used and their results. In the
concluding section, Discussion and Future Work (5), the results are summarized, and an outlook
on further work is given.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>A search for related work has led to identifying numerous publications that deal with drones
in logistics and use of digital technologies to optimize or even automate the inventory process.
Two main lines of research have been identified and are presented below.</p>
      <p>
        Locating and identification utilizing sensor technology: Radio-Frequency Identification
(RFID) tags are a widespread and established method to track products and load carriers in
a warehouse [14, 15, 16], and to achieve an increase in transparency regarding warehouse
movements [17]. Often RFID readers are attached to a drone, and the drone flies over storage
locations and tracks the individual goods [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">18, 15, 19, 9, 3, 20, 8</xref>
        ]. Drones can detect storage
locations that would otherwise be dificult to reach (e.g., very high storage locations in a high
rack) and minimize the risk for the employees [
        <xref ref-type="bibr" rid="ref7">16, 21</xref>
        ]. In summary, the identification of loads
using RFID, especially in combination with a drone, has proven to save costs and automate
inventory processes [17]. However, this always presupposes that the loads are equipped with
RFID tags, representing a further cost factor in logistics.
      </p>
      <p>
        Reading optical product features or characteristics: There is a necessary consensus that
image processing in logistics can be a great advantage for the traceability and monitoring of
goods [
        <xref ref-type="bibr" rid="ref8">22</xref>
        ]. If camera-drones are used in inventory management, they are often used as a
medium for reading optical product annotations, such as one-dimensional barcodes or
QRCodes [
        <xref ref-type="bibr" rid="ref10 ref9">23, 24, 8</xref>
        ]. However, these approaches assume that those annotations exist. Considering
packed products, barcodes are often available and do not pose a problem. But with unpackaged
goods or empties, optical annotations are rarely present. Here is a need for solutions that focus
on the optical representation of goods. Although both AI-based image processing and the use
of drones are considered to have great potential in logistics [
        <xref ref-type="bibr" rid="ref11">25, 10, 8</xref>
        ], there are hardly any
publications available that combine these two approaches. Freistetter and Hummel (2019) outlined
an approach to drone-based inventory in libraries. They flew of bookshelves and identified
book spines using computer vision techniques. As soon as a book is in the center of the image,
the title of the book is read [
        <xref ref-type="bibr" rid="ref12">26</xref>
        ]. This is a particular indoor use case. Despite the
fundamental similarity, a use case from a library cannot necessarily be transferred to larger industrial
warehouses. Especially concerning disruptive factors (e.g. environmental influences, such as
changing weather, which can lead to diferent lighting), recordings in libraries are less afected.
Dörr et al. show an approach that deals with a similar use case in the warehouse area. The goal
of the approach is product structure recognition based on image data. Diferent convolutional
neural networks are used based on top of each other. For the training of the models a separate
dataset was built [
        <xref ref-type="bibr" rid="ref13">27</xref>
        ]. The very recent publications show that the combination of drone-based
inventory and AI image processing is currently subject of research [16].
      </p>
      <p>
        Especially the approach of Dörr et al. shows many similarities to our approach presented
here, e.g., the hierarchical structure with several deep learning models [
        <xref ref-type="bibr" rid="ref13">27</xref>
        ]. However, there
are also several diferences. The approaches difer in their place of use since this thesis is an
outdoor use case. Also, the type of objects to be identified difers.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>The use case considered here focuses on the automated recognition of beverage pallets on
images. Since there is no public dataset for this specific problem, we developed a custom dataset.
The structure and further processing of the dataset is described in the following.
Data Acquisition: A drone of the model DJI Phantom pro 4 v22 was used to make the video
recordings in the warehouses of the practice partner. The drone was selected because it was
already available for the research project, so there was no need to purchase a new drone. In
addition, the characteristics of this drone are very common, which is why the findings are equally
transferable to other drone models. The video recordings were made in Full-HD resolution, a
frame rate of 30 frames per second and using the built-in image stabilization. To create the
greatest possible variance in the data, the recordings were made over four days and at two
locations of the beverage dealer. The aim was to record the weather’s influence on the images
(e.g., brightness) by recording the images under diferent conditions. During the project, two
days with sunny weather and one day each with cloudy and very cloudy weather were used for
the recording. Special attention was also paid to the recorded scenes. We tried to capture all
pallet locations of the outdoor warehouse. It could not be avoided that not all types of pallets
(e.g., diferent manufacturers of beverages) are represented equally often in the data because
the diferent pallets’ stock varies very much in reality.</p>
      <p>After finishing the video recordings, the video data had to be processed. For this purpose, the
iflmed sequences were viewed manually, and irrelevant scenes were removed (e.g., the drone’s
starting and landing sequences). Since there are only marginal diferences in subsequent frames
at 30 frames per second, only every 30th frame of the video clips was transferred to the final
data set when the training data set was created to avoid potential overfitting when training
the neural network. After the frames’ automated extraction, they were manually sighted, and
faulty or low-quality images were removed. The final dataset consists of 336 images, separated
into train and test set. The test set only contains images from pallet stacks that are not included
in the train set to avoid memorization. This includes pallets with beverages of previously
un2https://www.dji.com/de/phantom-4-pro-v2/specs
seen brands and breweries.</p>
      <p>
        Data Annotation: After image acquisition, annotation has been applied using the tool
labelstudio from Heartex3. This tool was used because it ofers many possibilities for annotating
images and was already used in another context within the research project. In this process,
each pallet that was largely visible was annotated using polygons instead of only bounding
boxes to make use of the annotation masks later on. Additionally, each polygon was given a
class to diferentiate between two types of pallets, pallets containing cases of beer and pallets
containing other beverages. The resulting annotations were exported and converted to the
COCO dataset format [
        <xref ref-type="bibr" rid="ref14">28</xref>
        ], as it is one of the standard formats for object detection and
segmentation in images that are widely supported by most frameworks. This results in a training
set containing 284 images with 5261 annotated polygons and a test set containing 52 images
and 1471 polygons.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment and Results</title>
      <p>The experiment was conducted as follows. First a baseline model was used to test the impact
of various modifications of the input data on the prediction accuracy. Then, several models
using diferent architectures, selected based on defined criteria, were trained to evaluate their
performance applying the identified modifications using the baseline model. Finally, the best
performing model was used to perform a qualitative evaluation and to identify possible errors.</p>
      <sec id="sec-4-1">
        <title>4.1. Baseline Model</title>
        <p>
          The model used during the following experiments is a Mask R-CNN using a ResNet50
Backbone, implemented using Detectron2 [
          <xref ref-type="bibr" rid="ref15 ref16">29, 30</xref>
          ]. The model was pre-trained using the MSCOCO17
dataset to try to make up for the low amount of images in the used dataset, described in 3. It
was then trained over up to 4000 iterations, where each iteration used a batch of twelve images.
Evaluations of the model were performed every 250 iterations during training, and once the
training was completed. During training and testing, the images were resized to 1000x750 and
not further modified. To validate each experiment’s results, it was repeated multiple times; the
following metrics are averages over all runs. For all experiments the same train-test-split was
used.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Initial Results and Adjustments</title>
        <p>The initial baseline model achieved an average precision of 27.06 and 27.59 evaluating bounding
boxes and segmentation respectively, using the test portion of the dataset. As these results are
far from suficient to predict the pallets’ position on images accurately, significant adjustments
were necessary to improve the performance of the model.</p>
        <sec id="sec-4-2-1">
          <title>4.2.1. Merging of classes</title>
          <p>During the first tests, it became clear that the model had dificulties in classifying the pallets
given the annotated classes (beer and other beverages). Not only was the per-class-precision
much higher on the pallets marked as containing cases of beer (33.647 compared to 20.463,
evaluating bounding boxes), some pallets were often wrongly classified. While this problem’s
origin likely lies within the training data that contained more annotations and variants of
pallets of beer cases, it is challenging to balance the classes to reduce this spread since nearly all
images contain pallets of both categories. The classes have been merged to create a model to
predict only the bounding box and segmentation mask of a pallet, not further classifying its
contents to circumvent this problem. If applied like this, the classification task must be
processed by a diferent system, possibly using classifier or brand detectors. This work has not yet
further pursued the creation of such a solution.</p>
          <p>A model trained on a classless dataset achieves an average precision of 45.95 and 46.70 on
bounding boxes and masks, as noted in table 1. A following manual inspection of the
predictions also confirmed the increase of the quantity and quality of the predictions.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Reduction of image area</title>
          <p>Another problem of the model was the detection of small pallets on the edges of the images.
It is likely caused by the low amount of small objects in the training data and distortion of the
camera lens on the edges of the image. This problem can be addressed by cutting of parts on
the left and right edges of the recordings. This should not harm the quality of the inventory
process as the drone flight is planned to fly so that each row is at least once near the center
of the recorded images and therefore not lost in this process. An example of such reduction of
the image contents is visible in figure 2 where 50% of the image have been removed, in equal
parts on each side. During the training and evaluation process, the input images were resized
to 750x750 instead of 1000x750, to better match the aspect ratio of the modified images. The
resulting model achieves values far better than before the removal. This is expected as the
problem is simplified significantly. The model achieves a mean Average Precision (mAP) of
47.68 and 46.70 evaluating predicted bounding boxes and segmentation, respectively.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>4.2.3. Image augmentation</title>
          <p>
            Image augmentation is widely used in many research projects [
            <xref ref-type="bibr" rid="ref17">31, 32, 33</xref>
            ], but can have
diferent efects depending on the problem [
            <xref ref-type="bibr" rid="ref13">27</xref>
            ]. Therefore, diferent image augmentation methods
were tested and evaluated based on the performance of the models trained using augmented
images. The augmentation of the original images was applied using imgaug4 during the
loading of the image batch of each iteration. Each image of the batch was augmented individually
with random parameters in given boundaries. As imgaug ofers a wide variety of methods to
augment images, some were selected to evaluate its performance on the given problem. In this
case we chose augmentations to simulate variations of real recordings, such as MotionBlur,
PerspectiveTransformation, Contrast and JpegCompression, to accommodate for movement of the
drone, varying lighting and weather conditions, while also considering methods used
traditionally to adapt the image, Rotation, ScaleXY, FlipLR and CropAndPad. The following methods
were chosen for comparison.
          </p>
          <p>(a) FlipLR - Performs a horizontal flip with a probability of 0.5.
(b) ScaleXY{150,125} - Scales the width and the height of the images with using random
values from [0.5, 1.5] for ScaleXY150 and [0.75, 1.25] for ScaleXY125.
(c) Rotate{10,20,30} - Rotates the image around its center using random degrees up to 10 for</p>
          <p>Rotate10, 20 for Rotate20 and 30 for Rotate30. Rotations can applied in both directions.
(d) Contrast - Increases or decreases the contrast of the image using random values from
[0.5, 1.4].
(e) JpegCompression - Reduce the quality of the image by applying Jpeg compression
using a random degree from [0.7, 0.95].
(f) MotionBlur - Creates a motion blur efect with a kernel of size 7x7.
(g) CropPad{25,50} - Remove random percentage of images from all edges and pad the image
to its original size. Using random values from [-0.25, 0.25] for CropAndPad25 and from
[-0.5, 0.5] for CropAndPad50.
(h) PerspectiveTransform - Transforms the image, as if the camera had a diferent
perspective using random scales from [0.01, 0.1].
(i) FlipLR-ScaleXY150 - Combines FlipLR and ScaleXY150.
(j) CropPad25-ScaleXY150 - Combines CropAndPad25 and ScaleXY150.</p>
          <p>(k) Rotate20-ScaleXY150 - Combines Rotate20 and ScaleXY150.</p>
          <p>An overview of the efect of these augmentations is displayed in figure 3. Some of the tested
augmentations were omitted in the figure as their efects are barely visible due to the size of
the individual images or are variations or combinations of already displayed efects.</p>
          <p>It has to be noted that after the application of one or multiple augmentations the bounding
boxes were recomputed according to the transformed mask, to make them a minimal fit to
the object again. Otherwise, some augmentations, such as Rotation, could create a bounding
box according to the transformed bounding box, which could be too large to accurately locate
the object. In some situations, this results in a small, but measurable, boost of performance of
models using afected augmentations.</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>4.2.4. Baseline results</title>
          <p>To evaluate the diferent possible modifications, we trained several models using the same
settings and evaluated them on the shared test set. For the evaluation of the diferent methods,
we compared both the mAP of the predicted bounding boxes and the mAP of the segmentations,
even though they are very similar, displayed in table 1.</p>
          <p>The best performing model without the use of image augmentation was the model trained
using the simplest variation of the images, utilizing the cutting of edges and merging of
diferent classes. It achieved a precision of more than 70 and therefore performs better than most
other models, including many models trained using image augmentation. Once image
augmentation (Table 1) is considered, the model trained on unaugmented images is outperformed by
several diferent models. While nearly all models using merged classes and cut images perform
better, only a few models using a diferent image base produce comparable or better results.</p>
          <p>Especially interesting are the efects of specific augmentation methods. While the Rotation
augmentation decreased the accuracy of models using images with uncut edges, it increased
the precision on images scaled to a 1:1 aspect ratio. Some methods seem almost always to
reduce the prediction quality, such as Contrast (d), JpegCompression (e), and MotionBlur (f).
While the idea behind using these methods was to make images slightly more corrupt to
increase the ability to learn from realistic variations of these images, it mostly hurt performance.
Other augmentation methods used, such as ScaleXY (b), FlipLR (a), and CropPad (g), seem to
always improve the results of the trained model when used alone or in combination with other
methods, contrary to the observation by Dörr et al.. This is supported by the fact that the
almost always best-performing method used the combination of ScaleXY (b) and FlipLR (a).</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Evaluation of other architectures</title>
        <p>While the results provided by the diferent Mask R-CNN models certainly provide valuable
information, the model itself is no longer state-of-the-art in terms of precision. Therefore we
selected three diferent models and tested their performance using the results gained using
the previous model. The first model we additionally tested is DetectoRS [34], whose
innovative characteristic is the use of Recursive Feature Pyramids. It achieves near state-of-the-art
performance on the MSCOCO17 dataset and was implemented and trained using
MMDetecion [35]. The second model selected is Yolact [36]. Yolact is an architecture that is able to
generate predictions of the recordings in near real time and was trained using the
MMDetection framework aswell. While real time predictions are not necessary during the stocktaking
process, Yolact was selected since these models could easily be used to serve diferent purposes
within the same domain. The third and last model evaluated is DETR [37] due to it’s
innovative approach. DETR utilizes the transformer architecture introduced in the domain of NLP to
generate instance segmentations. It was chosen to evaluate whether or not new and innovative
approaches can be applied to the domain of palettes, and trained using the code provided by
the authors5.
4.3.1. Results
Each of the additional models was tested and evaluated on the test dataset, using the merged
and cut variant and the augmentation method (i) that showed to increase performance the
most. The results are displayed in table 2. It is clear that DetectoRS outperforms all other
models by a significant margin in terms of precision. It achieves a mAP of 88.3 and 85.6 while
taking 0.13 s per frame using a NVIDIA® GeForce® RTX 2080 Ti, making it also slower than
all other models. In contrast, Yolact achieves the lowest mAP and, contrary to expectations,
5https://github.com/facebookresearch/detr
is not the fastest model, but is 0.01 s slower than Mask R-CNN. However, Yolact is by far the
fastest model when using a CPU. DETR, with it’s new approach, is both both slower and less
precise than Mask R-CNN and does not stand out in any metric.</p>
        <p>In addition to evaluation based on metrics, timings, and model size, a manual qualitative
evaluation was performed. This showed that the mAP value was consistent with the visual
impression. In terms of both bounding boxes and segmentation, DetectoRS provides the best
results. Mask R-CNN also delivers satisfactory results, while the quality of DETR and Yolact in
particular falls of sharply.</p>
        <p>To get an impression of the quality of the predictions of DetectoRS, some of its predictions
are visualized in figure 4. The generated predictions are predominantly of high to very high
quality. In a few cases, however, there are (partly) incorrect predictions (figure 5). Three typical
errors can be described as follows:
1. Recognition of side views of pallets: Despite the label strategy and the pre-processing
steps, in the case of images with a very specific acquisition angle, namely whenever
the side views occupy a large image area, the isolated, incorrect identification of pallets
occurs, in which side views are provided with a bounding box and mask.
2. Individual pallets are not recognized: The evaluation has shown that in rare cases
individual pallets are not recognized. The special feature here is that this error always
refers only to a maximum of two pallets standing next to each other. All other pallets
on these images were recognized completely and without errors. Further optimization
of the detector parameters (e.g., tresholding) will most likely solve this error.
3. Strongly overlapping bounding boxes: Pallets with boxes of diferent colors
sometimes have overlapping bounding boxes. This has no consequences for the pallet
recognition, but it could lead to problems in subsequent steps, such as the classification of the
pallets. To solve this problem, further methods could be used in the preprocessing of the
images.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Future Work</title>
      <p>
        In this work, we present an initial step for automatized inventory using images recorded by a
drone and an AI based object detector to identify the location of pallets on recorded images.
Various modifications have been tested to increase the accuracy of predicted bounding boxes
and segmentation masks, partly without compromising the quality and direct usability of the
results. In doing so, we also showed that image augmentation methods could increase the
precision of models significantly, contrary to the observations in related projects [
        <xref ref-type="bibr" rid="ref13">27</xref>
        ]. In
summary, it has been shown that horizontal flipping and image scaling as an image augmentation
technique can have a positive impact on performance during model training. In particular, the
model architecture DetectoRS showed very good results in the experiments, measured by mAP
(bounding boxes and segmentation), although not being the fastest in comparison. While
already delivering promising results, there are certain factors limiting the usage of the developed
models in practice.
      </p>
      <p>Firstly, a solution for the classification of the type or class of beverages on the pallets must
be developed and integrated. While it would have been preferable to predict its class using
the same model that also predicts its location, tests showed it impacted its localization
performance significantly. Secondly, while the detection of the front of the pallets gives valuable
information to automate the inventory process, additional information is needed to complete
it. This mainly includes the length-wise number of stacks of pallets in their row, which could
be recorded from above. Finally, regardless of the technical implementation, discussions and
tests with practical partners have shown that the organization in the warehouse must also be
changed if drones are to be used. Even though drones are very flexible and can reach storage
areas that are dificult for humans to reach, processes in the warehouse must be adapted to
ensure that drones can be used safely and eficiently [38].
[32] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, A. C. Berg, SSD: single
shot multibox detector, CoRR abs/1512.02325 (2015). URL: http://arxiv.org/abs/1512.02325.
arXiv:1512.02325.
[33] X. Chen, R. Girshick, K. He, P. Dollár, Tensormask: A foundation for dense object
segmentation, 2019. arXiv:1903.12174.
[34] S. Qiao, L.-C. Chen, A. Yuille, Detectors: Detecting objects with recursive feature pyramid
and switchable atrous convolution, 2020. arXiv:2006.02334.
[35] K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang,
D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi,
W. Ouyang, C. C. Loy, D. Lin, MMDetection: Open mmlab detection toolbox and
benchmark, arXiv preprint arXiv:1906.07155 (2019).
[36] D. Bolya, C. Zhou, F. Xiao, Y. J. Lee, Yolact: Real-time instance segmentation, in: ICCV,
2019.
[37] M. Zheng, P. Gao, X. Wang, H. Li, H. Dong, End-to-end object detection with adaptive
clustering transformer, 2020. arXiv:2011.09315.
[38] E. H. C. Harik, F. Guérin, F. Guinand, J. Brethé, H. Pelvillain, Towards an autonomous
warehouse inventory scheme, in: 2016 IEEE Symposium Series on Computational
Intelligence (SSCI), 2016, pp. 1–8. doi:10.1109/SSCI.2016.7850056.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hompel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <source>Essay zur Logistik 4.0</source>
          (
          <year>2015</year>
          ).
          <source>doi:10.13140/RG.2.1.2857.4245.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] pwc,
          <source>Five forces transforming transport logistics</source>
          ,
          <year>2019</year>
          . URL: https://www.pwc.pl/pl/pdf/ publikacje/2018/transport-logistics-trendbook
          <article-title>-2019-en</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cimini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lagorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pirola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <article-title>Exploring human factors in logistics 4.0: empirical evidence from a case study</article-title>
          ,
          <source>IFAC-PapersOnLine</source>
          <volume>52</volume>
          (
          <year>2019</year>
          )
          <fpage>2183</fpage>
          -
          <lpage>2188</lpage>
          . URL: http://www.sciencedirect.com/science/article/pii/S2405896319315137. doi:https:
          <year>2008</year>
          .4599787.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P.</given-names>
            <surname>Jhunjhunwala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shriya</surname>
          </string-name>
          , E. Rufus,
          <article-title>Development of hardware based inventory management system using uav and rfid</article-title>
          , in: 2019
          <source>International Conference on Vision Towards Emerging Trends in Communication and Networking (ViTECoN)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . doi:
          <volume>10</volume>
          .1109/ViTECoN.
          <year>2019</year>
          .
          <volume>8899488</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Rahmadya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Takeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kagoshima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Umehira</surname>
          </string-name>
          ,
          <article-title>A framework to determine secure distances for either drones or robots based inventory management systems</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>170153</fpage>
          -
          <lpage>170161</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2020</year>
          .
          <volume>3024963</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Beul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Droeschel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nieuwenhuisen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quenzel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Houben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Behnke</surname>
          </string-name>
          ,
          <article-title>Fast autonomous flight in warehouses for inventory applications</article-title>
          ,
          <source>IEEE Robotics and Automation Letters</source>
          <volume>3</volume>
          (
          <year>2018</year>
          )
          <fpage>3121</fpage>
          -
          <lpage>3128</lpage>
          . doi:
          <volume>10</volume>
          .1109/LRA.
          <year>2018</year>
          .
          <volume>2849833</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Škrinjar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Škorput</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Furdić</surname>
          </string-name>
          ,
          <article-title>Application of unmanned aerial vehicles in logistic processes</article-title>
          , in: I. Karabegović (Ed.), New Technologies,
          <source>Development and Application</source>
          , Springer International Publishing, Cham,
          <year>2019</year>
          , pp.
          <fpage>359</fpage>
          -
          <lpage>366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>H.</given-names>
            <surname>Borstell</surname>
          </string-name>
          ,
          <article-title>A short survey of image processing in logistics - how image processing contributes to eficiency of logistics processes through intelligence</article-title>
          ,
          <year>2018</year>
          . doi:
          <volume>10</volume>
          .13140/RG. 2.2.11060.76168.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kalinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Petrovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ilin</surname>
          </string-name>
          , E. Pristanskiy,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kurenkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ramzhaev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Idrisov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsetserukou</surname>
          </string-name>
          , Warevision:
          <article-title>Cnn barcode detection-based uav trajectory optimization for autonomous warehouse stocktaking</article-title>
          ,
          <source>IEEE Robotics and Automation Letters</source>
          <volume>5</volume>
          (
          <year>2020</year>
          )
          <fpage>6647</fpage>
          -
          <lpage>6653</lpage>
          . doi:
          <volume>10</volume>
          .1109/LRA.
          <year>2020</year>
          .
          <volume>3010733</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [24]
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Hong-ying, The application of barcode technology in logistics and warehouse management</article-title>
          ,
          <source>in: 2009 First International Workshop on Education Technology and Computer Science</source>
          , volume
          <volume>3</volume>
          ,
          <year>2009</year>
          , pp.
          <fpage>732</fpage>
          -
          <lpage>735</lpage>
          . doi:
          <volume>10</volume>
          .1109/ETCS.
          <year>2009</year>
          .
          <volume>698</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wawrla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Maghazei</surname>
          </string-name>
          , T. Netland,
          <article-title>Application of drones in warehouse operations (</article-title>
          <year>2019</year>
          ).
          <article-title>URL: www.pom.ethz.ch, whitepaper from ETH Zurich (Chair of Production and Operations Management</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Freistetter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Hummel</surname>
          </string-name>
          ,
          <article-title>Human-drone teaming: Use case bookshelf inventory</article-title>
          ,
          <source>in: Proceedings of the 9th International Conference on the Internet of Things, IoT</source>
          <year>2019</year>
          ,
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2019</year>
          . URL: https://doi.org/10. 1145/3365871.3365913. doi:
          <volume>10</volume>
          .1145/3365871.3365913.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>L.</given-names>
            <surname>Dörr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Brandt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pouls</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>Fully-automated packaging structure recognition in logistics environments</article-title>
          ,
          <source>in: 2020 25th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA)</source>
          , volume
          <volume>1</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>526</fpage>
          -
          <lpage>533</lpage>
          . doi:
          <volume>10</volume>
          .1109/ETFA46521.
          <year>2020</year>
          .
          <volume>9212152</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Belongie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. D.</given-names>
            <surname>Bourdev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Perona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Zitnick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Microsoft</surname>
            <given-names>COCO</given-names>
          </string-name>
          :
          <article-title>common objects in context</article-title>
          ,
          <source>CoRR abs/1405</source>
          .0312 (
          <year>2014</year>
          ). URL: http://arxiv.org/abs/1405.0312. arXiv:
          <volume>1405</volume>
          .
          <fpage>0312</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , G. Gkioxari,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mask</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          ,
          <source>CoRR abs/1703</source>
          .06870 (
          <year>2017</year>
          ). URL: http://arxiv.org/abs/1703.06870. arXiv:
          <volume>1703</volume>
          .
          <fpage>06870</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kirillov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Massa</surname>
          </string-name>
          , W.-Y. Lo,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          , Detectron2, https://github.com/ facebookresearch/detectron2,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Eficientdet: Scalable and eficient object detection</article-title>
          , CoRR abs/
          <year>1911</year>
          .09070 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1911</year>
          .09070. arXiv:
          <year>1911</year>
          .09070.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>