<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Docker to deploy computing software</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tatiana S. Demidova</string-name>
          <email>dem_tatiana@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton A. Sobolev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasia V. Demidova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Migran N. Gevorkyan</string-name>
          <email>gevorkyan_mn@rudn.university</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Applied Probability and Informatics Peoples' Friendship University of Russia Miklukho-Maklaya str.</institution>
          <addr-line>6, Moscow, 117198</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>38</fpage>
      <lpage>46</lpage>
      <abstract>
        <p>There are many ways to facilitate the creation of large-scale projects. One of the most commonly used methods is to create virtual machines that contain the program environment. However, software has recently been created to make this process even easier. One example is Docker, a software for automating the deployment and management of applications in an operating system-level virtualization environment. This paper discusses the Docker software, its features and benefits, which allows you to create images that contain the program and all the necessary components for its operation. The purpose of this work is to study the capabilities of Docker. And also, the creation of a container containing a software implementation of the neural network for recognition of various handwritten characters. Training and test data is a database of handwritten numbers and letters "MNIST" and "EMNIST". To teach the neural network to recognize numbers, a training set containing 60 thousand copies was used, and the test set includes 10 thousand copies. For letters - the training set contains 88800 copies, and the test set includes 14800 copies. The project was created on the basis of the image tensorflow downloaded from public Docker-registry Docker Hub. The program is written in Python3, using the service for interactive computing Jupyter Notebook.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Copyright © 2019 for the individual papers by the papers’ authors. Copying permitted for private and
academic purposes. This volume is published and copyrighted by its editors.</p>
      <p>In: K. E. Samouylov, L. A. Sevastianov, D. S. Kulyabov (eds.): Selected Papers of the IX Conference
“Information and Telecommunication Technologies and Mathematical Modeling of High-Tech Systems”,
Moscow, Russia, 19-Apr-2019, published at http://ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>Docker — container virtualisation system</title>
      <p>
        In this paper we use Docker to automate the deployment and management of
applications. Docker is collection of programs used to run processes in an isolated
environment based on special images [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It allows one to create containers that contain
the application and its entire environment, and provides an environment for their use.
      </p>
      <p>The main components of Docker are images, registry and containers.</p>
      <p>Docker image is a read-only template. For example, an image might contain an
operating system with installed applications. Images are used to create containers.
Docker makes it easy to create a new image and update existing ones or download
ready-only images.</p>
      <p>Docker-registry stores images. There are public and private registries. One can
download images from registry or upload images to registry. There is free public Docker
registry called Docker Hub. It is accessible for all users.</p>
      <p>Containers are similar to directories. They contain everything to run the application.
Each container is created from an image, which forms an isolated and a secure platform
for the application.</p>
      <p>Docker container is created by following command
$ docker run &lt;attributes&gt; &lt;image&gt; &lt;command&gt;</p>
      <p>One of the advantages of Docker is the ability to create a container using Dockerfile
and the docker build command. Dockerfile contains the base image. This image is
used to build the desired container by applying commands, specified in Dockefile.</p>
      <p>To demonstrate the capabilities of Docker, a project was created using a neural
network for handwriting recognition.</p>
    </sec>
    <sec id="sec-3">
      <title>Machine learning libraries</title>
      <p>
        Artificial neural network is a mathematical model built on the principle of biological
neural networks of nerve cells of a living organism. A neural network is constructed
in such a way that its nodes work like neurons in the human brain. The node collects
information, processes it, and passes it to the next node [
        <xref ref-type="bibr" rid="ref11 ref18 ref3">3, 11, 18</xref>
        ].
      </p>
      <p>For the implementation of neural networks, there is a huge amount of software. The
main diferences between them are their functionality. While some frameworks are used
as shells to extend functionality and facilitate the writing of neural networks, others are
full-fledged languages and can define neural networks of any level of complexity.</p>
      <p>TensorFlow is library designed for machine learning. This framework is created
by Google. It uses Python languages as back-end, but core parts are written in
C++ [?,13,14]. We use TensorFlow to train and build a neural network for classifying and
ifnding images that are close to human perception. All calculations in this environment
are performed using stream data graphs, where nodes represent diferent mathematical
operations and graph branches represent arrays of data.</p>
      <p>
        The Keras library [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is an add-in for TensorFlow to create high-level neural networks,
written in the Python programming language [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The main advantage of this library it
is easy usage when working with deep learning networks. Keras can be easily extended
with new modules in the form of classes and functions. This environment includes the
implementation of optimizers, layers, functions, and other tools for working with images
and text.
      </p>
      <p>
        The Scikit-learn [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] library, written in Python, has many algorithms for training
neural networks with and without a teacher. Developers pay much attention to the
usability of this environment and optimization issues to improve the speed of its operation.
Scikit-learn includes various classification, regression and clustering algorithms. It is
designed for interaction with numerical and scientific Python libraries such as NumPy
and SciPy.
2.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Neural network architecture</title>
      <p>
        The convolutional neural network [
        <xref ref-type="bibr" rid="ref17 ref6 ref9">6, 9, 17</xref>
        ] was chosen as the topology of the neural
network to solve the problem of handwriting recognition. Today convolutional neural
networks are considered to be the best for solving image recognition problems. The
architecture of the neural network, which is based on the convolution operation, was
ifrst developed in the late 1990s by Lekun et al.
      </p>
      <p>Convolutional neural network (CNN) consists of the following types of layers:
convolutional layers, subsampling (or pooling) layers and perceptron layers. The first two types
of layers, alternating with each other, form the input feature vector for the multilayer
perceptron.</p>
      <p>In convolutional neural network a convolution operation uses a limited matrix of
weights of small size, which moves through the processed layer, forming after each shift
activation signal for the next layer of the neuron with a similar position.</p>
      <p>This weight matrix is called the convolution kernel. In a convolutional neural
network, sets of weights encoding image elements are formed independently by training
the network. After a convolutional layer comes the pooling layer. It also has maps, the
number of which coincides with the previous layer. The purpose of this layer is to reduce
the dimension of the maps. In the previous convolution operation, signs were identified
that do not require such detail. Filtering already unnecessary parts helps the network
avoid overtraining. After several repetitions of convolutional layers and pooling layers, a
layer of the usual multilayer perceptron follows. The output layer is connected to all
neurons of the previous layer and the number of neurons corresponds to the number of
recognized classes. In the case of binary classification, a single neuron and hyperbolic
tangent can be used as an activation function. Then, the output of a neuron with a
value of 1 means belonging to a class, and the output of a neuron with a value of -1
means not belonging to a class.</p>
      <p>The following architecture was used for the convolution network to recognize digits
(fig. 1).</p>
      <p>This network consists of 6 layers:
1. Convolutional layer with 75 feature maps, convolution kernel size: 5x5.
2. Pooling layer (MaxPooling) with 2x2 poolsize.
3. Another convolution layer with 100 feature maps, convolution kernel size: 5x5.
4. Second pooling layer with 2x2 poolsize.
5. Fully connected layer with 500 neurons
6. Fully connected output layer with 10 neuron, which correspond to the classes of
handwritten digits from 0 to 9.</p>
      <p>The activation function in hidden layers is ReLU, and the output layer is softmax.</p>
      <p>As training and test data we use "MNIST" database of handwritten numbers. To
train the neural network to recognize numbers we use a training set containing 60
thousand copies and the test set containing 10 thousand copies.</p>
      <p>We use Matplotlib library to create an image of the symbols from the database 2.
3.</p>
    </sec>
    <sec id="sec-5">
      <title>Neural network classification quality assessment</title>
      <p>
        One of the concepts for describing metrics in terms of classification errors is the
error matrix. It is used if there are two classes and an algorithm predicting that each
object belongs to one of the classes, then the classification error matrix will look like
this [
        <xref ref-type="bibr" rid="ref10 ref16">10, 16</xref>
        ]:
      </p>
      <p>One of the simplest metrics is accuracy — the proportion of true results (both positive
and true negative) among the total number of cases considered, i.e. the probability that
the class will be predicted correctly.</p>
      <p>=</p>
      <p>+  
  +   +   +</p>
      <p>.</p>
      <p>To assess the quality of the algorithm on each of the classes precision and recall
metrics are used:
 =</p>
      <p>+  
,
 =</p>
      <p>+  
.</p>
      <p>The accuracy of the classification of positive results (precision) is the proportion of
positive results that are correctly identified by the classifier among the total number of
considered cases.</p>
      <p>Recall, also known as sensitivity, reflects the proportion of positive results that are
correctly identified by the classifier.</p>
      <p>Specificity — reflects the proportion of negative results that are correctly identified
by the classifier.</p>
      <p>To associate a precision with the fullness F-measure are used. F-measure is a harmonic
mean of precision and completeness:
 =
2 ·  · 
 + 
.</p>
      <p>One way to evaluate the model is AUC-ROC — square under the curve of error in
coordinates True Positive Rate (TPR) and False Positive Rate (FPR):
   =</p>
      <p>+  
,
   =</p>
      <p>+  
.</p>
    </sec>
    <sec id="sec-6">
      <title>4. Software implementation of neural network</title>
      <p>The neural network for handwriting recognition based on MNIST, showed an average
eficiency of about 99.24%. The average training time of one era is about 160 seconds</p>
      <p>Metrics were also calculated for each of the classes:</p>
    </sec>
    <sec id="sec-7">
      <title>5. Implementation of the project using Docker</title>
      <p>The following describes the process of downloading the created software package
in Docker Hub. The first step is to create an image from the container that contains
the program. We use the commit command to commit the changes to the new Docker
image.</p>
      <p>$ docker commit container_id repository/new_image_name</p>
      <p>To show container’s id we use command docker ps –a. As repository name we
use user’s Docker Hub login and assign a name to the image. This image is saved
locally and to make sure that the new image is saved successfully, we use the command
docker images.</p>
      <p>Next, we upload our image to Docker Hub. To do this we enter the DockerHub
account with command
$ docker login -u docker-registry-username
where docker-registry-username — Docker Hub user name. We enter the password
and upload the image to DockerHub by the command
$ docker push docker-registry-username/docker-image-name
After successfully uploading the image, we can see it in the account toolbar (Fig. 6).</p>
      <p>Thus, we can conclude that Keras library combined with Docker is a powerful toolset
for the development software implementation of various machine learning methods.
Keras is the most convenient library for writing neural networks, as it is easy to use
and has a high speed of creating neural network models. The ease of use of Keras is
due to the huge number of pre-installed functions designed to create diferent layers
of the neural network. And Docker makes it easier to develop applications, because
the installation of all the necessary libraries can be replaced by a single command —
download the desired image and in addition, is a universal way to deliver the developed
applications on local machines and run them in an isolated environment.
6.</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>The paper considered Docker software that allows you to create images that contain
the program and all the necessary components for its operation.</p>
      <p>Created image based on tensorflow containing the Jupyter notebook implemented in
Python convolutional neural network using bibltoteki Keras.</p>
      <p>This image was uploaded to the public image repository of Docker Hub containers.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The publication has been prepared with the support of the “RUDN University
Program 5-100”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Docker:
          <article-title>Enterprise Container Platform - URL: www</article-title>
          .docker.com
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Docker</surname>
          </string-name>
          Hub - URL: https://hub.docker.com/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Tariq</surname>
          </string-name>
          , Rashid.
          <source>Make Your Own Neural Network - Spb.: “Alfa-kniga”</source>
          ,
          <year>2017</year>
          . - 272 p.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>4. The MNIST database of handwritten digits -</article-title>
          URL: http://yann.lecun.com/exdb/mnist
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>The EMNIST</surname>
          </string-name>
          Dataset - URL: http://www.nist.gov/itl/iad/image-group/emnistdataset
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>LeCun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Gradient-based learning applied to document recognition / Y</article-title>
          . LeCun [et al.]
          <source>// Proc. of the IEEE. - 1998</source>
          . - Vol.
          <volume>86</volume>
          , No 11. - P.
          <fpage>1</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Scikit-Learn</surname>
          </string-name>
          :
          <article-title>Machine Learning in Python</article-title>
          . - URL: https://scikit-learn.org/stable/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Aurelien</given-names>
            <surname>Geron</surname>
          </string-name>
          .
          <article-title>Hands-On Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems</article-title>
          .:
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          -
          <year>2017</year>
          . - 564 p. - ISBN-
          <volume>13</volume>
          :
          <fpage>978</fpage>
          -
          <lpage>1491962299</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>F.</given-names>
            <surname>Rosenblatt</surname>
          </string-name>
          .
          <article-title>The perceptron: A probabilistic model for information storage and organization in the brain</article-title>
          . - Psychological
          <string-name>
            <surname>review</surname>
          </string-name>
          . - Vol.
          <volume>65</volume>
          , No.
          <volume>6</volume>
          <fpage>.</fpage>
          -
          <lpage>1958</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Werbos</surname>
            <given-names>P. J.</given-names>
          </string-name>
          , Beyond regression:
          <article-title>New tools for prediction and analysis in the behavioral sciences</article-title>
          .
          <source>Ph.D. thesis</source>
          , Harvard University, Cambridge, MA,
          <year>1974</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hopfield</surname>
            <given-names>J. J.</given-names>
          </string-name>
          <article-title>Learning algorithms and probability distributions in feed-forward and feed-back networks</article-title>
          .
          <article-title>- 1987.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Keras</surname>
          </string-name>
          :
          <article-title>The Python Deep Learning library</article-title>
          . - URL: https://keras.io/
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Aurelien</given-names>
            <surname>Geron</surname>
          </string-name>
          .
          <article-title>Hands-On Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems</article-title>
          .:
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          -
          <year>2017</year>
          . - 564 p.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>14. Tensorflow: oficial website - URL: https://www.tensorflow.org/</mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Richert</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coelho</surname>
            <given-names>L.</given-names>
          </string-name>
          <article-title>Building machine learning systems with Python</article-title>
          .
          <source>- Birmingham: Packt Publ. - 2013</source>
          . - 290 p.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Hastie</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tibshirani</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            <given-names>J.</given-names>
          </string-name>
          <article-title>The elements of statistical learning: data mining, inference, and prediction</article-title>
          . -2nd ed. -New York: Springer. -
          <year>2013</year>
          . - 745 p.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Hornick</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stinchcombe</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            <given-names>H</given-names>
          </string-name>
          .
          <article-title>Multilayer feedforward networks are universal approximators//Neural Networks</article-title>
          .
          <year>1989</year>
          . Vol.
          <volume>2</volume>
          , no. 5. P.
          <volume>359</volume>
          -
          <fpage>366</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Ben</surname>
            <given-names>J. A.</given-names>
          </string-name>
          <string-name>
            <surname>Kroese</surname>
            and
            <given-names>P. Patrick van der Smagt.</given-names>
          </string-name>
          <article-title>An introduction to neural networks</article-title>
          .
          <article-title>- 1993.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>