<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Usage of fully convolutional neural network for automation of extracting the left ventricle contour on the ultrasonic data images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrey A. Mukhtarov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasiliy V. Zyuzin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasia O. Bobkova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ural Federal University</institution>
          ,
          <addr-line>Yekaterinburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>75</fpage>
      <lpage>82</lpage>
      <abstract>
        <p>The article discusses experience of application of fully convolutional neural networks for automation of left ventricle contouring. Results of the quality analysis of contouring show that this approach can be used to automate the work of cardiologists with the echographic data.</p>
      </abstract>
      <kwd-group>
        <kwd>Contouring</kwd>
        <kwd>left ventricle</kwd>
        <kwd>neural networks</kwd>
        <kwd>ultrasonic images</kwd>
        <kwd>image segmentation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Cardiologists use the echographic data of patients to determine the left ventricle
(LV) area of the heart in order to study the contractility of the left ventricle
walls, restore the LV volume, and calculate various indicators. As a rule, the
contour is selected subjectively, and it depends on quali cation of the physician
performing the procession of medical images. Such diagnostics takes a long time
and is not always accurate.</p>
      <p>At the moment, there are no automated software tools that allow one to fully
automate the LV contouring on the heart ultrasonic data. Thus, the problem of
increasing the speed and quality of diagnostics by automating the LV contouring
is actual.</p>
    </sec>
    <sec id="sec-2">
      <title>Choice of neural network model</title>
      <p>To solve the problem, it was decided to use some machine learning method i.e.,
the neural networks. The results of literature research on the similar subjects
show that a fully convolutional neural network (FCN) gives the best results
for the problem of image segmentation. This network is similar to convolutional
neural network (CNN) where the last fully connected layer is replaced by another
convolution layer with a large susceptible eld. The idea is to capture the global
scene context, which gives information about objects on the image including
their localization.</p>
      <p>Analysis of existing studies of similar problems showed that the neural
network AlexNet is the most popular implementation of the CNN for the general
classi cation of objects. The AlexNet model outperforms competing approaches
based on traditional functions in solving a number of computer vision problems.</p>
      <p>Nox existing approach is based on the recent successful results of using the
deep networks for the image classi cation [1{3]. Chen Liang-Chieh, Evan
Shelhamer, and Phi Vu Tran [4{6] presented the experience of using fully
convolutional neural networks. Moreover, authors of article [5] described in details the
solution of the multiclass semantic segmentation problem. Our paper describes
how the CCN model of the AlexNet is converted to FCN-32 in order to be able
to train the network using images of arbitrary size. Furthermore, the authors
show how to tune maximally ne the neural network for the best classi cation
by converting the network to FCN-8. The experiments were carried out on
Pascal VOC data, which include color images of various types. By their research, it
was shown that such application of the fully convolutional neural networks gives
the best result.</p>
      <p>For the presented problem, it was decided to use the pre-trained model
FCN8-AlexNet-pascal. Since in our case just only one object is required to be
identied, there should be only two classes at the network outlet: the background and
the LV region. Therefore, an additional convolutional layer with two outputs was
added to the source network. In order to determine that the network has started
to relearn at an early stage, another set of images is used in training whith
volume about 10% of the training set. The network is not training on this test set.
Network predicts the result on the test set, and a test error is determined. If an
error on the training set is usually reduced, then the test error can increase in
this case. This means that the network has become more receptive to its
training set and the rest part of the images will not be recognized correctly. So, it is
necessary to change the training parameters, network structure, or training set.</p>
      <p>Figure 1 shows several intermediate consecutive convolutional layers (vertical
lines) for highlighting more complex maps of image feature. The pool layers are
indicated by a grid that shows the relative spatial dimension. The rst line
(FCN32) is a single-stream network that allows one to predict an area of 32 pixels in
one step. The second line (FCN-16) allows predicting the network more subtle
details while retaining the high-level semantic information by combining the
forecasts of the last layer and the pool4 layer (areas of 16 pixels). The third line
(FCN-8) provides additional classi cation accuracy from the pool3 projections.</p>
      <p>A convolutional layer has a set of matrix lters that are applied to images and
determine the feature. A combination of such several layers will build new
features according to the previous features of a lower order. In practice, this means
that the network is trained to see complex features, which are a composition of
simpler ones.</p>
      <p>The pooling layer represents a layer without training. Here, the images are
ltered highlighting the largest value of the pixel in the area and ignoring the
others. Thus, the image decreases in size and the most signi cant features are
left regardless of their location.</p>
      <p>The next three layers are a fully connected network. Here, each neuron takes
in the input all the outputs of the previous layer neurons. Then, the upsample
layer performs the image increase.</p>
      <p>Thus, it was decided to use the original FCN-8 model with some modi
cations. To avoid over tting of the network, it was decided to add layers of
normalization and dropout. Dropout randomly disconnects some neurons from
the fully connected layer during training.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Building a model</title>
      <p>The training was conducted on ultrasound images of patients, the total number
included 1895 images. From them 90% of the frames were used as a training
sample, and the remaining 10% as the test set. The study was conducted on
a GPU (graphics processing unit) using the ca e framework and the nVidia
GTX 1070 graphics card. The training lasted for 12 hours and was ful lled by
the backpropagation method. Classi cation is based on blocks of pixels, i.e., the
central pixel and 8 nearest to it.</p>
      <p>The following steps are necessary to train the neural network. The training
takes place through several iterations. Their number is set during the training of
the network. The network passes through all input data at each iteration. The
following steps are performed at each iteration:
1. upload the data and initialize the weights in random order;
2. perform the direct propagation;
3. calculate losses;
4. perform the reverse propagation;
5. update the weights using gradient descent;
6. repeat from step 2 until all the iterations run out.</p>
      <p>The loss function is a mathematical function (with the current set of
parameters) that shows the quality of classi cation. The selected pretrained model was
used to select simultaneously several objects in the images. In order to classify
each pixel of the image, a map of all detected objects in the image was used.
Each pixel has an assosiated label. In order to di er the objects, the colors were
indexed, so, each object had its own color. Therefore, the training of the network
required preliminary processing of the input data. For this purpose, the images
with expert contours were converted to RGB with a mask for color indexing.</p>
      <p>The network architecture is presented in Table 1 and in Fig. 2; examples of
the input data are shown in Fig. 3.
{ precision</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation results</title>
      <p>To quantify the quality of contouring, it was decided to use the following criteria:
P recision =</p>
      <p>S\ ;</p>
      <p>Scont
Recall =</p>
      <p>S\ ;
Sexp
where S\ is the intersection of the square area limited by expert contour
and area formed from classi ed pixels, Scont is the square area formed from
the classi ed pixels;
{ recall</p>
      <p>where Sexp is the area of the region bounded by the expert contour;
{ F-measure</p>
      <p>F =
2</p>
      <p>P recision Recall
P recision + Recall
;
{ proportion of erroneously classi ed pixels;
{ proportion of correctly classi ed pixels;
{ area under receiver operating characteristic curve (AUC).</p>
      <p>
        The results of network training are shown in Fig. 4.
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
      </p>
      <p>Figure 4 shows that the model has been trained quite well. The losses of
the training and test samples are close to zero, while the dice (the Sorensen
coe cient) on the test sample reaches the value of 94%.
Results of the research show that the method of fully convolutional neural
networks can be used to distinguish the LV heart contour on ultrasound data. This
method gives the best results in comparison with other researched methods. To
increase the accuracy of contouring, the increase in the train sample is required,
and the use of other neural networks models is also possible.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G. E.:
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>Advances in neural information processing systems</source>
          .
          <volume>1097</volume>
          {
          <issue>1105</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409</source>
          .
          <fpage>1556</fpage>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          et al.:
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <volume>1</volume>
          {
          <issue>9</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papandreou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kokkinos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuille</surname>
            ,
            <given-names>A. L.</given-names>
          </string-name>
          :
          <article-title>Semantic image segmentation with deep convolutional nets and fully connected CRFs</article-title>
          .
          <source>International conference on learning representations. (</source>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Shelhamer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Fully convolutional networks for semantic segmentation</article-title>
          .
          <source>Computer vision and pattern recognition</source>
          .
          <volume>39</volume>
          ,
          <issue>640</issue>
          {
          <fpage>651</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>P. V.</given-names>
          </string-name>
          :
          <article-title>A fully convolutional neural network for cardiac segmentation in ShortAxis MRI</article-title>
          .
          <article-title>Computer vision and pattern recognition</article-title>
          . (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zyuzin</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bobkova</surname>
            ,
            <given-names>A. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Porshnev</surname>
            ,
            <given-names>S. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukhtarov</surname>
            ,
            <given-names>A. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bobkov</surname>
            ,
            <given-names>V. V.:</given-names>
          </string-name>
          <article-title>The application of decision trees algorithm for selecting the area of the left ventricle on echocardiographic images</article-title>
          .
          <source>First International Workshop on Pattern Recognition</source>
          ,
          <volume>10011</volume>
          ,
          <fpage>100110I</fpage>
          -
          <lpage>1</lpage>
          {
          <fpage>100110I</fpage>
          -
          <lpage>7</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Porshnev</surname>
            ,
            <given-names>S. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukhtarov</surname>
            ,
            <given-names>A. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bobkova</surname>
            ,
            <given-names>A. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zyuzin</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bobkov</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          :
          <article-title>The study of applicability of the decision tree method for contouring of the left ventricle area in echographic video data</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          ,
          <volume>1710</volume>
          , 248{
          <fpage>258</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>