<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Segmentation: Using deep convolutional networks for coral reef pixel-wise parsing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aljoscha Steffens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Campello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James Ravenscroft</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrian Clark</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hani Hagras</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Filament AI</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Essex</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we describe a deep-convolutional network based method to segment coral reef images into different types of substrates. The method described in the paper includes data preparation, model summary, specific techniques to deal with class imbalance, and downstream post-processing computer vision tasks, such as morphological operations and polygon generation from pixel segmentation. We present the results of our method in the ImageCLEFcoral pixel-wise parsing task, evaluated across the different classes of substrate.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Semantic segmentation models have received significant attention in computer
vision due to their applicability in medical imagining, autonomous driving, and
full-scene understanding. In this paper, we consider the ImageCLEF 2019
pixelwise parsing competition [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which consists of segmenting pictures from coral
reefs into 13 different substrates. We are particularly interested in evaluating the
applicability of deep convolutional neural networks (DCNNs) to the coral images,
taken under real conditions in the ocean. A model that performs well on the task
of automatic segmentation of corals could be beneficial to the conservation of
reefs by measuring the amounts of different corals, their condition and other
characteristics.
      </p>
      <p>
        This paper is organised as follows. In section 2 we explore the
ImageCLEFcoral dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], how to split it into training and validation set in order to keep the
distributions balanced, as well as a data augmentation approach. In section 3, we
describe DeeplabV3 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a DCNN designed for semantic segmentation, alongside
with our pipeline. Our method includes post-processing tasks such as
morphological operations and polygon filling. Training, bootstrapping and inference are
also described. In section 4 we discuss possible routes to improve the results.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>The data that was provided consisted of 240 images of size 4032 3024 3
as well as a text file containing the polygons in the images for the following
classes/substrates:</p>
      <p>Hard Coral - Branching Soft Coral
Hard Coral - Submassive Soft Coral - Gorgonian
Hard Coral - Boulder Sponge
Hard Coral - Encrusting Sponge - Barrel
Hard Coral - Table Fire Coral - Millepora
Hard Coral - Foliose Algae - Macro or Leaves
Hard Coral - Mushroom</p>
      <p>We mapped the 13 substrates to integers f0; 1; : : : 13g (the background
corresponding to 0) and created a 4032 3024 integer matrix for each image defined
as</p>
      <p>Mikj = c if pixel i; j corresponds to substrate c,
(1)
with ties broken arbitrarily. The corresponding matrix acts as a per-class "mask"
for each substrate, an example can be seen in figure 1.</p>
      <p>Fig. 1: Substrate mask and image for id 2018_0714_112417_024
The class-distribution of the pixel reflects the naturally occurring composition
of corals in the photographed area and is thus highly imbalanced, see figure 2.
Also, the class-distributions between two images can vary strongly as only a
limited selection of coral types will be present in each image. Both the overall
class-imbalance and the difference in inter-image class-distributions need to be
taken into account and are discussed in section 3.2 and 2.1 respectively.
2.1</p>
      <sec id="sec-2-1">
        <title>Data Split</title>
        <p>The data was split into a training and validation set with 204 images (85%)
for training and 36 for validation. Due to reasons discussed in section 2.2, we
chose to split the data on a per-image basis. This posed a problem given the
difference in inter-image class-distributions and the relatively small number of
images: for a random split it would be highly likely that the overall, training, and
validation class-distribution would differ strongly which would have an effect on
the evaluation of the model performance.</p>
        <p>In order to achieve balanced distributions for the training and validation
set, we created N training and validation sets of constant size, with randomly
selected images and compared the resulting training/validation distributions by
using a cosine distance that is weighted by the overall class-distribution.
dist(v1; v2; w) =</p>
        <p>Pd</p>
        <p>i=1 wi2
qPd
i=1 wi
vi1</p>
        <p>vi1 vi2
qPd
i=1 wi
vi2
(2)</p>
        <p>From those N splits, we chose the split that resulted in the lowest distance.
Afterwards, we selected one item from each set that, if swapped, resulted in the
biggest decrease in the weighted cosine distance. Swaps were performed until
there was no further decrease possible by swapping individual items. The whole
procedure - N random splits, choosing the best split, optimise by swapping
was performed several times in order to increase the chance of finding an good
final split.</p>
        <p>While the given approach does not result in an optimal solution, it is fast and
did achieve satisfying result for the given task. The three distributions (overall,
training, validation) can be seen in figure 2.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Data Augmentation</title>
        <p>Splitting the data on image level was mainly done due to how we prepared
the data before feeding it into the Neural Network. With their large spatial
dimensions that resulted in both rich details and many different coral types per
image, it seemed sensible to use random cropping as a main preparation step.
For each image that was loaded into memory during training, 16 random crops
with square sizes between 400 400 and 1400 1400 were performed. The crops
were then bi-linearly scaled to 256 256 and randomly flipped in vertical and
horizontal direction before they were fed into the network. Both the flipping
and the random cropping served as a data augmentation method that ensured
that the network was exposed to a variety of relative coral sizes, orientations,
and image-compositions. This way, the network never saw the exact same image
twice which reduced overfitting.</p>
        <p>
          For the validation data, random crops and re-scaling were done prior to
training the network and thus always the same. This was done to make the
metrics that were computed after each epoch comparable.
We used DeeplabV3 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], a deep convolutional neural network that improves over
existing networks. In particular, DeeplabV3 extends Deeplab [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and avoids the
need of a post-processing machine learning model (such as conditional random
fields). Nevertheless, since the challenge is evaluated over polygons, we needed to
apply image processing techniques (in particular, morphological transformations
and region-filling algorithms) in order to generate the final file. Note that the
polygon-filling operations have more degrees-of-freedom than the training
preprocessing, therefore adding extra parameters to the model. These operations
are described in more details in 3.5. A high-level operational diagram of the
model and evaluation can be found below. More details will be described in the
next sections.
Deeplab V3 is a Deep Convolutional Neural Network (DCNN) for semantic
image segmentation proposed by Chen et al in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] in 2017. It is a state-of-the-art
network architecture that, with pretraining on the ImageNet [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and JFT-300M
[10] dataset resulted in a mIOU of 86.9% on the PASCAL VOC 2012 test set [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
The model consist of two parts: The first part is a feature extracting backbone
that is not strictly limited to be of a given type - in the paper the authors use
a ResNet-50 and ResNet-101 but other architectures can be employed as well.
The second part is where the novelty happens with an extensive use of atrous
convolutions. An atrous convolution filter kernel layout is defined by the normal
size of say 3x3 and in addition to that, an atrous rate that specifies how many
0-values are between the individual filter entries along the spatial dimensions.
With 0-values between, the atrous convolution would be the same as a normal
convolution; inserting one zero between two neighbouring values in a 3x3
convolution would make it have the same receptive field as a 5x5 convolution, while
only employing 9 weights instead of 25. Deeplab V3 uses atrous convolutions for
constructing feature pyramids. Feature pyramids are used to combine features
from different scales into one feature map, which is done by using atrous
convolutions with different rates on the same feature map and concatenating the
individual outputs to a new feature map. For our submission we used a PyTorch
implementation of DeepLab V3 found from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] with a ResNet101 backbone [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
and an output-stride of 16
        </p>
        <p>Prior to polygon-filling post-processing, the model outputs, for every pixel,
a probability that such pixel belongs to a class, or more formally:
fikj (c) = probability that pixel ij on image k belongs to class c
(3)
3.2</p>
      </sec>
      <sec id="sec-2-3">
        <title>Class imbalance and weighted loss function</title>
        <p>As discussed in section 2, the class-distribution of the data is highly skewed.
This is a problem, as the model will emphasize more on classifying frequent
classes correctly to achieve lower errors if no counter-measures are in place. One
approach that is often used - and that was used here as well - is to weigh the
loss function based on the class distribution. We used a pixel-wise cross-entropy
loss and weighted the individual components with following weights:
w(c) =</p>
        <p>1
log( + p(c))
(4)
with p(c) being the relative occurrence of class c and being a hyper-parameter
that scales the weights (in our submission = 1:025 yielded good results).
This was done as there are orders of magnitudes between the individual relative
occurrences, and the model would over-emphasize on the infrequent classes if we
just used the reciprocal.</p>
        <p>The final cross-entropy loss is as follows:
N</p>
        <p>1
height
width</p>
        <p>N height width C
X X X X
k=1 i=1 j=1 c=1
yikj (c)
w(c) log</p>
        <p>exp(fikj (c)) !
P exp(fikj (c))</p>
        <p>;
(5)
where fikj (c) and yikj (c) describe the predicted confidence (f ) and ground truth
(y) for image k at position ij and class c respectively and N is the number of
images that are included in the loss.
3.3</p>
      </sec>
      <sec id="sec-2-4">
        <title>Training and Bootstrapping</title>
        <p>We trained the Neural Network for 50 epochs, with a batch size of 32 (2 images
per batch, 16 crops per image) on a Nvidia GeForce GTX 1080 Ti. After the
training, we used the network to predict the training images and cropped out
areas where the network was particularly bad. The network was then trained on
those images for another 30 epochs.
3.4</p>
      </sec>
      <sec id="sec-2-5">
        <title>Inference</title>
        <p>In order to predict a full-sized image at inference time, we used a sliding
window approach. With window-sizes of 500 500, 1000 1000, and 1500 1500
corresponding step-sizes of 400, 800, and 1200 we cut each 4032 3024 image
into 112 partially overlapping sections. Each section was then scaled to 256 256
and fed into the Neural Network. The results were scaled to their original size
and added at their respective position to a 4032 3024 14 confidence matrix
C. For each position Ci;j the number of votes were stored (meaning how often
a given pixel was predicted) so that the average confidence could be calculated
subsequently. The final classification for pixel i; j was then given by
c = argmax Ci;j (k)
k2f0;:::;13g
(6)
By using sliding windows with different window-sizes we made sure that each
pixel was predicted at several different resolutions and thus also with different
amounts of context.
3.5</p>
      </sec>
      <sec id="sec-2-6">
        <title>Post-processing</title>
        <p>
          After calculating the classification mask for a predicted image, we used several
basic computer vision algorithm for post-processing and transforming the data
into the given submission format.
1. Find connected components.
2. Morphological opening with kernel size = (31; 31)
3. Morphological closing with kernel size = (31; 31)
4. Flood fill
5. Polygon approximation using Douglas-Peucker algorithm [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], with maximum
distance to correct output " equals 0:1% of the contour arc-length for that
connected component.
        </p>
        <p>
          We used the OpenCV 3.4.2 implementations [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] of the corresponding
algorithms.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Final results and further work</title>
      <p>The average intersection over union for all classes, as reported in the official
ImageCLEF 2019 Coral competition, over the test set, is described in Table 2.
Soft corals and hard corals (boulder) performed relatively well in comparison to
the other classes, which was expected, anticipated by the class abundance, as
shown in figure 2. Surprisingly, mushroom pixels have been correctly identified
21%, in spite of the small number of samples. A full analysis and visualisation
of the results per class is left for future investigation.</p>
      <p>There are a number of areas where our pipeline can potentially be improved
in the future and things that can be investigated:
– Increasing the input size to the model
– Increase the batch size
– Test different model backbones
– Tuning hyper-parameters (both for training and for post-processing)
– Investigate the impact of crop sizes
– Investigate the impact of bootstrapping
– Try different methods to counteract the class-imbalance
10. C. Sun, A. Shrivastava, S. Singh, and A. Gupta1. Revisiting unreasonable
effectiveness of data in deep learning era. ICCV, 2017.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>Open source computer vision library 4.1.0</article-title>
          . https://docs.opencv.
          <source>org/4</source>
          .1.0/,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <article-title>Opencv contour features</article-title>
          . https://docs.opencv.
          <source>org/3</source>
          .1.0/dd/d49/tutorial_ py_contour_features.html,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>3. pytorch-deeplab-xception</article-title>
          . https://github.com/jfzhang95/ pytorch-deeplab-xception,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Clift</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          .
          <article-title>Overview of ImageCLEFcoral 2019 task</article-title>
          .
          <source>In CLEF2019 Working Notes</source>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Papandreou, I. Kokkinos,
          <string-name>
            <given-names>K.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Yuille</surname>
          </string-name>
          . Deeplab:
          <article-title>Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach</source>
          . Intell.,
          <volume>40</volume>
          (
          <issue>4</issue>
          ):
          <fpage>834</fpage>
          -
          <lpage>848</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Papandreou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schroff</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Adam</surname>
          </string-name>
          .
          <article-title>Rethinking atrous convolution for semantic image segmentation</article-title>
          .
          <source>CoRR, abs/1706.05587</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei. ImageNet: A LargeScale Hierarchical Image</surname>
          </string-name>
          <article-title>Database</article-title>
          .
          <source>In CVPR09</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , pages
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          ,
          <year>June 2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. D.</given-names>
            <surname>Cid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Klimuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tarasau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          , D.-
          <string-name>
            <surname>T.</surname>
            Dang-Nguyen,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            , M.-
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Kavallieratou</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          <string-name>
            <surname>del Blanco</surname>
            ,
            <given-names>C. C.</given-names>
          </string-name>
          <string-name>
            <surname>Rodríguez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Vasillopoulos</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Karampidis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Campello</surname>
          </string-name>
          .
          <source>ImageCLEF</source>
          <year>2019</year>
          :
          <article-title>Multimedia retrieval in medicine, lifelogging, security and nature</article-title>
          .
          <source>In Experimental IR Meets Multilinguality</source>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 10th International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ), Lugano, Switzerland,
          <source>September 9-12 2019. LNCS Lecture Notes in Computer Science</source>
          , Springer.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>