<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Image Segmentation via Self-Supervised Learning Image Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea M. Storås</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SimulaMet</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norway andrea@simula.no</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the submission of team Medical-XAI for the Medico: Transparency in Medical Image Segmentation task held at MediaEval 2021. We propose an unsupervised method that utilizes tools from the field of explainable artificial intelligence to create segmentation masks. We extract heat maps, which are useful in order to explain how the 'black box' model predicts the category of a certain image, and the segmentation masks are directly derived from the heat maps. Our results show that the created masks can capture the relevant findings to a certain extent using only a small amount of image-level labeled data for the classification model and no segmentation masks at all for the training. This is promising for addressing diferent challenges within the intersection of artificial intelligence for medicine such as availability of data, cost of labeling and interpretable and explainable results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Medical image segmentation is one of the focus areas for researchers
working on artificial intelligence ( AI) and medicine. Especially since
the release of U-Net [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the field has somehow exploded, leading
to a myriad of publications of diferent segmentation approaches
within diferent medical specialisations. One of the areas that get
most of the attention is the segmentation of polyps in the colon.
An important motivation is that colon cancer is one of the most
prominent cancers worldwide and early detection by finding polyps
is an eficient method to reduce mortality. The MediaEval Medico
challenge 2021 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is using this important medical challenge as a
task in addition to adding an extra challenge by asking participants
to provide as transparent solutions as possible. Transparency within
AI is a rather new concept and has several sub-parts that contribute
to it ranging from open data, over explainable and interpretable
results to open source code. Our solution for solving this year’s task
is going a step further by also addressing the problem of dealing
with a low amount of labeled segmentation data because obtaining
accurate labels for the medical data is often dificult due to the
availability of medical experts. In the following, we provide a detailed
description of our approach, followed by experimental results and
a detailed discussion of the advantages and disadvantages of our
method.
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        Our method consists of several steps, as illustrated in Figure 1. First,
global features are extracted from the images, and clustering is
applied for labeling unlabeled medical images. A few labeled images
are clustered together with a high number of unlabeled images from
the HyperKvasir data set [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Unlabeled images that fall into the
cluster with the highest number of labeled images get the same
label as the labeled images in the cluster. The process is repeated
until a suficient number of labeled images is reached. The k-means
algorithm from scikit-learn [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is used for clustering, and diferent
numbers of clusters are tested. The k-means algorithm is always
initialized with init = ‘k-means++’, random_state = 0, max_iter =
300 and algorithm = ‘auto’. The rest of the hyperparameters are set
to default values. An EficientNet-b1 classifier [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] implemented in
Pytorch [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is trained on the labeled images to predict the correct
category.
      </p>
      <p>
        The Adam optimizer and the BCEWithLogitsLoss from PyTorch
with default hyperparameter settings are applied during model
training. The classifier is evaluated on images with known labels
to measure model performance as well as evaluating the clustering
technique for labeling images. Grad-CAM [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is applied to create
heat maps highlighting which pixels the classifier focuses on during
classification. Since the model is trained to detect polyps, we expect
the heat maps to highlight these as segments. Segmentation masks
are constructed from the heat maps. Several thresholds are tested
and the most promising values are selected based on visual
inspection. All created models and source code can be found publicly
online1.
      </p>
      <p>
        For each task we submitted five diferent runs with diferent
configurations of our system. Run 1 applied the development data
set provided in the challenge. In order to get images without polyps,
1https://github.com/kelkalot/Medico-2021-Team-Medical-XAI
(a)
(e)
(b)
(f)
(c)
(g)
(d)
(h)
each image was split into 4 tiles. The corresponding segmentation
masks were used to label 100 tiles as ‘polyps’ and 100 as
‘nonpolyps’. The rest of the tiles were unlabeled. For Runs 2 - 5 we
used 100 images labeled as ‘polyps’ and 100 images labeled as
‘non-polyps’ from the Kvasir data set [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Unlabeled images from
the HyperKvasir data set [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] were included for labeling using the
clustering technique described above.
      </p>
      <p>
        For all runs, the following global features were extracted using
the LIRE [
        <xref ref-type="bibr" rid="ref3 ref8">3, 8</xref>
        ] library: EdgeHistogram, Tamura, LuminanceLayout
and SimpleColorHistogram. The internal evaluations of the
classiifers were performed on 60 images not used for training. All models
were trained for 25 epochs on 1, 000 labeled images for a fair
comparison. Heat maps were generated by passing the images through
the final model. We explored extracting heat maps from several
layers in order to identify the most appropriate model. The diferent
configurations are the following: Run 1: 50 clusters, layer 14; Run
2: 200 clusters, layer 13; Run 3: 200 clusters, layer 14; Run 4: 250
clusters; layer 20; Run 5: 250 clusters, layer 22. ‘Layer’ corresponds
to the layer in the model that the heat maps were extracted from.
The segmentation masks were constructed from the heat maps. The
thresholds were set as &gt; 0.4 and &lt; 0.7 for Runs 1 - 3, while the
threshold was &gt; 0.28 for Runs 4 and 5.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>RESULTS AND ANALYSIS</title>
      <p>Segmentation masks for two of the runs are illustrated in Figures 2a
- 2h. Table 1 shows the results for Subtask 1, which is polyp
segmentation. Overall, we observe that our method does not reach
perfect scores, but taking into account that the segmentation is
performed unsupervised, it still achieves acceptable results. We also
observe that the number of clusters is connected to how good the
performance is. This is most probably due to the fact that a higher
number of clusters leads to more specialized clusters, which then
have a positive efect on the performance of the following model.
However, the choice of layer for the heat maps and thresholds for
the segmentations could also afect the performance.</p>
      <p>A. Storås
Precision</p>
      <p>In terms of how fast the inference is performed, we actually
observe that the model based on the larger number of clusters is
faster than the others. The reason for this is not clear, although
it seems that higher performance is connected to faster inference
time. This needs to be investigated more in future work.
4</p>
    </sec>
    <sec id="sec-4">
      <title>DISCUSSION AND CONCLUSION</title>
      <p>To cluster the medical images, we used global features, which are
easy to interpret and increase the transparency of our system. The
heat maps are useful in order to explain how the ‘black box’ model
predicts the category of a certain image, and the segmentation
masks are directly derived from the heat maps. We believe the level
of transparency of our system is quite high. The overall performance
of the approach is for sure on the lower scale, but taking into
account that it is completely unsupervised segmentation, it can
still be considered as good. Overall the presented method seems
promising for medical applications and opens up several directions
of future work.
5</p>
      <p>FUTURE WORK
1, 000 images were used to train the classifiers. For future work, we
will test how the performance changes with an increasing number
of images. Moreover, we want to explore other global features for
clustering the images as well as deep features. Other clustering
algorithms should also be tested. By connecting the knowledge we
have about the global features with the deep features, we might get
an interpretation about what features (color, texture) are important
for a certain disease. We will also look into other techniques for
generating segmentation masks from heat maps. We plan to test the
system on other types of medical data sets, including non-image
data.
6</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGEMENTS</title>
      <p>I thank Michael A. Riegler for the assistance with experiments,
methodology and writing.</p>
      <p>Medico: Transparency in Medical Image Segmentation</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Borgli</surname>
          </string-name>
          , Vajira Thambawita, Pia H Smedsrud, Steven Hicks, Debesh Jha,
          <string-name>
            <surname>Eskeland Sigrun</surname>
            <given-names>L</given-names>
          </string-name>
          , Kristin Ranheim Rand l, Konstantin Pogorelov, Mathias Lux, Duc Tien Dang Nguyen, Dag Johansen, Carsten Griwodz,
          <string-name>
            <surname>Stensland Håkon</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enrique</surname>
          </string-name>
          Garcia-Ceja, Peter T Schmidt, Hugo L Hammer,
          <article-title>Michael A Riegler, Pål Halvorsen</article-title>
          , and Thomas de Lange.
          <year>2020</year>
          .
          <article-title>HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy</article-title>
          .
          <source>Scientific Data</source>
          <volume>7</volume>
          ,
          <issue>1</issue>
          (
          <year>2020</year>
          ),
          <volume>283</volume>
          . https://doi.org/10.1038/s41597-020-00622-y
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Steven</given-names>
            <surname>Hicks</surname>
          </string-name>
          , Debesh Jha, Vajira Thambawita, Hugo Hammer, Thomas de Lange, Sravanthi Parasa,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2021</year>
          . Medico Multimedia Task at MediaEval 2021:
          <article-title>Transparency in Medical Image Segmentation</article-title>
          .
          <source>In Proceedings of MediaEval 2021 CEUR Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Lux</surname>
          </string-name>
          , Michael Riegler, Pål Halvorsen, Konstantin Pogorelov, and
          <string-name>
            <given-names>Nektarios</given-names>
            <surname>Anagnostopoulos</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Lire: open source visual information retrieval</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Multimedia Systems. 1-4.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Luke</given-names>
            <surname>Melas-Kyriazi</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>EficientNet-Pytorch</article-title>
          . https://github.com/ lukemelas/EficientNet-PyTorch.
          <article-title>(</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Adam</given-names>
            <surname>Paszke</surname>
          </string-name>
          , Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang,
          <string-name>
            <surname>Zachary</surname>
            <given-names>DeVito</given-names>
          </string-name>
          , Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and
          <string-name>
            <given-names>Soumith</given-names>
            <surname>Chintala</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>PyTorch: An Imperative Style, High-Performance Deep Learning Library</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          32,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Beygelzimer</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>d'Alché-</article-title>
          <string-name>
            <surname>Buc</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Fox</surname>
          </string-name>
          , and R. Garnett (Eds.). Curran Associates, Inc.,
          <fpage>8024</fpage>
          -
          <lpage>8035</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine Learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          ),
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>KVASIR: A Multi-Class Image Dataset for Computer Aided Gastrointestinal Disease Detection</article-title>
          .
          <source>In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys'17)</source>
          . ACM, New York, NY, USA,
          <fpage>164</fpage>
          -
          <lpage>169</lpage>
          . https://doi.org/10.1145/3083187.3083212
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Martha Larson, Mathias Lux, and
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Kofler</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>How 'how' reflects what's what: content-based exploitation of how users frame social images</article-title>
          .
          <source>In Proceedings of the 22nd ACM international conference on Multimedia</source>
          .
          <volume>397</volume>
          -
          <fpage>406</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Olaf</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Fischer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>
          . In International Conference on
          <article-title>Medical image computing and computer-assisted intervention</article-title>
          . Springer,
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Ramprasaath R Selvaraju</surname>
          </string-name>
          , Michael Cogswell,
          <string-name>
            <surname>Abhishek Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ramakrishna Vedantam</surname>
            , Devi Parikh, and
            <given-names>Dhruv</given-names>
          </string-name>
          <string-name>
            <surname>Batra</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Grad-cam: Visual explanations from deep networks via gradient-based localization</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          . 618-
          <fpage>626</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>