<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Matching the Clinical Reality: Accurate OCT-Based Diagnosis From Few Labels</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valentyn Melnychuk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evgeniy Faerman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilja Manakov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Seidl</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer Institute for Integrated Circuits IIS</institution>
          ,
          <addr-line>Erlangen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ludwig Maximilian University of Munich</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Unlabeled data is often abundant in the clinic, making machine learning methods based on semi-supervised learning a good match for this setting. Despite this, they are currently receiving relatively little attention in medical image analysis literature. Instead, most practitioners and researchers focus on supervised or transfer learning approaches. The recently proposed MixMatch and FixMatch algorithms have demonstrated promising results in extracting useful representations while requiring very few labels. Motivated by these recent successes, we apply MixMatch and FixMatch in an ophthalmological diagnostic setting and investigate how they fare against standard transfer learning. We find that both algorithms outperform the transfer learning baseline on all fractions of labelled data. Furthermore, our experiments show that Mean Teacher, which is a component of both algorithms, is not needed for our classification problem, as disabling it leaves the outcome unchanged. Our code is available online: gitlab.com/Valentyn1997/oct_diagn_semi_supervised.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Semi-supervised image classification</kwd>
        <kwd>Transfer learning</kwd>
        <kwd>Optical Coherence Tomography</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>annotation, etc.) result in exponentially more labelling
expenses. Additionally, most clinics lack the tools to
In recent years deep learning techniques have taken label vast amounts of data. Secondly, and perhaps more
the field of AI by storm. Virtually all state-of-the-art fundamentally, there is an epistemic problem in
genersystems in computer vision (CV) rely on some form ating accurate labels. For any given diagnostic problem,
of deep learning. This paradigm shift has sparked the the inter-expert agreement is well below 100%. This
imagination of many practitioners and researchers in discrepancy stems from the fact that medicine is
comthe medical image analysis domain. Computer-aided plex and does not always fit neatly into a classification
diagnosis appeared to be next-in-line to benefit from formulation. Additionally, each expert comes with his
the advancements made in CV, as the amount of data in or her own set of experiences and knowledge.
clinical diagnostics is increasing rapidly. The research Instead of solely relying on supervised learning,
community has proposed a plethora of new algorithms semi-supervised learning (SSL) should discover the bulk
and systems for the automated diagnosis of a wide of the knowledge required for solving a diagnostic task
range of diseases. However, clinical adoption has been on its own, with labels only serving as additional
guidslow. One crucial reason is that supervised learning, ance. The idea of SSL is to train a machine learning
which forms the basis for the vast majority of deep algorithm on vast amounts of unlabeled data and a
learning approaches, is ill-suited to the medical domain. small set of labelled samples. SSL is a much better</p>
      <p>
        This mismatch is two-fold. For one, the labelled data match for the clinical setting, as unlabeled data is
ofneeded for supervised learning is prohibitively costly ten abundant since it is acquired as part of the clinical
to generate for medical applications. With a shortage of routine.
medical practitioners, diverting medical experts’ time In this work, we apply two recently proposed SSL
and energy to labelling eforts becomes exceedingly ex- methods, MixMatch [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and FixMatch [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], to a
diagpensive. More fine-grained problem formulations (e.g. nostic problem in ophthalmology. We test which
persingle-label vs. multi-label, volume level vs. slice level forms better in classifying optical coherence
tomography (OCT) b-scans into four classes (one healthy and
PGraolcweeadyi,nIrgeslaonfdthe CIKM 2020 Workshops, October 19- 20, 2020, three pathological) at diferent fractions of labelled data.
email: v.melnychuk@campus.lmu.de (V. Melnychuk); We compare the two SSL methods to a baseline
transfaerman@dbs.ifi.lmu.de (E. Faerman); ilja.manakov@gmx.de (I. fer learning approach, similar to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. After going over
Manakov); seidl@dbs.ifi.lmu.de (T. Seidl) related work in the next section, we explain the basis
(oTr.cSiedi:d0l)000-0002-2401-6803 (V. Melnychuk); 0000-0002-4861-1412 for our experiments in Section 3, covering MixMatch,
© 2020 Copyright for this paper by its authors. Use permitted under Creative FixMatch and the transfer learning baseline. In Section
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmUmoRns WLiceonrsekAsthtriobuptioPnr4o.0cIneteerdnaitniognasl ((CCC EBYU4R.0)-.WS.org) 4 we describe the dataset and present the results of our
investigation. We conclude with Section 5 by summa- of labels. Yosinski et al. [21] discovered how unfreezing
rizing our findings and discussing how they apply to diferent parts of the network while fine-tuning afects
the clinical setting. the target performance.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>
        Semi-supervised learning. State-of-the-art meth- Transfer learning and semi-supervised learning are two
ods for image classification concentrate on finding the main approaches for predictive modelling when dealing
right combination of SSL paradigms. One of the early with data with few labels. Transfer learning approaches
approaches – Mean-Teacher [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] – uses exponential reuse knowledge from previously learned tasks. On the
moving average (EMA) of model parameters. Virtual other hand, the SSL approaches allow learning with
Adversarial Training (VAT) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], tries to find a minimal small labelled datasets by utilizing unlabeled data from
perturbation and fit a robust model against it. Mix- the same distribution in the learning process. In the
folMatch [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and RealMix [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] encompass mixing and over- lowing, we first discuss our transfer learning baseline
laying labelled and unlabelled images to obtain con- and afterwards describe the SSL approaches we have
sistent predictions. Unsupervised Data Augmentation chosen for this study.
(UDA) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] uses strongly augmented images to force
consistency among unlabeled images. ReMixMatch [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] 3.1. Transfer Learning
uses so-called “augmentation anchoring”, i.e. strong
and weak augmentations, to enforce consistency. In- When applying transfer learning techniques, the user
spired by UDA and ReMixMatch, the authors of Fix- has to choose how to adapt the model from the
auxilMatch [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] significantly simplify SSL by relying only on iary to the primary task. In our experiments, we use
augmentations and pseudo-labelling with a confidence a network, which was pre-trained on ImageNet [20].
threshold. We provide a broader overview of applied For adapting the model to OCT classification we try
SSL methods in Section 3.2. two common approaches. In the feature extraction
ap
      </p>
      <p>
        Surprisingly, there exists only a little amount of liter- proach, we freeze all parameters except for the final
ature on SSL applied to ophthalmological data. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and fully connected layer, analogous to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Alternatively,
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] utilize SSL for OCT segmentation. In the domain of we use the pre-trained network as initialization and
automated diagnosis, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] employ an autoencoder with allow all parameters to change. We refer to this as
an additional classification module on the latent code in ifne-tuning hereafter.
the detection of retinopathy from colour fundus images.
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] tackle the same problem by extending the GAN 3.2. Semi-supervised Learning
framework [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to one “fake” and six “real” classes, i.e.
the labeled classes. Recent works [14] and [15] apply In our study we compare two of recent
state-of-thethe same principle to the classification of OCT b-scans. art algorithms for SSL MixMatch [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and FixMatch [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Most recently, [16] applied SSL methods to glaucoma Both algorithms combine several pre-existing
techdetection by imputing missing visual field (VF) mea- niques from SSL. In this chapter, we review the main
surements through nearest-neighbour identification in ideas and compare their utilization in both algorithms.
the latent space of a pre-trained classification CNN. We refer the reader to Appendix A for the detailed
Afterwards, [16] train a multi-task network jointly on algorithm descriptions.
glaucoma classification and VF measurement
prediction. To the best of our knowledge, we are the first to Data Augmentation. Data Augmentation is a
reguapply consistency regularization based SSL techniques larization technique which is often used in supervised
(see Section 3) to the problem of automated diagnosis learning. The goal is that the model’s prediction is
in ophthalmology. not afected by the certain transformation of data
instances. Therefore additional training data is added to
Transfer learning. Among numerous approaches the dataset by applying various perturbations to the
existing in the deep transfer learning [17], we choose data while keeping original labels. Most of the data
the fine-tuning or network-based transfer learning to augmentations are domain-specific and require domain
be the most promising. [18, 19] proposed to use Ima- knowledge.
geNet [20] pre-trained CNN as the initialization for
different visual recognition tasks with the limited amount
      </p>
      <sec id="sec-3-1">
        <title>MixMatch uses random flip-and-shift augmentations</title>
        <p>(horizontal flips and random crops) for both labelled
and unlabeled data.</p>
        <p>FixMatch distinguishes between weak and strong
data augmentations. Flip-and-shift augmentations are
considered as weak augmentations, whereas afine
trasformations and color-jittering are examples of
strong augmentations (originally – 14 diferent
transformations from RandAugment [22]).
makes the predictions for the training data and is
updated based on the training loss. Both MixMatch and</p>
      </sec>
      <sec id="sec-3-2">
        <title>FixMatch employ Mean Teacher for the computation</title>
        <p>of pseudo-labels. Note, that keeping a second model in
memory and updating its parameters results in higher
memory requirements and computation costs.</p>
        <sec id="sec-3-2-1">
          <title>MixUp. MixUp [26] is another regularization tech</title>
          <p>nique to avoid overfitting. MixUp linearly combines
training instance pairs and their prediction targets.</p>
          <p>Pseudo-Labelling. Pseudo-Labeling or self-training Therefore it tries to impose linear behaviour between
loss [23] is the process of using the trained model to training samples. MixMatch does not diferentiate
beobtain labels for unlabeled instances. The predicted tween pseudo targets predicted for the unlabeled
inlabels are used to guide the further learning process, stances and ground truth labels and mixes all
possie.g. by using generated labels as new targets. ble target pairs. Therefore a resulting instance used</p>
          <p>MixMatch applies diferent augmentations for an un- in training may be a combination of two pseudo
tarlabeled instance and computes the class distribution gets, two ground truth labels or of pseudo-target with
for each augmentation. Therefore, instead of hard one ground truth label.
hot label MixMatch defines a probability distribution
as the target. To sharpen the distribution and to reduce
its entropy, the temperature of distribution is adjusted 4. Experiments
[24].</p>
          <p>FixMatch uses a “classic” version of pseudo-labelling
with hard labels and fixed confidence. The class
probability distribution is taken from model outputs after
a weak augmentation. If the probability of the most
probable class exceeds a predefined threshold the label
is assigned to a strongly augmented version of the same
instance and used in the loss calculation.</p>
          <p>Our work follows the principles of the fair SSL
evaluation framework, defined by [ 27]. The authors highlight
the importance of using the same classifying model
structure for comparison. The evaluation is also
meaningful for the real use-case if SSL methods are
compared with well-fine-tuned transfer learning and fully
supervised models.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>For the evaluation we use the UCSD dataset pub</title>
          <p>
            lished by Kermany et al. [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]. It contains 84,495
optiConsistency Regularization. Consistency regular- cal coherence tomography (OCT) b-scans pertaining
ization [25] imposes the constraint that the model to four categories; “normal”, “drusenoid” (DRUSEN),
should make similar predictions for the same instance “choroidal neovascularization” (CNV) and “diabetic
under diferent data augmentations. Both MixMatch macular edema” (DME). The images vary in size, where
and FixMatch apply data augmentation on labelled and the median image has a size of 496×512 pixels. The
unlabeled data and enforce similar prediction for the height of the images ranges between 496 and 512 and
same instance under diferent augmentations. For the the width between 384 and 1536. The dataset is also
obunlabeled instances, the pseudo-label is used as a target. tainable through Kaggle1. For better comparability, we
FixMatch uses soft augmentations to compute pseudo- use the same train/validation/test split as in the Kaggle
labels for hard augmentations of the same training challenge. There are several images for each patient
sample. in the dataset and splits are done patient-wise, there
are no images of the same patients in diferent splits.
          </p>
          <p>
            Mean Teacher. Another popular consistency re- Test and validation are balanced, there are 8 and 242
quirement in SSL is a similar prediction over time or images per class respectively (see Fig. 1). In our
experipunishing the behaviour when the model changes its ments, we vary the number of labelled data, which we
decisions rapidly. The Mean Teacher algorithm [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ] sample randomly from the training subset. We sample
maintains two models. The teacher model stores an the same number of labelled training instances from
exponential moving average of student’s parameters each class. For SSL approaches the rest of the train set
and is used to make the predictions to compute the is used as unlabeled data.
pseudo-labels. Therefore pseudo-labels computed by We compare the performance of transfer learning
the teacher can be considered as a weighted combina- and SSL models using the same Wide ResNet-50-2
tion of decisions of previous models. The student model
30000
t
un20000
o
c
10000
0
          </p>
          <p>Train subset
CNV</p>
          <p>E
DM</p>
          <p>DRUSEN</p>
          <p>NORMAL</p>
          <p>CNV</p>
          <p>E
DM</p>
          <p>DRUSEN</p>
          <p>NORMAL</p>
          <p>CNV</p>
          <p>E
DM</p>
          <p>DRUSEN</p>
          <p>NORMAL
DRUSEN</p>
          <p>NORMAL
[28] backbone. Since the images are monochrome we Method
duplicate the channel three times for RGB channels.</p>
          <p>
            For each model, we perform hyperparameter search,
described in Appendix B.1 and B.2. For all experiments, Alqudah [29]
we report the model performance on test data in the
epoch with the lowest validation loss.
4.1. Comparison of transfer learning
and SSL approaches
 
Kermany et al. [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] All 96.6%
          </p>
          <p>Accuracy</p>
          <p>Notes</p>
          <p>All 97.1%
Wu et al. [30] All 97.5%
Chetoui et al. [31] All 98.46%
Tsuji et al.[32] All 99.6%
WideResNet-50-2 All 99.69%
(our backbone)</p>
          <p>Original paper
Extended UCSD
with 5 classes
With EMA decay
( EMA = 0.999)
First, in Table 1 we compare the performance of our He et al. [14] 835 87.25±1.44% * *Average precision
backbone model trained with all labelled instances to
the results reported previously in the literature for the Table 1
same UCSD dataset. As we can see, the backbone model Reported test accuracies for UCSD dataset. Methods have
achieves almost perfect performance when trained with tdhifeerepnrtopboascekdboSnLeLs
amnedththoudss.areNnevoetrftuhlelylecsos,mopuarrabbelsetwfiutlhlyenough labels. supervised model outperforms previously reported
meth</p>
          <p>Next, in Fig. 2b we compare two transfer learning ods.
approaches. Note, that the hyperparameter search was
done for each number of labels for each approach. We
discover that, contrary to our expectations, the fine- almost perfect performance. The Fix-Match algorithm
tuning variant outperforms feature extraction approach also outperforms Mix-Match in almost all settings and
in all label settings. We believe that a thorough selec- with only 50 labelled points per class achieves the
action of hyperparameters with representative validation curacy of 98.14%. We also observe a small SSL
perforset reduces the risk of overfitting. Furthermore, since mance drop for 25 labelled images per class – mainly
the original models are trained on the dataset with RGB because methods require even more epochs to fit (we
channels, we believe that the model can better adapt to employ a heuristical formula for defining the maximum
the monochrome setting when all model weights are number of epochs based on the number of labels, see
allowed to be changed. Appendix B.2, 3).</p>
          <p>In the Fig. 2a we present the results of both SSL Finally, since practitioners have often to deal with
algorithms and compare them with the best perform- the resource constraints and actual running times are
ing transfer learning setting. We find that the SSL ap- rarely reported in the literature, we report them in
proaches outperform transfer learning on all fractions Table 2. Note, that all methods are implemented in
of labelled data. The gap between SSL and transfer the same framework and the experiments are done on
learning widens significantly for smaller fractions of the same machine with two Tesla V100 Nvidia GPUs.
labelled data. With only 10 labelled representatives per To use the same batch size as recommended in the
class, the FixMatch achieves an accuracy of over 86%, original publications, we have used both GPUs to train
while transfer learning reaches only 59%. We also see, Fix-Match. Other models are trained on a single GPU.
that with about 2000-4000 labels all methods achieve
1.0
0.9
y
c
ra
ccu0.8
tea
s
T0.7
0.6
1.0
0.9
y
c
a
rcu0.8
c
ta
s
eT0.7
0.6
 
Transfer Learning 10m 9m 12m 15m 24m 39m 1h 39m
Mix-Match 1d 16h 5m 9h 12m 6h 13m 2h 30m 2h 37m 2h 24m 2h 26m
FixMatch 5d 9h 36m 1d 19h 4m 1d 40m 9h 58m 10h 40m 9h 50m 7h 51m</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion</title>
      <sec id="sec-4-1">
        <title>In this work, we have demonstrated the eficacy of</title>
        <p>MixMatch and FixMatch, when applied to an
ophthalmological diagnostic problem on OCT data. The two
algorithms were able to attain high accuracy, achieving
well over 80% on as little as 40 labelled samples (i.e.
ten per class). Both algorithms outperformed transfer
learning in the few labelled data settings. This study
emphasizes the use of SSL methods in the clinical
adoption of AI. Although both MixMatch and FixMatch are
more computationally expensive than transfer learning,
the amount of labelling efort saved by using them is
immense. With labelling being one of the biggest
factors hindering clinical use of AI methodology, we argue
that smarter use of the abundance of unlabeled data
already present at the clinic will be a major strategy
for overcoming this hurdle.</p>
      </sec>
      <sec id="sec-4-2">
        <title>As part of future work, we propose to also compare</title>
      </sec>
      <sec id="sec-4-3">
        <title>SSL approach with the few-shot deep learning.</title>
        <p>(a) Best models, maximum performance among 4 runs
(Semi-supervised) / 8 runs (Transfer learning) per
each  
40
100 200 800 2000 4000
nl - number of labelled datapoints, log-scaled
20000
(b) Study of layers freezing for Transfer learning,
maximum performance among 4 runs per each  
4.2. Mean Teacher</p>
        <sec id="sec-4-3-1">
          <title>The Mean Teacher is inherent part of Fix-Match algo</title>
          <p>rithm and is also optionally recommended for
MixMatch. We observe learning curves to be more stable
for both train and validation subsets for all the models
when models are trained using it. However, we assume
that with the right chosen validation subset, the
variability could be advantageous and one can find a better
ift. Usage of Mean Teacher causes additional
computation and memory costs and as can be seen in Table 3
most of the time models without it perform better.
0.0
0.999
0.0
Acknowledgments
denotes the set of weak augmen- do not need the linear ramp-up for   .
As we use  for filtering confident pseudo-labels, we</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Appendix A. MixMatch &amp; FixMatch – algorithm details</title>
      <p>Let  = {(  ,   ),  ∈ (1, ...,  )} be the batch of
laand MixUp [26]. Let 
tations and (⋅) – strong augmentations.  ̂ =  M( ;  ))
is the prediction of backbone classifier, parametrized
is unsupervised loss weight.
by  .</p>
      <p>(⋅, ⋅) denotes categorical cross-entropy and  
MixMatch employs only weak augmentations  (⋅)
= {  ,  ∈ (1, ...,  )} be the
unlabeled data batch. The model outputs of 
dom weak augmentations  (⋅) of the same unlabelled
sample are treated as soft pseudo-labels   . These soft
pseudo-labels are averaged and sharpened with the
rantemperature  for each image in  to yield a
pseudolabel for that image. Then, images from both randomly
augmented  and  are concatenated and shufled,
resulting in set  . Afterwards, samples in  and  are
weakly augmented and linearly interpolated with
samples from  . This results in ̂ and ̂ – "mixed-up"
versions of augmented labelled and 
Coeficients of MixUp are sampled from
unlabelled batches.
tribution. The final loss is the sum of categorical
crossifnal unsupervised part of the loss. They are filtered
with the threshold  . The loss of FixMatch is then the
sum of two categorical cross-entropies for labelled and
unlabelled images:
 =
1</p>
      <p>∑
 (, )∈
+</p>
      <p>∑
(,, ̂ )∈̂</p>
      <p>H(,  M( ( );  ))+</p>
      <p>1max( )&gt; H( ̂ ,  M(( );  )) (2)</p>
    </sec>
    <sec id="sec-6">
      <title>B. Experiments</title>
      <p>B.1. Transfer learning
We took a version of Wide ResNet-50-2 pre-trained
on ImageNet from PyTorch.2 Transfer learning was
much computational budget:
ifne-tuned for every individual   , as it did not require
• learning rate ∈ {1 ∗ 10−3, 5 ∗ 10−4}
• optimizer weight decay ∈ {0.0, 0.0001}
• layers freezing ∈ {Fine-tuning,</p>
      <p>Feature extraction} (see Section 3.1)
Beta( , 
) dis- use Adam optimizer [33],</p>
      <p>= 32, number of epochs =
50. Additionally, early stopping with the patience of 25</p>
      <p>Further hyperparameters are kept fixed, namely we
entropy for images from ̂ (supervised part) and Brier epochs was applied to avoid overfitting.
score for ̂ images (unsupervised part):
 =

1</p>
      <p>∑
(, )∈̂</p>
      <p>H(,  M( ;  ))+
+</p>
      <p>∑
(, )∈̂
|| −  M( ;  )||2, (1)</p>
      <p>2
MixMatch linearly ramps up  
after each batch to reduce the influence of unsupervised</p>
      <p>from 0 to its maximum
part during early stages of training.</p>
      <p>FixMatch is a more simplified method. Unlabeled
data batch 
= {  ,</p>
      <p>∈ (1, ..., 
bigger. Given the model’s prediction
augmented unlabelled sample   , method yields hard
pseudo-labels  ̂ = argmax(  ) and ̂ = {(  ,   ,  ̂ )}.</p>
      <p>Afterwards, the model predicts labels for both a batch
of weakly augmented labelled images and a batch of
strongly augmented unlabelled images. Only the
confident predictions for unlabelled samples are used in the</p>
      <p>)} is now  -times
for a weakly</p>
      <p>B.2. MixMatch &amp; FixMatch</p>
      <sec id="sec-6-1">
        <title>Hyperparameter fine-tuning for both SSL methods was</title>
        <p>two-fold: firstly, we fine-tuned more general
parameters on 200 labelled samples (  = 200) with respect to
the validation loss (see Table 4). Secondly, for each
spe</p>
      </sec>
      <sec id="sec-6-2">
        <title>We omit using cosine learning rate decay.</title>
        <p>cific   , we tuned subset-size-dependent parameters.</p>
        <p>The labeled batch size was  = 16 for both
algorithms. Additionally, we fix  = 4,  = 0.7 for FixMatch.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Regarding secondary fine-tuning, after the increase</title>
        <p>of   , each epoch becomes proportionally longer. Thus,
we propose the following inverse formula to define the
number of epochs:</p>
        <p>Number of epochs = round
 
(   div  )
(3)</p>
      </sec>
      <sec id="sec-6-4">
        <title>2https://pytorch.org/hub/pytorch_vision_wide_resnet/.</title>
        <p>Learning rate
 
Grid-search size
Primary hyperparameter search grid for SSL methods. Best
value is marked with bold font. SGD – stocastic gradient
by maximum number of batches in labelled subset.
descent with momentum ( = 0.9) [34]. An epoch is defined
where  
while training.</p>
        <p>denotes total number of labelled batches, used</p>
      </sec>
      <sec id="sec-6-5">
        <title>While secondary fine-tuning, we vary:</title>
        <p>•  EMA ∈ {0.0, 0.999} (EMA decay)
•   ∈ {12000, 15000} for MixMatch /
  ∈ {24000, 30000} for FixMatch</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Berthelot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Carlini</surname>
          </string-name>
          , I. Goodfellow,
          <string-name>
            <given-names>N.</given-names>
            <surname>Papernot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <article-title>Mixmatch: A holistic approach to semi-supervised learning</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>5049</fpage>
          -
          <lpage>5059</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Berthelot</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-L. Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Carlini</surname>
            ,
            <given-names>E. D.</given-names>
          </string-name>
          <string-name>
            <surname>Cubuk</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kurakin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , C. Raffel, Fixmatch:
          <article-title>Simplifying semi-supervised learning with consistency and confidence</article-title>
          , ArXiv abs/
          <year>2001</year>
          .07685 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Kermany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goldbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Valentim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Baxter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McKeown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yan</surname>
          </string-name>
          , et al.,
          <article-title>Identifying medical diagnoses and treatable diseases by image-based deep learning</article-title>
          ,
          <source>Cell</source>
          <volume>172</volume>
          (
          <year>2018</year>
          )
          <fpage>1122</fpage>
          -
          <lpage>1131</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Tarvainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Valpola</surname>
          </string-name>
          ,
          <article-title>Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results</article-title>
          ,
          <volume>40</volume>
          100
          <fpage>200</fpage>
          800
          <year>2000</year>
          4000 nl
          <article-title>- number of labelled datapoints</article-title>
          ,
          <source>log-scaled 20000 in: Advances in neural information processing neural information processing systems</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>systems</fpage>
          ,
          <year>2017</year>
          , pp.
          <fpage>1195</fpage>
          -
          <lpage>1204</lpage>
          .
          <fpage>2672</fpage>
          -
          <lpage>2680</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Miyato</surname>
          </string-name>
          , S.-i. Maeda,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ishii</surname>
          </string-name>
          , Virtual [14]
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rabbani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Retinal adversarial training: a regularization method for optical coherence tomography image classificasupervised and semi-supervised learning, IEEE tion with label smoothing generative adversarial transactions on pattern analysis and machine in- network,</article-title>
          <string-name>
            <surname>Neurocomputing</surname>
          </string-name>
          (
          <year>2020</year>
          ).
          <source>telligence 41</source>
          (
          <year>2018</year>
          )
          <fpage>1979</fpage>
          -
          <lpage>1993</lpage>
          . [15]
          <string-name>
            <surname>V. Das</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Dandapat</surname>
            ,
            <given-names>P. K.</given-names>
          </string-name>
          <string-name>
            <surname>Bora</surname>
          </string-name>
          ,
          <article-title>A data-eficient</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Alonso</surname>
          </string-name>
          , T. Beltramelli, Realmix: To-
          <article-title>approach for automated classification of oct imwards realistic semi-supervised deep learning al- ages using generative adversarial network, IEEE gorithms</article-title>
          , arXiv preprint arXiv:
          <year>1912</year>
          .
          <volume>08766</volume>
          (
          <year>2019</year>
          ).
          <source>Sensors Letters</source>
          <volume>4</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Hovy</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>T.</given-names>
            <surname>Luong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          , [16]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Ran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <article-title>Unsupervised data augmentation for consistency C. C</article-title>
          . Tham,
          <string-name>
            <given-names>R. T.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Mannil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Chetraining</surname>
          </string-name>
          , arXiv: Learning (
          <year>2019</year>
          ). ung, P. A.
          <string-name>
            <surname>Heng</surname>
          </string-name>
          ,
          <article-title>Towards multi-center glaucoma</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Berthelot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Carlini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Cubuk</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Kurakin,
          <article-title>OCT image screening with semi-supervised joint</article-title>
          K. Sohn,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. Rafel,
          <article-title>Remixmatch: Semi- structure and function multi-task learning, Medsupervised learning with distribution alignment ical Image Analysis 63 (</article-title>
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1016/j. and augmentation anchoring,
          <source>arXiv preprint media</source>
          .
          <year>2020</year>
          .
          <volume>101695</volume>
          . arXiv:
          <year>1911</year>
          .
          <volume>09785</volume>
          (
          <year>2019</year>
          ). [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          , C. Liu,
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>A survey on deep transfer learning</article-title>
          , in: InternaJ. Liu,
          <article-title>Semi-supervised automatic segmentation of tional conference on artificial neural networks, layer and fluid region in retinal</article-title>
          optical coherence Springer,
          <year>2018</year>
          , pp.
          <fpage>270</fpage>
          -
          <lpage>279</lpage>
          .
          <article-title>tomography images using adversarial learning</article-title>
          , [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Oquab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bottou</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Laptev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sivic</surname>
          </string-name>
          ,
          <source>Learning IEEE Access 7</source>
          (
          <year>2018</year>
          )
          <fpage>3046</fpage>
          -
          <lpage>3061</lpage>
          .
          <article-title>and transferring mid-level image representations</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sedai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Antony</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ishikawa</surname>
          </string-name>
          ,
          <article-title>using convolutional neural networks</article-title>
          , in: ProceedJ. Schuman,
          <string-name>
            <given-names>W.</given-names>
            <surname>Gadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Garnavi</surname>
          </string-name>
          ,
          <article-title>Uncertainty ings of the IEEE conference on computer vision guided semi-supervised segmentation of retinal and pattern recognition</article-title>
          ,
          <year>2014</year>
          , pp.
          <fpage>1717</fpage>
          -
          <lpage>1724</lpage>
          . layers in oct images, in: International Confer- [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Donahue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Hofman,</surname>
          </string-name>
          <article-title>ence on Medical Image Computing</article-title>
          and
          <string-name>
            <given-names>Computer- N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , E. Tzeng, T. Darrell, Decaf: A deep
          <string-name>
            <surname>Assisted</surname>
            <given-names>Intervention</given-names>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>282</fpage>
          -
          <lpage>290</lpage>
          .
          <article-title>convolutional activation feature for generic vi-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <article-title>Semi-supervised adver- sual recognition, in: International conference on sarial learning for diabetic retinopathy screening</article-title>
          ,
          <source>machine learning</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>647</fpage>
          -
          <lpage>655</lpage>
          . in: International Workshop on Ophthalmic Medi- [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Feical Image Analysis</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>60</fpage>
          -
          <lpage>68</lpage>
          . Fei, ImageNet:
          <string-name>
            <given-names>A</given-names>
            <surname>Large-Scale Hierarchical</surname>
          </string-name>
          Image
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wan</surname>
          </string-name>
          , G. Chen,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lei</surname>
          </string-name>
          , Retinopathy Database, in: CVPR09,
          <year>2009</year>
          .
          <article-title>diagnosis using semi-supervised multi-channel [21]</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Yosinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clune</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lipson</surname>
          </string-name>
          ,
          <article-title>How generative adversarial network, in: International transferable are features in deep neural networks?</article-title>
          , Workshop on Ophthalmic Medical Image Analy- in
          <source>: Advances in neural information processing sis</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>182</fpage>
          -
          <lpage>190</lpage>
          . systems,
          <year>2014</year>
          , pp.
          <fpage>3320</fpage>
          -
          <lpage>3328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pouget-Abadie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mirza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          , [22]
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Cubuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          , RanD. Warde-Farley,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ozair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <article-title>Ben- daugment: Practical automated data augmentagio, Generative adversarial nets, in: Advances in tion with a reduced search space</article-title>
          , arXiv: Com-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>