<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The impact of averaging logits over probabilities on ensembles of neural networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cedrique Rovile Njieutcheu Tassi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jakob Gawlikowski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Auliya Unnisa Fitri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rudolph Triebel</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>German Aerospace Center (DLR), Institute of Data Science</institution>
          ,
          <addr-line>Mälzerstraße 3-5, 07745 Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>German Aerospace Center (DLR), Institute of Optical Sensor Systems</institution>
          ,
          <addr-line>Rutherfordstraße 2, 12489 Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>German Aerospace Center (DLR), Institute of Robotics and Mechatronics</institution>
          ,
          <addr-line>Münchener Straße 20, 82234 Wessling</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Model averaging has become a standard for improving neural networks in terms of accuracy, calibration, and the ability to detect false predictions (FPs). However, recent findings show that model averaging does not necessarily lead to calibrated confidences, especially for underconfident networks. While existing methods for improving the calibration of combined networks focus on recalibrating, building, or sampling calibrated models, we focus on the combination process. Specifically, we evaluate the impact of averaging logits instead of probabilities on the quality of confidence (QoC). We compare combined logits instead of probabilities of members (networks) for models such as ensembles, Monte Carlo Dropout (MCD), and Mixture of Monte Carlo Dropout (MMCD). Comparison is done using experimental results on three datasets using three diferent architectures. We show that averaging logits instead of probabilities increase the confidence thereby improving the confidence calibration for underconfident models. For example, for MCD evaluated on CIFAR10, averaging logits instead of probabilities reduces the expected calibration error (ECE) from 12.03% to 5.44%. However, the increase in confidence can bring harm to confidence calibration for overconfident models and the separability between true predictions (TPs) and FPs. For example, for MMCD evaluated on MNIST, the average confidence on FPs due to the noisy data increases from 51.31% to 94.58% when averaging logits instead of probabilities. While averaging logits can be applied with underconfident models to improve the calibration on test data, we suggest to average probabilities for safety- and mission-critical applications where the separability of TPs and FPs is of paramount importance.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Model averaging</kwd>
        <kwd>Combination process</kwd>
        <kwd>Logit averaging</kwd>
        <kwd>Probability averaging</kwd>
        <kwd>Ensemble</kwd>
        <kwd>Monte Carlo Dropout (MCD)</kwd>
        <kwd>Mixture of Monte Carlo Dropout (MMCD)</kwd>
        <kwd>Quality of confidence (QoC)</kwd>
        <kwd>Confidence calibration</kwd>
        <kwd>Separating true predictions (TPs) and false predictions (FPs)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>produce more underconfident networks. For example, [ 7]
showed that averaging networks trained with modern
Recently, averaging the predictions of multiple stochas- regularization techniques resulted in more
underconfitic or deterministic networks has become a standard ap- dent networks and therefore miscalibrated predictions.
proach for improving accuracy [1, 2] and uncertainty [12] supported this argument by theoretically and
empirestimates [3]. Generally, the quality of uncertainty es- ically showing that averaging calibrated networks do not
timates (e.g.: QoC) is assessed by the degree of calibra- always lead to calibrated confidences. Calibrating
confition and/or the ability to detect FPs. Model averaging dences of averaged networks has received little attention
can yield well-calibrated confidence [ 4, 5] and is one of in the literature. Generally, post-processing calibration
the state-of-the-art methods for detecting FPs caused methods, such as temperature scaling [13], can be used
by out-of-distribution examples [4, 3]. However, recent to recalibrate the confidences of averaged networks, as
ifndings [ 6, 7, 8] show that model averaging does not nec- demonstrated in [8, 12]. From [14] and further supported
essarily lead to calibrated confidence, especially when by [8], confidence calibration in model averaging is
corthe networks are built using modern regularization tech- related to diversity inherent in individual networks and
niques, such as mixup [9] or label smoothing [10, 11]. the more diverse the networks, the better the calibration.
This is because modern regularization techniques can Motivated by this observation, [14] promoted model
di(strongly) regularize networks, resulting in underconfi- versity using structured dropout to reduce calibration
dence. Furthermore, averaging underconfident networks errors. [7] proposed class-adjusted mixup that trains
The IJCAI-ECAI-22 Workshop on Artificial Intelligence Safety less confident networks by evaluating the diference
be(AISafety 2022) tween accuracy (estimated on a validation dataset after
$ Cedrique.NjieutcheuTassi@dlr.de (C. R. N. Tassi); each training epoch) and the confidence of each
trainJakob.Gawlikowski@@dlr.de (J. Gawlikowski); Auliya.Fitri@dlr.de ing sample to activate or deactive mixup training for
(A. U. Fitr©i)2;02R2 Cuodpyorilgphth20.T22rfoiertbhiespl@aperdblyri.tdsaeuth(oRrs.. UTserpieerbmiettle)d under Creative overconfidence (average confidence &gt; accuracy) or
unCPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org) derconfidence (average confidence &lt; accuracy),
respectively. All these methods for improving the calibration TPs and FPs. Therefore, FPs can be made with high
confiof combined networks focus on recalibrating, building, dence similar to TPs. For example, for MMCD evaluated
or sampling the calibrated networks. However, this work on FashionMNIST (see Table 4), the average confidence
focuses on combining the networks. Specifically, we ad- on FPs due to the noisy data increased from 51.31% to
dress the question: What is the impact of averaging 94.58% when averaging logits instead of probabilities. In
logits instead of probabilities of multiple (stochastic summary, we provide empirical evidence demonstrating
or deterministic) networks on the QoC? how combining logits instead of probabilities of multiple</p>
      <p>We hypothesized that averaging logits instead of prob- (stochastic or deterministic) networks
abilities of multiple networks increases the confidence
of the averaged network. This is because logits (inputs • preserves accuracy, but increases the confidence
to softmax), which can be interpreted as found evidence on TPs and FPs.
for possible classes [15], are continuous values normal- • reduces the calibration error (given
underconfiized using the softmax to produce discrete probabilities. dent networks), but increases the calibration error
The softmax normalization of continuous values (log- (given overconfident networks).
its) to discrete values (probabilities) causes information • can harm the separability between TPs and FPs.
loss and possible robustness to changes in the
magnitudes of logits. This implies that the softmax function 2. Related works
is a nonlinear function that maps multiple logit vectors
with large diferences in magnitudes to the same discrete The combination process describes how multiple
memprobability vector. We evaluated the impact of the in- bers are combined and the information type (e.g., logits
crease in confidence caused by averaging logits instead or probabilities) that is combined. Several approaches
of probabilities on the QoC. Specifically, we evaluated the such as stacking [16] and voting [17, 18, 19]) have been
QoC by assessing the degree of confidence calibration, reported for aggregating multiple predictions. Some of
which measures the diference between the predicted these approaches have been reviewed and discussed in
(average confidence) and true probabilities (empirical ac- [20, 21] and experimentally compared in [18, 16] to find
curacy). Furthermore, we evaluated the QoC by assessing the one with the best accuracy. It was found that one
its ability to seperate TPs and FPs. To provide empirical approach improves accuracy better than another
dependevidence for evaluating the QoC, we considered the logit ing on several factors, such as the number of members,
averaging against probability averaging and compared diversity inherent in individual members, and accuracy
both approaches using diferent averaged models, such of individual members. However, in [22], we compared
as ensemble, MCD, and MMCD. The comparison was approaches such as averaging, plurality voting, or
majorbased on results from diefrent experiments conducted ity voting to find the one that better captures uncertainty.
on three datasets, namely, MNIST, FashionMNIST, and We found that the averaging approach captures
uncerCIFAR10 evaluated on VGGNet, ResNet, and DenseNet, tainty better than voting approaches. Before our work,
respectively. [23] argued that simple averaging approaches are more</p>
      <p>Results show that averaging logits instead of probabil- robust than voting approaches. This argument was
furities preserves accuracy, but increases confidence. For ther supported by [24]. This is because the averaging
example, for MCD evaluated on CIFAR10 (see Table 2), approach considers all members’ predictions, whereas
the accuracy remained around 85.36% while the aver- plurality/majority voting ignores uncertain predictions
age confidence increased from 73.35% to 80.04% when and therefore, reduces the uncertainty in the combined
we averaged logits instead of probabilities. Furthermore, members’ prediction. Although various combination
apgiven underconfident models, the increase in the degree proaches have been presented and compared in the
literof confidence reduces the calibration error on the test ature, the information type that is combined has received
data. For example, for MCD evaluated on CIFAR10, ECE relatively little attention. [25] showed that averaging
dropped from 12.04% to 5.40% when the average confi- quantiles rather than probabilities improve the
predicdence increased from 73.35% to 80.04%. However, given tive performance. Generally, for neural networks and
overconfident models, the increase in the degree of con- classification problems in particular, multiple members
ifdence increased the calibration error on the test data. (networks) are combined by averaging probabilities [16].
For example, for the ensemble evaluated on CIFAR10 (see [16] evaluated the impact of combining logits instead of
Table 3), ECE increased from 3.03% to 7.40% when the av- probabilities on accuracy, however, the impact on the
erage confidence increased from 89.43% to 96.17%. Finally, QoC remains unclear. Thus, we investigated the impact
for underconfident or overconfident models, the increase of combining logits instead of probabilities on the QoC.
in the degree of confidence can harm the separability
between TPs and FPs. This is because averaging logits
instead of probabilities increases the confidence of both</p>
    </sec>
    <sec id="sec-2">
      <title>3. Background</title>
      <p>In the context of image classification, let the training
data  = { ∈ R×  ×  ,  ∈   }∈[1,] be a
realization of independently and identically distributed
random variables (, ) ∈  ×  , where  denotes
the ℎ input and  its corresponding one hot encoded
class label from the set of standard unit vectors of R ,
  .  and  denote the input and label spaces.  ×
 ×  denotes the dimension of input images, where
,  , and  refer to the height, weight, and number of
channels, respectively.  and  denote the numbers of
possible output classes and samples within the training
data, respectively.</p>
      <sec id="sec-2-1">
        <title>3.1. Convolutional neural network (CNN)</title>
        <sec id="sec-2-1-1">
          <title>A CNN is a nonlinear function  parameterized by model</title>
          <p>parameters  , called the network weights. Here, it maps
input images  ∈ R×  ×  to class labels  ∈   ,
 :  ∈ R×  ×  →  ∈ [0, 1] ;  () =  (1)</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>The network parameters are optimized on the train</title>
          <p>ing dataset, . Given a new data sample  ∈
R×  ×  , a trained CNN  predicts the corresponding
target  =  () using the set of trained weights  . The
network output (logit) is given by  =  (), from which
a probability vector (|, ) =  (), can
be computed. In the following, this probability
vector will be abbreviated by  and its entries by  with
 = 1, . . . ,  and ∑︀</p>
          <p>=1  = 1. Further, we get the
predicted confidence  = max() and predicted class
label  = arg max()</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. Monte Carlo Dropout (MCD)</title>
        <p>MCD was investigated in [26, 27, 28] for uncertainty
estimation. It is one of the most widespread Bayesian
methods reviewed in [3]. It approximates the
prediction (|, ) using the mean of  stochastic
forward passes, (|,  1), ..., (|,   ), representing 
stochastic CNNs parameterized by samples  1,  2,..., and
  . That is
(|, ) ≈
1 ∑︁ (|,  ) ≈
 =1
1 ∑︁   (). (2)
 =1
Specifically, MCD approximates the prediction with a
dropout distribution realized by sampling weights with
masks drawn from known distributions, such as
Gaussian, Bernoulli, or a cascade of Gaussian and Bernoulli
distributions [22]. For example, given the activation
vector  fed to a MCD layer (placed for example at the input
of the first fully-connected layer) and assuming that
sampling is realized with masks drawn from a cascade of</p>
        <sec id="sec-2-2-1">
          <title>Gaussian and Bernoulli distribution, the MCD layer sam</title>
          <p>ples the ℎ element of  as  =  *   *   with
  ∼  (1,  2 = /(1 − )) and   ∼ ().
Here,  denotes the dropout probability. In this work, we
can refer to MCD as an average of  stochastic CNNs.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>3.3. Ensemble</title>
        <sec id="sec-2-3-1">
          <title>An (explicit) ensemble was investigated in [4, 27, 28] for</title>
          <p>uncertainty estimation. It approximates the prediction
(|, ) by learning diferent settings. Given a set
of CNNs   for  ∈ 1, 2, ...,  , the ensemble
prediction is obtained by averaging over the predictions of the
CNNs. That is,
(|, ) :=
:=</p>
          <p>1 ∑︁ (|,  )

=1

1 ∑︁   ().

=1</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>In this work, we can refer to an ensemble as an average</title>
          <p>of  deterministic CNNs.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>3.4. Mixture of Monte Carlo Dropout (MMCD)</title>
        <p>MMCD was investigated in [29, 30, 31] for uncertainty
estimation. It combines both MCD and ensemble. For
prediction estimation, MCD evaluates a single feature
representation, but additionally considers the uncertainty
associated with the feature representation. However,
an ensemble evaluates multiple feature representations
without considering the uncertainty associated with
individual feature representations. Hence, MMCD applies
MCD to an ensemble to evaluate multiple feature
representations and consider the uncertainty associated with
individual feature representations. Given a set of CNNs
  for  ∈ 1, 2, ...,  , the MMCD prediction is
obtained by averaging over the predictions of all stochastic
CNNs. That is,
(|, ) ≈
≈</p>
        <p>1 ∑︁ ∑︁ (|,   )
 ·  =1 =1</p>
        <p>1 ∑︁ ∑︁   ().
 ·  =1 =1
(3)
(4)</p>
        <sec id="sec-2-4-1">
          <title>In this work, we can refer to MMCD as an average of</title>
          <p>·  stochastic CNNs.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Combining logits instead of probabilities</title>
      <p>The output layer of a CNN-based classifier includes
 output neurons with a softmax activation function,
 = [︀ 1 . . .  ]︀  as
 =  ()</p>
      <p>1
= ∑︀
=1 exp()</p>
      <p>︀[ exp(1) . . . exp( )︀]  .</p>
      <sec id="sec-3-1">
        <title>From Figure 1, given an ensemble of  deterministic</title>
        <p>CNNs with logits , the average logit  can be estimated
as
 :=</p>
        <p>1 ∑︁  :=

=1</p>
        <p>1 ∑︁   ().

=1
and the predicted probability vector of the ensemble of
deterministic CNNs can be reformulated as</p>
      </sec>
      <sec id="sec-3-2">
        <title>Given MCD representing an ensemble of  stochastic</title>
        <p>CNNs with logits , we can estimate the average logit 
as
 ≈</p>
        <p>1 ∑︁  ≈
 =1</p>
        <p>1 ∑︁   (),
 =1</p>
        <p>(8)
(a) Logit averaging
(b) Probability averaging
which normalizes its inputs (continuous values) to pro- and reformulate the predicted probability vector of MCD,
duce discrete probabilities  (with  = 1, . . . ,  and as shown in (7). Similarly, given MMCD representing an
∑︀=1  = 1) representing the probability that the ensemble of  ·  stochastic CNNs with logits  , we
input image belongs to the class associated with the can estimate the average logit  as
tioℎn oauretpluogtintseaunrodni.nteTrhpereitnedpuats teovitdheencseofftomr apxosfsuibnlce-  ≈ 1 ∑︁ ∑︁  ≈ 1 ∑︁ ∑︁   (), (9)
classes [15]. The discrete probability  is interpreted  ·  =1 =1  ·  =1 =1
as the model confidence that the input belongs to the
class associated with the ℎ output neuron. Given the
logit vector  = [︀ 1 . . .  ]︀  , the softmax estimates
and reformulate the predicted probability vector of
MMCD, as shown in (7). From Figure 2, averaging logits
instead of probabilities of multiple stochastic or
deterministic CNNs increases the confidence of the averaged
CNNs. Intuitively, logit averaging provides the best
evidence (characterized by a low level of uncertainty caused
by the reduction of inductive biases inherent in
individ(5) ual logits) for making decisions. However, probability
averaging provides the best confidence associated with
decisions made using weak evidence (characterized by
a high level of uncertainty caused by inductive biases
inherent in individual logits). This implies that a decision
made using probability averaging considers more
uncer(6) tainty than that made using logit averaging. In this work,
we evaluated the impact of the possible increase in the
degree of confidence caused by applying logit averaging
instead of probability averaging on the QoC.</p>
        <sec id="sec-3-2-1">
          <title>5.1. Experimental setup</title>
          <p>We hypothesized that the QoC of CNNs (strongly)
depends on the task-dificulty (specified using the
training data), the underlying architecture, and/or the
training procedure (mostly influenced by the regularization
(|, ) =  ().
(7)</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Experiments</title>
      <p>Softmax</p>
      <sec id="sec-4-1">
        <title>5.2. Evaluation metrics</title>
        <p>QoC was evaluated by assessing the degree of confidence
calibration. Specifically, we evaluated the calibration
error using measures, such as the negative log likelihood
(NLL) applied in [4, 5, 31], expected calibration error
strength). Therefore, we compared logits and probabili- (ECE) applied in [13, 8, 12], and Brier score (BS) applied
ties averaging on three datasets to evaluate the impact of in [4]. Low values of NLL, ECE, and BS indicate low
calithe task-dificulty on the QoC. Moreover, we compared bration error and vice versa. Furthermore, we evaluated
logits and probabilities averaging using three diferent QoC by assessing its ability to separate TPs and FPs. Here,
architectures to evaluate the impact of the underlying ar- we evaluated the average confidence on evaluation data
chitecture on the QoC. Specifically, we evaluated MNIST causing TPs or FPs. Given evaluation data causing TPs,
[32] on VGGNets [1], FashionMNIST [33] on ResNets [2] we expect the average confidence on the evaluation data
and CIFAR10 [34] on DenseNets [35]. Finally, we com- to be high. However, for the evaluation data causing FPs,
pared logits and probabilities averaging on CNNs trained we expect low average confidence on the evaluation data.
using two regularization strengths (strong and weak reg- Moreover, we evaluated the ability to separate TPs and
ularization summarized in Table 1) to evaluate the impact FPs by evaluating the area under the receiver operator
of the regularization strength on the QoC. We observed characteristic (AU-ROC) applied in [37, 5]. AU-ROC
sumstrong and weak regularization results in underconfident marizes the trade-of between the fraction of TPs that are
and overconfident CNNs, respectively. All CNNs were correctly detected and those of FPs that are undetected
regularized using batch normalization [36] layers placed using diferent thresholds. In summary, in addition to the
before each convolutional activation function. All CNNs NLL, ECE, and BS, we evaluated the accuracy, average
were randomly initialized and trained with random shuf- confidence, and AUC-ROC.
lfing of training samples. All CNNs were trained using
the categorical cross-entropy and stochastic gradient de- 5.3. Evaluation data
scent with momentum of 0.9, learning rate of 0.02, batch
size of 128, and epochs of 100. All images were standard- We used five evaluation data for diferent purposes,
ized and normalized by dividing pixel values by 255. For namely test data, subsets of the correctly classified test
all MCD and MMCD, we sampled activations of the first data, out-of-domain data, swapped data, and noisy data.
fully-connected layer using masks drawn from a cascade
of Bernoulli and Gaussian distributions [22] and using
a dropout probability of 0.5. We performed 100
stochastic forward passes ( = 100) and considered ensembles
consisting of five deterministic CNNs (  = 5).</p>
        <sec id="sec-4-1-1">
          <title>Test data represent the test data from the experimental</title>
          <p>data, namely, MNIST, CIFAR10, and
FashionMNIST. These datasets include both correctly
classified and misclassified test data. Test data are
used for estimating the accuracy, NLL, ECE, and
(a) Test data
(b) Swapped data
(c) Noisy data</p>
          <p>BS. We expect the accuracy to be high and NLL, 5.4. Experimental results
ECE and BS to be low on test data.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>We evaluate the conducted experiments with respect to</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Subsets of the correctly classified test data accuracy and QoC.</title>
        <p>include 1000 correctly classified test data from Table 2 and Table 3 summarize the accuracy,
averthe experimental data. Since CNNs will make TPs age confidence, NLL, ECE, and BS of diferent models
on these data, we used these data for evaluating using the two averaging approaches and CNNs trained
the average confidence on TPs. using strong regularization (causing underconfidence)
and weak regularization (causing overconfidence). The
Swapped data were simulated using subsets of the cor- results show that averaging logits instead of
probabilrectly classified test data structurally perturbed ities do not strongly afect the accuracy. This means
by dividing images into four regions and diago- that averaging logits can preserve accuracy.
Furthernally permuting the regions. From Figure 3b, the more, averaging logits instead of probabilities
signifiupper left and right are permuted with the bottom cantly increases the average confidence. Figure 2
illusright and left regions, respectively. Swapped data trates why the confidence increases. Further, Table 2
include structurally perturbed objects within the shows that averaging logits instead of probabilities
siggiven images. We expect CNNs to make FPs on nificantly decreases the NLL, ECE, and BS for
underswapped data. Therefore, we used these data for condifent CNNs (trained using strong regularization).
evaluating the average confidence on FPs caused This means that averaging logits, unlike averaging
probby structurally perturbed objects. abilities, reduces the calibration error for undercondifent
Noisy data were simulated using subsets of the cor- CNNs. This is because the stronger the regularization,
rectly classified test data perturbed by applying the lower the confidence and the higher the gap between
additive Gaussian noise with a standard devia- accuracy and average confidence. Here, the increase
tion of 500. From Figure 3c, noisy data include in the degree of confidence caused by averaging
lognoise within the given images. We expect CNNs its instead of probabilities reduces the gap between
acto make FPs on these data. Therefore, we used curacy and average confidence. For example, Table 2
these data for evaluating the average confidence shows that averaging logits instead of probabilities of
on FPs caused by noisy objects. the ensemble reduces the gap between accuracy and
average confidence from 18.24(= |88.75 − 70.51|)% to
Out-of-domain data were simulated using 1000 test 9.52(= |88.94 − 79.42|)% on CIFAR10.
data of CIFAR100 [34]. Since CNNs will make However, the increase in the degree of confidence caused
FPs on these data, we used these data for evalu- by averaging logits instead of probabilities increases the
calating the average confidence on FPs caused by ibration error for overconfident CNNs (trained using weak
unknown objects. regularization). Table 3 provides empirical evidence for
this claim by showing that, on CIFAR10 and
FashionMIn general, we expect the average confidence to be high NIST, NLL, ECE, and BS of the ensembles increase when
on TPs and to be low on FPs. the logits are averaged instead of probabilities. We
argued that the more overconfident the CNNs, the higher
the confidence and the higher the gap between accuracy
and average confidence. Here, the increase in the
degree of confidence caused by averaging logits instead of
probabilities further increases the gap between the
accuracy and average confidence and therefore, increases
the calibration error. For example, Table 3 shows that, on
CIFAR10, averaging logits of the ensemble increases the
gap between the accuracy and average confidence from
0.76(= |88.67− 89.43|)% to 7.29(= |88.88− 96.17|)%.</p>
        <p>In Table 4, the average confidence on TPs and FPs is
shown for underconfident models using both averaging
approaches. The results show that averaging logits instead
of probabilities increases the confidence level on TPs and
FPs. The increase in the average confidence is sometimes
very large for FPs due to the noisy data. For example, for
MMCD evaluated on FashionMNIST, the average
confidence on the noisy data increases from 51.31% to 94.58%
when averaging logits. This is because noisy data can
increase the magnitude of logits and averaging logits is
more sensitive to changes in the magnitude of logits than
averaging probabilities (see Figure 2). The increase in the
degree of confidence caused by averaging logits can harm
the separability of TPs and FPs. For example, the increase
in the average confidence on the noisy data from 51.31%
to 94.58% causes the AUC-ROC obtained based on the
evaluation of the degree of confidence to decrease from</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Discussion</title>
      <p>The term ‘combination process’ encompasses how
multiple networks are combined and the information type
combined. It was found in [23, 24, 22] that simple
averaging is more robust and captures uncertainty better than
voting approaches. This is because the simple
averaging equally weights all predictions, while voting ignores
uncertain predictions. In this work, we compared the
process of averaging logits instead of probabilities. We
empirically showed that averaging logits instead of
probabilities increases the confidence while preserving the
accuracy for underconfident or overconfident networks.</p>
      <p>This might be because logit averaging preserves the
position of the max element of individual logit vectors, but
is more sensitive to the magnitude of logit values than
probability averaging. Thus, logit values with a large
magnitude contribute the most to the average logit. In
this way, the magnitude of logit values induces a
nonuniform weighting (for logit averaging), which is lost
(for probability averaging). Furthermore, we provided
empirical evidence showing that for underconfident
networks (trained using strong regularization), the increase
in the confidence caused by averaging logits instead of
probabilities reduces the calibration error on the test data. fidence calibration and the ability of their proposed method
This is because the increase in the degree of confidence to separate TPs and FPs. Finally, for mission- and
safetyreduces the gap between accuracy and average confi- critical applications where the separability of TPs and FPs
dence. However, the increase in confidence caused by is of paramount importance, we suggest to average
probaveraging logits instead of probabilities for overconfident abilities to avoid the negative impact of logits averaging
networks (trained using weak regularization) increases on the ability to separate TPs and FPs.
the calibration error on the test data. This is because the
increase in the confidence further increases the gap
between the accuracy and average confidence. This finding 7. Conclusion
suggests that for underconfident networks, we can
average logits instead of probabilities to reduce the calibration Due to averaging logits instead of averaging
probabilierror. However, we should average probabilities instead ties of stochastic or deterministic networks, the degree
of logits for overconfident networks to avoid increasing of confidence on TPs and FPs increased. This reduces
the calibration error. Although the increase in the confi- the calibration error on the test data for underconfident
dence caused by averaging logits reduces the calibration networks but afects the separability of TPs and FPs. Our
error on the test data for underconfident networks, we empirical results show that there is a trade-of between
empirically showed that it can harm the separability of improving calibration on the test data and improving
TPs and FPs. This is because averaging logits increases the separability of TPs and FPs. Additionally, the
inthe confidence on both TPs and FPs. Therefore, FPs can crease in the degree of confidence increases the
calibraalso be made with high confidence similar to TPs. These tion error on the test data for overconfident networks.
ifndings suggest that reducing the calibration error on Therefore, averaging logits should only be applied when
the test data and improving the separability of TPs and combining underconfident networks. For example, we
FPs can be two contradicting goals. Improving one may can average logits instead of probabilities of an ensemble
be at the detriment of the other. Furthermore, for two of networks trained with mixup or other modern data
augmentation techniques to improve calibration on the
dmooedsenlsotnaencedssa,riiflyseispabreattteerTcPasliabnradteFdPsthbaentter, tthhaenn . test data. Notwithstanding this, for mission- and
safetyThis implies that calibration methods may be insuficient critical applications where the separability of TPs and
for separating TPs and FPs and therefore, ensuring safe FPs is essential, we suggest traditionally average
probdecision-making. Additionally, existing methods for con- abilities. However, it remains unclear if the findings of
ifdence calibration may not help in separating TPs and this paper will change if the given networks or the
average logit are calibrated, for example, with temperature
FPs. Subsequently, future work will evaluate the ability scaling [13]. This suggests a new research direction.
of existing methods for confidence calibration to separate
TPs and FPs. We also recommend researchers to evaluate
both the calibration error of their proposed method for
conating scalable bayesian deep learning methods for
robust computer vision, in: Proceedings of the
IEEE/CVF conference on computer vision and
pattern recognition workshops, 2020, pp. 318–319.
[29] G. Kahn, A. Villaflor, V. Pong, P. Abbeel, S. Levine,</p>
      <p>
        Uncertainty-aware reinforcement learning for
collision avoidance, arXiv preprint arXiv:1702.01182
(
        <xref ref-type="bibr" rid="ref1">2017</xref>
        ).
[30] B. Lütjens, M. Everett, J. P. How, Safe
reinforcement learning with model uncertainty estimates,
in: 2019 International Conference on Robotics and
      </p>
      <p>
        Automation (ICRA), IEEE, 2019, pp. 8662–8668.
[31] A. G. Wilson, P. Izmailov, Bayesian deep learning
and a probabilistic perspective of generalization,
Advances in neural information processing systems
33 (
        <xref ref-type="bibr" rid="ref19">2020</xref>
        ) 4697–4708.
[32] Y. LeCun, L. Bottou, Y. Bengio, P. Hafner,
Gradientbased learning applied to document recognition,
      </p>
      <p>
        Proceedings of the IEEE 86 (1998) 2278–2324.
[33] H. Xiao, K. Rasul, R. Vollgraf, Fashion-mnist:
a novel image dataset for benchmarking
machine learning algorithms, arXiv preprint
arXiv:1708.07747 (
        <xref ref-type="bibr" rid="ref1">2017</xref>
        ).
[34] A. Krizhevsky, G. Hinton, et al., Learning multiple
      </p>
      <p>
        layers of features from tiny images (2009).
[35] G. Huang, Z. Liu, L. Van Der Maaten, K. Q.
Weinberger, Densely connected convolutional networks,
in: Proceedings of the IEEE conference on computer
vision and pattern recognition, 2017, pp. 4700–4708.
[36] S. Iofe, C. Szegedy, Batch normalization:
Accelerating deep network training by reducing internal
covariate shift, in: International conference on
machine learning, PMLR, 2015, pp. 448–456.
[37] D. Hendrycks, K. Gimpel, A baseline for
detecting misclassified and out-of-distribution examples
in neural networks, Proceedings of International
Conference on Learning Representations (
        <xref ref-type="bibr" rid="ref1">2017</xref>
        ).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <year>2017</year>
          , pp.
          <fpage>1321</fpage>
          -
          <lpage>1330</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Dalca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Sabuncu</surname>
          </string-name>
          , Con[1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Very deep convolu- ifdence calibration for convolutional neural net-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          in: International Conference on Learning Represen- arXiv:
          <year>1906</year>
          .
          <volume>09551</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>tations</surname>
          </string-name>
          ,
          <year>2015</year>
          . URL: http://arxiv.org/abs/1409.1556. [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sensoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kandemir</surname>
          </string-name>
          , Evidential deep [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learn- learning to quantify classification uncertainty</article-title>
          , Ad-
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>IEEE conference on computer vision and pattern 31</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>recognition</surname>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bibaut</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. van der Laan</surname>
          </string-name>
          , The relative [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gawlikowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R. N.</given-names>
            <surname>Tassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Lee</surname>
          </string-name>
          <article-title>, performance of ensemble methods with deep con-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Roscher</surname>
          </string-name>
          , et al.,
          <article-title>A survey of uncertainty in deep</article-title>
          <source>Journal of Applied Statistics</source>
          <volume>45</volume>
          (
          <year>2018</year>
          )
          <fpage>2800</fpage>
          -
          <lpage>2818</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>neural networks</article-title>
          ,
          <source>arXiv preprint arXiv:2107</source>
          .03342 [17]
          <string-name>
            <surname>L. I. Kuncheva</surname>
          </string-name>
          , Combining pattern classifiers: meth-
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          (
          <year>2021</year>
          ).
          <article-title>ods and algorithms</article-title>
          , John Wiley &amp; Sons,
          <year>2014</year>
          . [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lakshminarayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pritzel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Blundell</surname>
          </string-name>
          , Sim- [18]
          <string-name>
            <surname>M. Van Erp</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Vuurpijl</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Schomaker</surname>
          </string-name>
          , An overview
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>mation processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ). Workshop on Frontiers in Handwriting Recogni[5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Thulasidasan</surname>
          </string-name>
          , G. Chennupati,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Bilmes</surname>
          </string-name>
          , tion, IEEE,
          <year>2002</year>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Michalak</surname>
          </string-name>
          , On mixup training: [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tajti</surname>
          </string-name>
          ,
          <article-title>New voting functions for neural network</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>deep neural networks</article-title>
          ,
          <source>Advances in Neural Infor- cae</source>
          , volume
          <volume>52</volume>
          ,
          <string-name>
            <surname>Eszterházy Károly</surname>
          </string-name>
          Egyetem Líceum
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>mation Processing Systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ). Kiadó,
          <year>2020</year>
          , pp.
          <fpage>229</fpage>
          -
          <lpage>242</lpage>
          . [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Beutel</surname>
          </string-name>
          , E. Chi, Improving cal- [20]
          <string-name>
            <given-names>T. G.</given-names>
            <surname>Dietterich</surname>
          </string-name>
          ,
          <article-title>Machine-learning research</article-title>
          , AI
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>ibration through the relationship with adversar- magazine 18 (</article-title>
          <year>1997</year>
          )
          <fpage>97</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>ial robustness</article-title>
          , in: A.
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            , [21]
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Tulyakov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Jaeger</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Govindaraju</surname>
          </string-name>
          , D. Doer-
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>ral Information Processing Systems</source>
          ,
          <year>2021</year>
          . URL:
          <article-title>Machine learning in document analysis and recog-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          https://openreview.net/forum?id=NJex-5TZIQa. nition (
          <year>2008</year>
          )
          <fpage>361</fpage>
          -
          <lpage>386</lpage>
          . [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wen</surname>
          </string-name>
          , G. Jerfel,
          <string-name>
            <given-names>R.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. W.</given-names>
            <surname>Dusenberry</surname>
          </string-name>
          , [22]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rovile</surname>
          </string-name>
          , Bayesian convolutional neural
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>your calibration</article-title>
          ,
          <source>arXiv preprint arXiv:2010.09875 on Pattern Recognition and Artificial Intelligence</source>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          (
          <year>2020</year>
          ). Springer,
          <year>2019</year>
          , pp.
          <fpage>118</fpage>
          -
          <lpage>132</lpage>
          . [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rahaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Thiery</surname>
          </string-name>
          , Uncertainty quantifi- [23]
          <string-name>
            <given-names>R. T.</given-names>
            <surname>Clemen</surname>
          </string-name>
          ,
          <article-title>Combining forecasts: A review and</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>Information Processing Systems</source>
          <volume>34</volume>
          (
          <year>2021</year>
          ).
          <source>forecasting 5</source>
          (
          <year>1989</year>
          )
          <fpage>559</fpage>
          -
          <lpage>583</lpage>
          . [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cisse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Dauphin</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
            Lopez-Paz, [24]
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kittler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hatef</surname>
            ,
            <given-names>R. P.</given-names>
          </string-name>
          <string-name>
            <surname>Duin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Matas</surname>
          </string-name>
          , On combin-
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>International Conference on Learning Representa- and machine intelligence</source>
          <volume>20</volume>
          (
          <year>1998</year>
          )
          <fpage>226</fpage>
          -
          <lpage>239</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>tions</surname>
          </string-name>
          ,
          <year>2018</year>
          . [25]
          <string-name>
            <surname>K. C. Lichtendahl</surname>
            Jr,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Grushka-Cockayne</surname>
            ,
            <given-names>R. L.</given-names>
          </string-name>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iofe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wo- Winkler</surname>
          </string-name>
          , Is it better to average probabilities or
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>jna, Rethinking the inception architecture for com- quantiles?</article-title>
          ,
          <source>Management Science</source>
          <volume>59</volume>
          (
          <year>2013</year>
          )
          <fpage>1594</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <article-title>puter vision</article-title>
          ,
          <source>in: Proceedings of the IEEE conference 1611.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>on computer vision</article-title>
          and pattern recognition,
          <year>2016</year>
          , [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <article-title>Dropout as a bayesian ap-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          pp.
          <fpage>2818</fpage>
          -
          <lpage>2826</lpage>
          . proximation: Representing model uncertainty in [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kornblith</surname>
          </string-name>
          , G. Hinton,
          <article-title>When Does La- deep learning</article-title>
          , in: international conference on ma-
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>bel Smoothing</surname>
            <given-names>Help?</given-names>
          </string-name>
          , Curran Associates Inc.,
          <article-title>Red chine learning</article-title>
          ,
          <source>PMLR</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1050</fpage>
          -
          <lpage>1059</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Hook</surname>
            ,
            <given-names>NY</given-names>
          </string-name>
          , USA,
          <year>2019</year>
          . [27]
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Beluch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Genewein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nürnberger</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. M.</surname>
          </string-name>
          [12]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gales</surname>
          </string-name>
          ,
          <article-title>Should ensemble members be Köhler, The power of ensembles for active learn-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>calibrated?</source>
          ,
          <source>arXiv preprint arXiv:2101.05397</source>
          (
          <year>2021</year>
          ).
          <article-title>ing in image classification</article-title>
          , in: Proceedings of the [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Guo</surname>
          </string-name>
          , G. Pleiss,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <source>On IEEE Conference on Computer Vision</source>
          and Pattern
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>calibration of modern neural networks</article-title>
          ,
          <source>in: Inter- Recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>9368</fpage>
          -
          <lpage>9377</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>national Conference on Machine Learning</source>
          , PMLR, [28]
          <string-name>
            <given-names>F. K.</given-names>
            <surname>Gustafsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Danelljan</surname>
          </string-name>
          , T. B.
          <string-name>
            <surname>Schon</surname>
          </string-name>
          , Evalu-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>