<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring the Hyperparameters of XGBoost Through 3D Visualizations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ole-Edvard Ørebaek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marius Geitle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>stfold University Collage</institution>
          ,
          <addr-line>Halden</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Optimizing the hyperparameters is one of the most important and time-consuming activities to do when training machine learning models. But the lack of guidance available to optimization algorithms means that finding values for these hyperparameters is left to black-box methods. Black-box methods can be made more eficient by incorporating an understanding of where good hyperparameter values might be located for a specific model. In this paper, we visualize hyperparameter performance-landscapes in several datasets to discover how the XGBoost algorithm behaves for many combinations of hyperparameter values across these datasets. Using this knowledge, it might be possible to design more eficient search strategies for optimizing the hyperparameters of XGBoost.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Hyperparameters</kwd>
        <kwd>hyperparameter optimization</kwd>
        <kwd>visualizations</kwd>
        <kwd>performance-landscapes</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Hyperparameter optimization is the task of optimizing machine learning algorithms’
performance by tuning the input parameters that influence their training procedure and model
architecture, referred to as hyperparameters. While essential to most machine learning problems,
hyperparameter optimization is a highly non-trivial task as a given algorithm can have many
hyperparameters of diferent datatypes, and efective values for these difer from one dataset to
another [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. This makes it dificult to determine efective values for specific problems and
even more so universally, which in practice results in the use of black-box search algorithms
[
        <xref ref-type="bibr" rid="ref1 ref2">2, 1</xref>
        ]. Black-box algorithms are designed to find solutions without exploitable knowledge
regarding the problem and are often based on principles similar to brute-forcing or random
guessing. While black-box algorithms for hyperparameter optimization are often quite
sophisticated and can be empirically proven to return efective values, they cannot provide much, if
any, insight into what makes these efective compared to others. Obtaining insight into how
diferent hyperparameter values, individually and in combination, impact performance under
diferent circumstances would, however, be extremely useful. With such insights,
hyperparameter optimization methods can be designed to exploit preexisting knowledge in combination
with black-box methods, which is likely more eficient than pure black-box algorithms.
      </p>
      <p>A useful method of obtaining insight into complex problems is to visualize them, as this
allows them to be comprehensively presented. In this paper, we gain insight into the behavior
of the hyperparameters of the XGBoost algorithm by visualizing and comparing landscapes of
prediction performance generated based on hyperparameter combinations.</p>
      <p>The remainder of the paper is structured as follows: In Section 2, we present related work
relevant to the visualization of performance-landscapes. The theory behind the XGBoost
algorithm is outlined in Section 3. In Section 4, we document the methodology used for generating
and visualizing samples of hyperparameter-based performance-landscapes. In Section 5, we
present the findings of comparing the generated landscape-samples, and we discuss these
findings in Section 6. Finally, in Section 7, we conclude the paper and discuss future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Visualizing problem subjects can be an efective method of intuitively obtaining many types of
insight, as demonstrated with many articles within machine learning research. For instance,
Li et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] studied the loss landscapes of artificial neural networks through a proposed
visualization method based on the principle of random directions and filter-wise normalization.
With this method, they provided valuable insight into the nature of artificial neural networks;
specifically, how skipped connections afect the sharpness and flatness of loss landscapes and
why these are necessary when training very deep networks. Smilkov et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presented in
their study Tensorflow Playground, a tool for providing users an intuitive understanding of
neural networks by allowing direct manipulation of the neural networks through visual
representations. Other papers [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ], though not directly focused on visualizations, utilize some as a
method of demonstrating concepts to the reader.
      </p>
      <p>
        Despite their usefulness, research exploring visualizations directly related to
hyperparameters are quite limited. The perhaps most relevant papers are the ones that present tools using
visualizations as a method of aiding hyperparameter tuning/analysis processes [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
However, these papers primarily focus on designing the tools instead of using them practically to
obtain insight into hyperparameters.
      </p>
      <p>
        There are, however, several interesting papers that have investigated performance-landscapes,
though not necessarily in the context of hyperparameters. Performance-landscapes are
relevant to our paper because they can be visually analyzed to obtain insights into
hyperparameters’ efects. Performance landscapes have previously been primarily investigated in the
context of neural network loss. The most prominent example of this is the earlier mentioned paper
by Li et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which yielded valuable insight in this context. This study also inspired further
studies proposing similar methods [
        <xref ref-type="bibr" rid="ref10 ref7">10, 7</xref>
        ]. Of these, Fort and Jastrzebski [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], among other
things, demonstrated similarities in the efects of diferent neural network hyperparameters.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. XGBoost</title>
      <p>
        XGBoost, developed by Chen &amp; Guestrin [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], is a gradient boosting decision tree algorithm
designed for both regression and classification problems. Being state-of-the-art, the algorithm
is also regularly featured in winning solutions of, e.g., Kaggle1 competitions. XGBoost is trained
through minimizing a regularized objective function (Eq. 1) by iteratively adding base learners
  , in the form of decision trees, to an ensemble [
        <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
        ].
      </p>
      <p>
        ( ) = ∑  (  ,  ̂ 
the diference between them. The   that best minimizes the loss between  
penalized to avoid overfitting, as denoted by the regularization term
iteration’s prediction  ̂ 
( −1) is greedily added. Additionally, the complexity of the added   is
and the previous
denote the prediction and the target, and  is a loss function that measures
Much of XGBoost’s regularization and model architecture is defined through
hyperparameters. Some of the most impactful are the learning rate, the number of base learners in the
ensemble, and the base learners’ maximum depth. In this paper we refer to these as learning_rate,
n_estimators, and max_depth. The learning_rate originates from the principle of shrinkage [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
and is a value that scales base learner weights to reduce their individual influence on the
ensemble predictions. n_estimators and max_depth are quite natural regularization parameters,
as they directly influence the architecture of XGBoost’s produced models, and therefore
significantly impact performance. Other hyperparameters include gamma, l1 and l2 regularization,
Ω
.
and various parameters for subsampling and column sampling.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Visualizing Hyperparameter based Performance-landscapes</title>
      <p>The goal with this paper was to investigate how the hyperparameters of XGBoost afect its
prediction performance on diferent datasets and to investigate potential similarities between
these efects. Our motivation was to gain a general insight into XGBoost’s hyperparameters
compared to simply analyzing datasets individually. To accomplish this, we compared 3D
visualizations of performance-landscapes relative to each selected dataset, where each landscape
was generated based on a combination of two hyperparameters.
(1)</p>
      <sec id="sec-4-1">
        <title>4.1. Generating the Visualization Data</title>
        <p>To visualize the hyperparameter performance-landscapes, we selected three of XGBoost’s
hyperparameters and set their respective value-ranges, as tabulated in Table 1. Generated
performancelandscapes were based on the combination of two of these hyperparameters at a time. This
resulted in three diferent performance-landscapes for a given dataset, based on the
combination of learning_rate and n_estimators, learning_rate and max_depth, and n_estimators and
max_depth. For the classification datasets, the used performance metric was accuracy, while
for the regression datasets, Mean Absolute Error (MAE) was used.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Adaptive Zoom</title>
          <p>To reduce the amount of data needed to provide highly detailed visualizations, we developed a
novel algorithm, named Adaptive Zoom, for adaptively generating more data points in regions
with better predictive performance, thereby keeping computation time spent on regions of
low performance minimal. The main idea behind the algorithm was to iteratively "zoom" in
on the region of apparent best performance. "Zooming" here refers to adaptively determining
the region of known best performance, based on pre-generated points, and generating more
points within the determined region. Using this algorithm, we could eficiently obtain highly
detailed landscapes by only generating high numbers of points in these specific regions while
leaving the remaining regions at lower details. .</p>
          <p>The algorithm first identifies the p% best performing hyperparameter configurations in the
landscape. From this set of configurations, the region of best performance, defined by
hyperparameter value-ranges, is then determined by the lowest and highest value per hyperparameter.
Finally, a new landscape-sample can be generated based on the determined best performing
region.</p>
          <p>The Adaptive Zoom algorithm is visually demonstrated in Fig. 1, and its pseudocode is
contained in Fig. 2.</p>
          <p>(a)
(b)</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Interpolation</title>
          <p>To accurately compare the performance landscapes’ characteristics, we needed the same
hyperparameter combinations across all ranges. However, due to using Adaptive Zoom, a varying
amount of landscapes-samples, with varying ranges, were generated for each hyperparameter
 
 
ℎ

← 
← 
← 
← 
end for
return 
end procedure
procedure AdaptiveZoom( , ,

 ∈ ℕ</p>
          <p>)
sorted by performance
 ← points</p>
          <p>← empty array
for all ℎ ∈  ( ) do
 ( ) ←  percentage best performing points
 ← values of hyperparameter ℎ
( )
( )
(, ℎ  )
(</p>
          <p>,   ,  )
 ← generated landscape-sample based on 
scipy.interpolate.griddata2 with the "linear" method, which takes a list of points and a list of
their corresponding values, and returns a grid with a specified resolution of interpolated values.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Landscape Generation</title>
          <p>The landscapes, with performance values obtained through two-fold cross validation, were
generated with standard values of non-investigated hyperparameters at an initial resolution of
20 x 20. Further landscape-samples based on adaptive zoom were generated at an individual
resolution of 50 x 50. The number of samples for each landscape was dynamically determined
by running Adaptive Zoom until the returned hyperparameter ranges were no diferent from
the previous iteration. Finally, all generated landscape-samples of each hyperparameter
combination were merged into singular landscapes through linear interpolation.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Datasets</title>
        <p>To explore efects on performance landscapes, we selected several datasets for both
classification and regression problems, as tabulated in Table 2. The datasets were selected to be varied
in dataset characteristics, such as size, and number of categorical and continuous attributes,
to ensure that generated landscapes would be as varied as possible. This was important to
ensure that obtained findings could be generalized and would represent the general relationship
between the hyperparameters and performance as accurately as possible.</p>
        <p>In terms of preprocessing, all datasets were randomly shufled to ensure that their contained
data was dispersed, categorical features with text values were one-hot encoded, and id-columns
were removed.</p>
        <p>2https://docs.scipy.org/doc/scipy/reference/generated/scipy.interpolate.griddata.html</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Findings</title>
      <p>To obtain insight into the hyperparameters’ efects on performance, we analyzed the
landscapes by looking at their general convexity, the compared dominance of the
hyperparameters’ impact on performance, the locations of optima, the numbers of local optima, and general
observations of similarities between the landscapes. The motivation behind analyzing the
landscapes’ convexity and local optima was to investigate how applicable methods using gradient
descent are to hyperparameter optimization. The dominance of the diferent hyperparameters
was investigated to gain insight into which hyperparameters are more important to optimize
to achieve the best results. Comparisons of each landscape’s optima were investigated to gain
insight into the predictability and consistency of their locations.</p>
      <p>Note that the landscapes for the regression datasets are flipped to ensure visual consistency.
For this reason, the MAE-values are presented as negative.</p>
      <sec id="sec-5-1">
        <title>5.1. Convexity</title>
        <p>We found that the landscapes of the learning_rate and n_estimators combination (Figure 3)
were generally lacking in convexity and were instead quite flat and jagged in shape. There
were, however, some exceptions to this. The landscape of the forestfires dataset had a clear
convex shape in the learning_rate axis, where lower values of learning rates yielded better
3http://archive.ics.uci.edu/ml/datasets/QSAR+biodegradation
4https://archive.ics.uci.edu/ml/datasets/Contraceptive+Method+Choice
5https://archive.ics.uci.edu/ml/datasets/Soybean+(Large)
6https://datahub.io/machine-learning/vehicle
7https://archive.ics.uci.edu/ml/datasets/Breast+Cancer+Wisconsin+(Diagnostic)
8https://archive.ics.uci.edu/ml/machine-learning-databases/wine-quality/
9https://archive.ics.uci.edu/ml/datasets/auto+mpg
10https://archive.ics.uci.edu/ml/datasets/Forest+Fires
11https://archive.ics.uci.edu/ml/machine-learning-databases/housing/
(d)
biodegredation
(e) contraceptive
(d) biodegreda- (e) contraceptive
tion
performance than higher values. The landscapes of biodegredation, wdbc, housing, and
autompg also seemed to have some convexity in this axis, though they were still relatively flat in
overall shape.</p>
        <p>For the landscapes of the learning_rate and max_depth combination (Figure 4) the convexity
was quite varied from one dataset to another. The contraceptive dataset was relatively convex
in the max_depth axis, which was also the case for forestfires and winequality-red. Auto-mpg
and housing were somewhat convex in both the learning_rate and max_depth axis.</p>
        <p>Most of the landscapes of the n_estimators and max_depth combination (Figure 5) were flat
in overall shape. However, contraceptive, winequality-red, and perhaps forestfires, had some
convexity to them.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Dominance</title>
        <p>Based on the landscapes, it seemed that n_estimators had, compared to learning_rate, a
considerably larger impact on performance. However, this only seemed the case for
n_estimatorsvalues lower than approximately 100, which materialized as a "wall" in the landscapes. This
was apparent for all dataset-relative learning_rate and n_estimators combination landscapes
(Figure 3), except for forestfires, which did not seem to contain this wall. For
n_estimatorsvalues over 100, we found that the combination of learning_rate and n_estimators resulted in
quite jagged landscapes, with learning_rate appearing to be the most dominant. We also
observed that jaggedness in the n_estimators axis was not equal for all learning_rate values. Only
the landscapes of the auto-mpg and housing datasets were lacking in visible jaggedness.</p>
        <p>Comparing learning_rate to max_depth (Figure 2), it seems that these hyperparameters tend
to have a nearly equal impact on performance except for with a few datasets, such as forestfires,
contraceptive and winequality-red. For these, max_depth appeared to have a larger efect,
creating a wall similar to those observed in the combinations of learning_rate and n_estimators
(Figure 3). We also found that max_depth tended to stop impacting performance beyond certain
values. These values seemed to never exceed max depth = 15, but were observed for several
datasets as being less.</p>
        <p>For the combination of n_estimators and max_depth (Figure 5), we found that the most
dominant hyperparameter changed from one dataset to another, with there sometimes being
performance walls in the n_estimators axis, sometimes in the max_depth axis, and sometimes
in both. Beyond these performance walls, the two hyperparameters seemed about equally
dominant.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Optima</title>
        <p>For all combinations and datasets except vehicle, the optima were located within the
learning_rate value-range of 0.1 to 1.0. For vehicle, the optima were located around the learning_rate
value of 1.5. We also observed that the optima were always located between the n_estimators
value-range of roughly 100 to 250.</p>
        <p>The optima for each dataset was observed to generally be located around the same
learning_rate value within all relevant hyperparameter combinations (Figure 3 and 4), with
winequality being the only exception to this. For this dataset, the optimum was an entirely diferent
learning_rate value for the combination of learning_rate and n_estimators (Figure 3) compared
to learning_rate and max_depth (Figure 4). Relative to this, n_estimators optima seemed a bit
more variable. Max_depth optima were located within values less than 15.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Local Optima</title>
        <p>In terms of local optima, landscapes based on the combination of learning_rate and n_estimators
(Figure 3) seemed to contain many local optima. These did, however, seem to be predominantly
formed from the influence of learning_rate.</p>
        <p>Local optima were also observed to be plentiful for the combination of learning_rate and
max_depth (Figure 4). Here, it seemed that local optima were about equally based on
learning_rate and max_depth, resulting in much more chaotic landscapes. This was, however, only
the case for max_depth values less than 15 due to the observations outlined in Section 5.2.</p>
        <p>For the combination of n_estimators and max_depth (Figure 5), we observed that local
optima seemed less common. Most of the landscapes were relatively flat and only datasets like
biodegredation, contraceptive, vehicle and winequality-red had anything that could be referred
to as local optima. But even for these, local optima were small in size.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Generally</title>
        <p>For the learning_rate and n_estimators combination (Figure 3), we found that most landscapes
were quite similar in characteristics, such as general shape, convexity, and local optima. There
were, however, some diferences between the landscapes of the classification and regression
datasets. Most notably that regression datasets generally seemed to contain less local optima,
and that forestfires’ landscape shape was completely unique compared to the others.</p>
        <p>The diferent learning_rate and max_depth combination landscapes (Figure 4) also seemed
to have reoccurring characteristics. There seemed to generally be a sharp dip in performance
close to learning_rate = 2.0 and max_depth = 1. The only landscape that was lacking this dip
was the one generated from the vehicle dataset. We also found that biodegreadation, soybean
and wdbc had very similar landscapes for the combination of learning_rate and n_estimators,
which can also be said for winequality-red and forestfires.</p>
        <p>We also observed that several datasets resulted in similar landscapes for all hyperparameter
combinations. The clearest example of this was housing and auto-mpg, which seemed nearly
identical. This was also the case for soybean and wdbc. In terms of landscape characteristics,
the two datasets that stuck out the most were contraceptive and forestfires. For contraceptive,
max_depth seemed to have a much larger influence on performance than for other datasets,
while for forestfires, n_estimators appeared to have no significant performance impact.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>Our goal with this paper was to obtain insights into how and under what conditions the
hyperparameters of XGBoost afect its performance by analyzing and comparing
hyperparameterbased performance landscapes for various datasets. The motivation was to use such insights
to aid the eficiency of hyperparameter optimization strategies for XGBoost.</p>
      <p>The findings showed several indications that analyzing the performance-landscape
visualizations helped gain insight into efective search ranges of the investigated
hyperparameters; learning_rate, n_estimators, and max_depth. For learning_rate, the fact that optima
were within the value-range of 0.1 to 1.0 for all datasets except one, indicates that this range
might generally be optimal when searching for values of this hyperparameter. Similarly, with
n_estimators, the search range of 100 to 250 may be more optimal than other alternatives
under 500 estimators, based on the optima generally being located within this range. We also
discovered that n_estimators values less than 100 formed a "wall" of performance-increase,
most notably in landscapes of the learning_rate and n_estimators combinations (Figure 3). This
could imply that n_estimators less than 100 can reasonably be excluded from hyperparameter
searches. Regarding max_depth, we observed that values greater than 15 for this
hyperparameter seemed not to afect XGBoost’s performance, implying that limiting the search range
from 1 to 15 might be reasonable. These implications are useful for standardizing searches
regarding these hyperparameters, which makes their optimization more eficient by reducing
computation time.</p>
      <p>By analyzing and comparing the general characteristics between the performance-landscapes,
we also found indications that the general behavior of XGBoost’s hyperparameters is
somewhat predictable. Specifically, we found that landscapes based on the same hyperparameter
combination had general similarities across the datasets, and that certain datasets resulted in
almost identical landscapes, optima included, across all combinations. This was observed to
be the case for datasets both similar and dissimilar in general characteristics. Based on this, it
is possible that performance-landscapes based on XGBoost’s hyperparameters can largely be
predicted from aspects of the datasets, though any such specific aspects were not obvious from
the findings. However, if such landscape predictions are possible and reliable, it could mean
that efective search methods and hyperparameter values for XGBoost can potentially be made
somewhat deterministic. Based on this assumption, such predictions would be a prime target
to take advantage of when designing future hyperparameter optimization methods relevant to
XGBoost. Worth noting, however, is that this is still unlikely to completely remove the need
for black-box methods, due to the large amount of local optima and general flatness observed
in the landscapes. Nevertheless, a combination of the two would likely be beneficial.</p>
      <p>
        Compared to earlier studies, our visualization method seemed efective for obtaining
hyperparameter insight despite its simple premise and implementation. For instance, compared
to previously suggested tools using visualizations for hyperparameter tuning processes [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ],
our method does not need to be tied to a specific tuning process but instead focuses on
several datasets to obtain more general insight into the hyperparameters. It is also apparent that
exploring performance landscapes, similarly to earlier papers investigating neural networks
[
        <xref ref-type="bibr" rid="ref10 ref4 ref7">4, 7, 10</xref>
        ], is also useful for investigating hyperparameters of other algorithms/methods.
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion and Future Work</title>
      <p>In this paper, we attempted to gain insight into how XGBoost’s hyperparameters afect
performance by analyzing hyperparameter-based performance-landscape visualizations. The
findings indicate that the method of visualizing performance-landscapes, based on combinations
of two hyperparameters at a time, is a useful tool for gaining insight into, e.g., efective ranges
of XGBoost’s hyperparameters. This was derived from, e.g., how optima were generally
located between 0.1 and 1.0 learning_rate, and 100 and 250 n_estimators. Also, how max_depth
was never observed to have an efect on performance for values greater than 15, and how
n_estimators typically had a wall of performance within the value range of 1 to 100. We also
found indications that the visualization method is efective for gaining insight into the general
and specific behavior of XGBoost’s hyperparameters with diferent datasets. This was derived
from how we observed that most datasets had common landscape-characteristics and how
some dataset-landscapes were nearly identical regardless of similarity/dissimilarity of
datasetcharacteristics. Thus, we conclude that 3D visualizing hyperparameter-based
performancelandscapes is an efective tool for obtaining various types of insight into the behavior of
XGBoost’s hyperparameters. And that these insights can possibly be used to design more eficient
hyperparameter optimization strategies for this algorithm.</p>
      <p>
        While the findings and their implications are promising, several points of investigation
should be explored in future work. For instance, larger datasets should be investigated to
ensure that the findings are generalizable. During the process of generating the visualizations,
we also noticed that the interpolation method would produce artifacts for certain landscapes.
Minor aspects of the visualizations might therefore be inaccurate and should be fixed. The
utilized visualization method is limited in that it can only visualize two hyperparameters at
a time. However, it might be possible to use hyperparameter-based vectors to generate the
visualizations, similar to how Li et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] did with neural network weights, which would solve
this limitation. As discussed in Section 6, we observed that certain datasets, of similar and
dissimilar characteristics, produced strikingly similar landscapes. The cause of this is likely a
worthwhile point of further research.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lücke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt-Thieme</surname>
          </string-name>
          ,
          <article-title>Beyond manual tuning of hyperparameters</article-title>
          ,
          <source>KIKünstliche Intelligenz</source>
          <volume>29</volume>
          (
          <year>2015</year>
          )
          <fpage>329</fpage>
          -
          <lpage>337</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Feurer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <article-title>Hyperparameter optimization</article-title>
          ,
          <source>in: Automated Machine Learning</source>
          , Springer, Cham,
          <year>2019</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bergstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Random search for hyper-parameter optimization</article-title>
          ,
          <source>Journal of machine learning research 13</source>
          (
          <year>2012</year>
          )
          <fpage>281</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , C. Studer, T. Goldstein,
          <article-title>Visualizing the loss landscape of neural nets</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>6389</fpage>
          -
          <lpage>6399</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Smilkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sculley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. B.</given-names>
            <surname>Viégas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wattenberg</surname>
          </string-name>
          ,
          <article-title>Direct-manipulation visualization of deep networks</article-title>
          ,
          <source>arXiv preprint arXiv:1708.03788</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Frosst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Papernot</surname>
          </string-name>
          , G. Hinton,
          <article-title>Analyzing and improving representations with the soft nearest neighbor loss</article-title>
          , arXiv preprint arXiv:
          <year>1902</year>
          .
          <year>01889</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jastrzebski</surname>
          </string-name>
          ,
          <article-title>Large scale structure of neural network loss landscapes</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6706</fpage>
          -
          <lpage>6714</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Convertino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Most</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zajonc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Tsai</surname>
          </string-name>
          , Hypertuner:
          <article-title>Visual analytics for hyperparameter tuning by professionals</article-title>
          ,
          <source>in: Proceedings of the Machine Learning from User Interaction for Visualization and Analytics Workshop</source>
          at IEEE VIS,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Choo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-W.</given-names>
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sung</surname>
          </string-name>
          , Visualhypertuner:
          <article-title>Visual analytics for user-driven hyperparameter tuning of deep neural networks</article-title>
          ,
          <source>in: Demo at SysML Conference</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Scherlis</surname>
          </string-name>
          ,
          <article-title>The goldilocks zone: Towards better understanding of neural network loss landscapes</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>3574</fpage>
          -
          <lpage>3581</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          ,
          <source>in: Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>Greedy function approximation: a gradient boosting machine</article-title>
          ,
          <source>Annals of statistics</source>
          (
          <year>2001</year>
          )
          <fpage>1189</fpage>
          -
          <lpage>1232</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>