<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>İ. Po-
Journal of Physical Chemistry A 126 (2022) lat</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1162/neco_a_01199</article-id>
      <title-group>
        <article-title>Problems in Kinetic Modelling of Catalytic Reactions with Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandr Fedorov</string-name>
          <email>Aleksandr.Fedorov@catalysis.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Linke</string-name>
          <email>David.Linke@catalysis.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leibniz-Institut für Katalyse e. V.</institution>
          ,
          <addr-line>Albert-Einstein-Straße 29a, 18059 Rostock</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>445</volume>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>In the present study we investigated the behaviour of diferent types of neural networks in data extrapolation. The application is modelling a tubular chemical reactor in which a heterogeneous catalytic reaction (CO 2 hydrogenation to methanol) is performed. Since data are slow and expensive to measure we focused on small data sets for training. The diferent models (feed-forward neural network (NN), physics-informed NN, neural ordinary diferential equation (ODE), kinetics-constrained neural ODE) were trained in a way to achieve approximately the same values of loss function in the cross-validation. Although the obtained models have the same generalization ability, the extrapolation capability varies significantly. Wherein, a neural network model that is additionally constrained by the general chemical and chemical engineering knowledge demonstrated much better extrapolation ability compared to unconstrained models. Methods how to validate the generalization of the neural network kinetic models without using additional experimental data were suggested and discussed. neural ODE, physics-informed neural networks, kinetics-constrained neural ODE, kinetic modelling, CO2 hydrogenation timated from the data in data-driven kinetic modelling, the architecture of the models [11, 12, 13]. In our recently Workshop Proceedings Neural networks are ones of popular methods of ma- extrapolation we mean in general the ability to predict Attribution 4.0 International (CC BY 4.0).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The kinetic model development of catalytic reactions is
one of the dificult but important part of chemical
engineering for the process and industrial plant simulation
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The kinetic models of chemical reactions are
ideally based on the knowledge of the reaction mechanisms.
Due to complexity of the reaction mechanisms and the
existence of a lot of parameters which have to be
esseveral assumptions (postulating rate-determining stages,
quasi-equilibrium approximation etc.) are usually used
to convert the mechanisms to a set of reaction rate
expressions for describing the dynamic of the process [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
The complexity of this traditional approach is related
tion mechanism as well as the dificulty of performing
to the necessity of having deep knowledge of the reac- of CO
the screening of diferent possible assumptions. Notably,
solving the inverse kinetic task (the estimation of kinetic
model parameters from the data) is still a nontrivial task
in the modelling process [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For these reasons, a
machine learning approach is an attractive alternative for
kinetic model development because only the data are
needed to develop the models.
chine learning due to their flexibility which led to the
ITAT’24: Information technologies - Applications and Theory,
SeptemnEvelop-O
(D. Linke)
ber 20–24, 2024, Drienica, Slovakia
∗Corresponding author.
†These authors contributed equally.
as for testing the model ability to extrapolate. The
followCEUR
      </p>
      <p>
        ceur-ws.org
neural networks [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], physics-informed neural networks
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], neural ordinary diferential equations (neural ODEs)
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], kinetics-constrained neural ODE [13]. To compare
the neural networks based modelling approach with the
traditional one, we also developed a simple power-law
kinetic model assuming first order with respect to the
reagents in the reaction rate expressions.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Experimental part</title>
      <sec id="sec-2-1">
        <title>2.1. Data generation</title>
        <p>For the data generation, the kinetic model of 
drogenation to methanol was used from the work [14].
This kinetic model was based on the following chemical
reactions:
 2 + 3 2 ⇆ 
3 + 
2
,
 2 +  2 ⇆  + 
2
,
 2 + 4 2 ⇆</p>
        <p>4 + 2 2.</p>
        <p>A dual-site Langmuir-Hinshelwood kinetic model was
used for describing the dynamics of the suggested
reactions:
log10  ,1</p>
        <p>= −
log10  ,2
=</p>
        <p>3066
+ 2.029,
2 hy- ergy of the reaction constant  ; R - universal gas constant;
temperature  0 that was set to 573.15 K;   - activation
enT - temperature. The adsorption equilibrium constants
were presented in the form of Van’t Hof equation:
reaction  ;  ,
 ;   2
where   is the rate of the chemical reaction  ;   is the
partial pressure of compound  ;   is the rate constant of the
- the equilibrium constant of the reaction
- the equilibrium constant of 
2 adsorption;   2
- the equilibrium constant of dissociative adsorption of
 2. It is worth mentioning that the thermodynamic part
of the reaction 3 was removed from the equation in the
present study because the equilibrium constant of this
reaction is very high (Δ  ∘ &lt; 0 and |Δ  ∘|/ ≈ 19.6 ≫ 1 ,
 ,3
= exp (−Δ  ∘/</p>
        <p>)≈ 3.3 ⋅ 108). For comparison, the
values of equilibrium constants of the reactions 1 and 2
are 2.6 ⋅ 10−2 and 5.7 ⋅ 10−6, respectively. The values are
given for 573.15 K. The following equations were used to
estimate the values of the equilibrium constants of the
reactions 1 and 2:
(7)
(8)
(9)
(10)
(11)
,
(6) (at diferent residence time).
ating the training data. For each reaction conditions, 5
points of the residence time (0.05, 0.1, 0.2, 0.4, 1.0) were
used for the data generation. Thus, the total number of
diferent reaction conditions was 13 ⋅5 = 65. The total
inlet molar flow was set to a constant value of 1 for each
reaction conditions (only CO
2 and H
2 were used in the
inlet flow). Thus, the dataset represent the dependencies
between the outlet molar flows of each compound
(obtained by the integration of the system of ODE 11 ) and
the reaction condition (initial molar flows of 
temperature, total pressure, and the residence time). The
partial pressure of compounds (required for the
estimation the reaction rates 4, 5, and 6.) was estimated by the
2 and  2,
following equation assuming the ideal gas law:
(1)
(2)
(3) from the original work [14].</p>
        <p>( ) =  ,0 exp (
Δ  (

1
 0
1

−
)),
where  ,0 is the adsorption equilibrium constant at the
reference temperature  0 and Δ  is the molar change of
the enthalpy. The parameters of the model were taken</p>
        <p>To generate the data for the present study, the
following system of ODE was integrated:

  = ∑ 

=1
   ,
 1 =
 2 =
 3 =
 1 (  2  2
 3
−   3
   2 )
 ,1
  22 (1 +   2 ⋅   2)(1 + √  2 ⋅   2)
 2 (  2  2
−     2 )</p>
        <p>,2
√  2 (1 +   2 ⋅   2)(1 + √  2 ⋅   2)</p>
        <p>3√  2  2
(1 +   2 ⋅   2)(1 + √  2 ⋅   2)
,
(4) idence time,  - is the number of the chemical reactions;
where   is the molar flow of the compound  ;  is the
res, (5)

 is the stoichiometric coeficient of the compound
the reaction  (positive for the products and negative for
 in
reagents). The system of ODE represents the
mathematical model of 1D plug flow reactor for catalytic reaction
assuming the absence of heat and mass transfer
limitations. The integration of the system 11 allows one to
estimate the molar flows of compounds along the reactor
,
compound  ;   - the molar flow of compound  .
where  is the total pressure;   - the partial pressure of
The 
2 conversion   2
and selectivity of 
compound   , which are the commonly used measures of
reactor and catalyst performance, were estimated by the
following equations that relate the state at reactor inlet
to the state at reactor outlet:
  2 =
  =



 2</p>
        <p>−   2

 2




.</p>
        <p>,

 2</p>
        <p>−   2</p>
        <p>To validate and investigate the extrapolation ability
of the kinetic models, 4 diferent test datasets were
generated. The 4 test datasets difered in the range of the
residence time selection (0-1, 1-2, 2-5, and 5-10) to
investigate the extrapolation ability of the investigated kinetic
models. Each dataset had 1000 points which were
generated by randomly selecting values from the ranges of the
reaction conditions (temperature in 200-400 ∘C, pressure
in 20-40 bar, H2:CO</p>
        <p>2 ratio in 1.5-6.0) using a uniform
distribution. To train physics-informed neural network
model, an additional dataset (labeled as the zero set) was
created. The zero set was generated by a similar
procedure applied for the test data generation except for
setting the residence time to be 0. The zero dataset had
also 1000 points.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Model architecture and training</title>
        <p>Diferent variants of artificial neural networks were used
to fit the data to CO 2 hydrogenation. It is important to
The reaction conditions used for the training data generation. (the number of layers and neurons, a type of activation</p>
        <sec id="sec-2-2-1">
          <title>Temperature, ∘C</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>Pressure, bar</title>
          <p>H2:CO2 ratio
  = exp [
CH4, and CH3OH. The NN model has 1 hidden layer with
20 nodes and hyperbolic tangent activation function as
well as the output layer with sigmoid activation function.
To train the NN model, the following loss function was
used:</p>
          <p>( 

−</p>
          <p>2
) ,

(17)

 
(18)
is
note that the architecture and parameters of NN models
function etc.) were chosen to achieve approximately the
same value of loss function (Equation 17 ) for 10-fold
cross-validation during the model training.</p>
          <p>The first model was the simple feed-forward neural
network denoted as NN. This model has 4 inputs
(temperature, pressure, the inlet fraction of CO2 and the residence
time). Temperature was transformed in the following
form and used as an input for neural network models:
 
  
=</p>
          <p>∑</p>
          <p>⋅   2

 
where</p>
          <p>is the  -outlet of PINN model; 
the  -outlet of NN model. For the training PINN model,
we used the knowledge that there is no change in the
molar flows of compounds when the residence time is to
equal to 0. For this, an additional zero dataset was also
used for training of PINN model and the corresponding
loss function was:

2 = 
( 


) .</p>
          <p>
            The next model (NODE) was based on the neural ODE
[
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] that uses neural networks for approximating the right
models.
part of ODEs:
conventional kinetic model that is based on the
assumption that first order kinetics can suficiently describe the
(19) reactions. The PL model was selected as baseline model
and used for the comparison with the NN models. The
loss function 17 was used for training KCNODE and PL

  = ANN(  ,  ),
where ANN is the feed-forwards neural network. The
inputs of ANN were the partial pressure of compounds
and the re-parameterized temperature (Equation 15). The
NODE model has 7 inputs (partial pressure of compounds
- CO2, CO, CH4, CH3OH, H2, H2O as well as the
reparameterized temperature (Equation 15)). The NODE
model consists of two hidden layers with hyperbolic
tangent and exponential activation functions. The number
of nodes was 3 for both layers. The output of the ANN
was the linear function. The loss function 17 was used
for training NODE model.
          </p>
          <p>The kinetics-constrained neural ODE approach was
used for developing the fourth model (KCNODE). The
idea of the approach is to use the general knowledge
about the process and approximate the rates of chemical
reactions by the following equation:
(20)</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Hardware and software specifications</title>
        <p>Dormand-Prince-Shampine method (DOPRI5) method
was used for the integration of the neural ODE [16].</p>
        <p>Training neural network models was carried out by
minimizing the loss functions using ADAM [17] optimizer
with a learning rate of 0.005. L2-regularization was used.</p>
        <p>The parameter of the regularization was set to 10−6. All
calculation were implemented in the Python
programming language (version 3.9.12) [18]. The scientific
libraries NumPy [19] (version 1.23.0), SciPy [20] (version
1.8.1), Pandas [21] (version 1.23.0), Scikit-learn [22]
(version 1.1.1) were used for data analysis and evaluation.</p>
        <p>Pytorch [23] (version 1.12.0) and Torchdyn [24] (version
1.0.3) were used for building and training neural
networks models. Matplotlib [25] (version 3.7.2) was used
to visualize the results.</p>
        <p>3. Results and Discussion
  =   ( ) ⋅ ∏ (   ) ⋅ (1 −</p>
        <p>,</p>
        <p>) ⋅ ANN(  ,  ),
where   ( ) is the rate constant of the reaction  defined
by Equation 9;    is the partial pressure of the reagent 
in the reaction  .  - the reaction quotient. For example,
in the case of reaction 1 the corresponding quotient is:
 1 =
  3</p>
        <p>2 .
  2  2</p>
        <p>3
(backward reaction).</p>
        <p>The KCNODE model consists of 1 hidden layer with
3 nodes (hyperbolic tangent activation function) and the
outlet layer which has sigmoid activation function. The
sigmoid function is chosen for the output layer because
it ensures positive value, and, thus, Equation 21 aligns
with thermodynamic. So, if ( &lt;   ), the rate is positive
(forward reaction), and if ( &gt;   ), the rate is negative</p>
        <p>It is worth noting that the activation energies and the
rate constants were also parameters of the KCNODE
model and were varied during the training of the neural
network model, along with the weights and biases of the
neural network layers.
places the ANN part in the KCNODE model by simple
power law rate expressions. Thus, it represents a simple</p>
        <p>The first step of the present study was to obtain the
(21) trained neural network models with the similar
gener(22) the dependency between the value of 10-folds CV of the
alization ability. In our work, we have chosen the value
of the loss function in the 10-folds cross validation (CV)
as a metric of the generalization. Firstly, we estimated
the CV value for the PL model which was 1.7 ⋅10−3. To
achieve a similar CV value of the loss function for all
neural network models, we trained our models by
minimizing the corresponding loss function. Figure 1 shows
loss function and the number of epochs. It can be seen
that the CV loss function decreases with increasing
number of epochs for all the neural network models except
for the NODE one where we observe a significant rise in
the value of the loss function after around 2000 epochs.</p>
        <p>From the obtained data we have found the number of
iterations which is needed for training to achieve the
corresponding value of loss function (1.7 ⋅10−3 in our case).</p>
        <p>It is worth noting that we did not manage to achieve
this target value for the NODE model. For this model,
the minimum value of the loss function (2.1 ⋅10−3) was
therefore chosen which is still close to the targeted value.</p>
        <p>Thus, we managed to obtain a set of diferent neural
network models with similar generalization ability.</p>
        <p>To validate the generalization ability of machine
learn</p>
        <p>However, for kinetic model development, chemical and
The last model denoted as Power-Law (PL) model re- ing models, an additional test dataset is typically used.
models describe the solution of the system of ODE but
not the kinetic model represented in the form of rate
equations. This imposes serious restrictions on using
the resulting models for up-scaling (modelling another
type of reactor or extending the models by adding the
difusion/heat transfer).</p>
        <p>The neural ODE approach does not have such a
limitation since the neural ODE models represent the
approximation of the right part of ODE. From Table 2 one
can see that better extrapolation ability is observed for
neural ODE models (NODE and KCNODE) compared
to NN and PINN ones. Wherein, the loss function for
the dataset generated with a residence time range of
510 was 0.66 for KCNODE model and around 5 times
lower then the one for NODE model (3.4). Thus, the
neural ODE model additionally constrained by the general
Figure 1: The dependencies between the value of 10-fold chemical and chemical engineering knowledge
demonCV and the number of epochs during training for diferent strates much better extrapolation ability. To compare
models. In the case of neural ODE-based models (NODE and the KCNODE model with a traditional approach,
TaKCNODE), the number of epochs was multiplied with 10 for ble 2 shows the values of loss function for PL model.
tboetetsetrimvisautealtihzaetisotnan.dTahredtrdaeivniiantgiopnrotcheastsiwsashsorewpneaatseder5rotrimbeasr One can see that both KCNODE and PL models show
in the figure. similar predicting ability in fitting the test data. In
addition, another kinetics-constrained neural ODE model
(KCNODE*) was obtained by training after 10000 epochs
to achieve the minimum value of the CV loss function to
engineering knowledge can be utilized for validation. compare with other models. The values of loss function
We simulated virtual data using the obtained models and for diferent test data are also presented in Table 2. On
compared them in Figure2. The analysis revealed predic- can see, that the the loss function for the dataset
genertion issues with the NN and NODE models, particularly ated with a range of residence time in 5-10 is around 0.2
with the carbon balance and molar flow predictions. The and decreased compared to KCNODE model.
NN model predicted non-zero CH3OH flow at zero
residence time, and the NODE model predicted negative
molar flows. These issues are illustrated by plotting CO 2 4. Conclusions
conversion and product selectivity, which should range
from 0 to 1 but did not for NN and NODE models. In In the present work, we investigated the behaviour of
difcontrast, knowledge-integrating models (PINN and KC- ferent neural network models (feed-forward NN,
physicsNODE) did not exhibit these problems. informed NN, neural ODE, kinetics-constrained neural</p>
        <p>To assess the approximation ability of the obtained ODE) when applied for modelling catalytic data
describmodels, the series of the diferent test datasets was gen- ing the process of CO2 hydrogenation to methanol. To
erated. The datasets difered in the residence time range compare models under the same conditions, we trained
(from 0-1 to 5-10). The values of estimated loss function the models in a way to achieve a similar value of the
for the test datasets are presented in Table 2. One can loss function in the 10-folds cross validation. Although
see that the values of the loss function for the dataset the obtained models had similar generalization
capabilwith the residence time range of 0-1 are similar for all the ity, we showed that only neural ODE model additionally
models. However, when we try to predict data outside of constrained by the general chemical and chemical
engithe training dataset range, we observe that the values of neering knowledge demonstrate a good fitting of the data
the loss function increase with increasing residence time in the context of the residence time. Moreover, its
caparange for all the models. This increase varies significantly bility to extrapolate was comparable to the traditional
from model to model. The highest errors in the data pre- modelling approach. Due to the extrapolation ability, the
diction for a residence time range of 5-10 were found for application of neural ODE models for accelerated kinetic
the NN and PINN models. This is due to the fact that the model development can be expected to grow.
models were trained on the dataset where the residence
time varied in the range of 0-1. Another limitation of the
NN and PINN models is that both models only represent
the solution of the reactor model. It means that these</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Acknowledgments</title>
      <p>Financial support from German Research Foundation
(DFG) through the project NFDI4Cat (DFG no. 441926934)
is gratefully acknowledged.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yablonskii</surname>
            ,
            <given-names>G. S.</given-names>
          </string-name>
          <string-name>
            <surname>Yablonskii</surname>
            ,
            <given-names>V. I.</given-names>
          </string-name>
          <string-name>
            <surname>Bykov</surname>
            ,
            <given-names>V. I.</given-names>
          </string-name>
          <string-name>
            <surname>Elokhin</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          <string-name>
            <surname>Gorban</surname>
          </string-name>
          ,
          <article-title>Kinetic models of catalytic reactions</article-title>
          ,
          <source>Comprehensive Chemical Kinetics, Elsevier Science &amp; Technology</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Brübach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hodonj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pfeifer</surname>
          </string-name>
          ,
          <article-title>Kinetic analysis of co2 hydrogenation to long-chain hydrocarbons on a supported iron catalyst</article-title>
          ,
          <source>Industrial &amp; Engineering Chemistry Research</source>
          <volume>61</volume>
          (
          <year>2022</year>
          )
          <fpage>1644</fpage>
          -
          <lpage>1654</lpage>
          . doi:
          <volume>10</volume>
          . 1021/acs.iecr.1c04018.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Matera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. F.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Heyden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Savara</surname>
          </string-name>
          ,
          <article-title>Progress in accurate chemical kinetic modeling, simulations, and parameter estimation for heterogeneous catalysis</article-title>
          ,
          <source>ACS Catalysis 9</source>
          (
          <year>2019</year>
          )
          <fpage>6624</fpage>
          -
          <lpage>6647</lpage>
          . doi:
          <volume>10</volume>
          .1021/acscatal.9b01234.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Bebis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Georgiopoulos</surname>
          </string-name>
          ,
          <article-title>Feed-forward neural networks</article-title>
          ,
          <source>IEEE Potentials 13</source>
          (
          <year>1994</year>
          )
          <fpage>27</fpage>
          -
          <lpage>31</lpage>
          . doi:
          <volume>10</volume>
          . 1109/45.329294.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Raissi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Perdikaris</surname>
          </string-name>
          , G. Karniadakis,
          <article-title>Physicsinformed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial diferential equations</article-title>
          ,
          <source>Journal of Computational Physics</source>
          <volume>378</volume>
          (
          <year>2019</year>
          )
          <fpage>686</fpage>
          -
          <lpage>707</lpage>
          . doi:https://doi.org/10.1016/j. jcp.
          <year>2018</year>
          .
          <volume>10</volume>
          .045.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. T. Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rubanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bettencourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Duvenaud</surname>
          </string-name>
          ,
          <article-title>Neural ordinary diferential equations</article-title>
          , in: S. Bengio,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Grauman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cesa-Bianchi</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>31</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2018</year>
          . URL: https:
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>