<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Coupling quantum classification and quantum distance estimation in continual learning⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Corrado Loglisci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Donato Malerba</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Universita' degli studi di Bari</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Quantum machine learning recently gained prominence due to the promise of quantum computers in solving machine learning problems that are intractable on a classical computer. Nevertheless, several studies on problems which remain challenging for classical computing algorithms are emerging. One of these is classifying continuously incoming data instances according to the continual learning paradigm, which is studied in this paper through a hybrid computational solution that combines classical and quantum techniques. Hybrid approaches represents one of the current ways for the use of quantum computation in practical applications. In this paper, we show how typical issues of continual learning can be equally addressed with the properties of quantum mechanics, until to ofer often better results. We propose the combined use of quantum classification and quantum distance estimation to update the classification capabilities as new data instances are processed. Experiments are performed on real-world datasets with quantum simulators.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Quantum machine learning has been introduced with the promise to handle machine learning
problems that are intractable on a classical computer, especially those characterized by huge
amounts of data. In the research on Quantum computing technologies, the current status sees
the era of noisy intermediate scale quantum (NISQ) computers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which are devices able to deal
with low-middle size data problems. An approach which seems bringing practical advantages is
instead the one of hybrid frameworks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that combine classical and quantum methods and
allow to exploit quantum physics properties while limiting the impact of the existing restrictions
of the quantum devices.
      </p>
      <p>
        One of the categories of data-intensive problems in which the research on classical computing
dedicates still many eforts is learning of models from continuously incoming sequential data.
Even the accurate solutions of Deep Learning find challenging working on that data scenario.
This is demonstrated by the long list of recentest studies addressing the so-called catastrophic
forgetting [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which is the tendency of an artificial neural network to abruptly and drastically
forget previously learned information upon learning new information. In those cases, it is not
important designing algorithms for massive computation, but keeping the quality of the models
high over unbounded sequences of data.
      </p>
      <p>
        We investigate these points through a quantum-classical framework which builds a
classiifcation model in the supervised setting and works in continual learning [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ] by acquiring
XXX
continuously incoming data instances. The framework adapts continually a classification model
and keeps on learning over time. More precisely, it trains and updates a classifier on (sub-)
sequences of incoming data instances (data blocks) marked as labelled. Then, the classifier is
used to estimate the class-value of unlabelled incoming data instances. The update is performed
only when the properties and distribution of the labelled data changes, which is what typically
happens in continual learning. To detect such changes, we rely on a quantum centroid distance
estimation technique, often used in quantum clustering. The centroids are considered by their
characteristic of synthesizing clusters, which in this work are produced for each class-value on
the labelled data. So, the changes are detected when the properties of the clusters change and
this happens when the labelled data (assigned to the respective clusters) change. Consequently,
the classification model needs to be updated on those data.
      </p>
      <p>
        The framework has been tested on the binary classification task by using two real-world
datasets. The data size of these experiments is of the same order of the magnitude, or even
higher, of the one used in the related works [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. It has been also compared against a classical
computing algorithm working in continual setting. The experimental results are encouraging
and show the potential superiority in terms of accurate estimations over diferent experimental
configurations.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Quantum-classical framework for binary classification</title>
      <p>The overall framework (illustrated in Figure 1) relies on the classical computing techniques of
feature selection, data sampling, normalization and model optimization. On the quantum side,
it integrates quantum encoding techniques, quantum neural networks in the form of variational
quantum circuits for the problem of binary classification and quantum distance estimation.The
framework faces a binary classification problem that can be formulated as follows. We have data
instances described by X ∪  , X are the descriptive attributes/features, while  ∈ {−1, 1} denotes
the class label. By operating in continual learning, the framework alternates training sessions,
where we have labelled data blocks, with prediction sessions, where the data instances have
no class label. The succession of training sessions and prediction sessions is not predefined,
coherently with the realistic assumption according to which the distribution of labelled and
unlabelled data instances is not previously established and therefore not all the data instances
are labelled.</p>
      <p>In the following, we first provide a short description of each component and then describe
how these work in the whole framework.</p>
      <p>Feature selection operates only at the beginning and selects the subset of descriptive features
which we will consider afterwards. It exploits a classical computing technique based on the
mutual information between the class labels. This operation has been used to perceive the most
relevant characteristics and alleviate the problem of choosing the dimension of the quantum
circuits in terms of qubits. In fact, the number of the features of the input data determines the
number of qubits.</p>
      <p>
        Normalization scales the values of the previously selected features within the range of
[
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] by using the standard min-max function on the original ranges. It is performed for each
incoming data block, both those of training and those of prediction.
      </p>
      <p>Data sampling selects a subset of the labelled data instances within the previously data
block of training session. The samples will contain data instances of both the class labels and,
for each class label, the component takes data instances with simple random techniques without
replacement. The sample size is fixed.</p>
      <p>
        Quantum Centroid Distance Estimation operates only on the labelled data instances and
allows us to detect drift within data. Distance estimation is a standard operation in the realm
of the distance-based clustering and is used to build clusters and determine centroids, but, in
this work, there is no clustering procedure as it is typically defined. Indeed, we determine two
centroids as synthesis of the properties of the data instances of the two class labels, and use
distance estimation to compute the distance between the centroids and labelled data instances.
The centroids represent class prototypes and are used to identify the data instances underlying
the drift. When the drift occurs, the two centroids are re-determined. To implement these
operations we resort to the notion of fidelity
between quantum states [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and, inspired to
the algorithmic decisions of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], design the computation of the distances between single data
instances and the two centroids in superposition. The estimation of the fidelity can be performed
through the SWAP test implemented in the following circuit which presents three groups of
qubits, the first one for the single data instance, the second one for the centroids and the third
one as ancillary. We described it by tensor-product terms:
where,  3 stands for the unitary matrix
Γ( ) =  3( )|0⟩ Γ(  ) =  3(  )|0⟩
 3(, , ) =
      </p>
      <p>2
cos ( )</p>
      <p>−  sin ( )
(

 sin ( )  (+) cos (2 ))
2
2
(1)
(2)</p>
      <p>X stands for the data instance represented through Γ() with the angle encoding and   is the
j-th centroid out of  centroids (k=2).</p>
      <p>So, the state preparation is defined as:</p>
      <p>2
∏ Γ(  ) ⊗ Γ( ) ⊗  ⊗
=1
while the SWAP test is implemented with the following term:
( ⊗ ⊗  ⊗2 )   ,Γ(  ),Γ( ) ( ⊗ ⊗  ⊗2 )
(3)
(4)
(5)
(6)
(7)
(8)
where,  is the Fredkin gate and   is the j-th ancillary qubit which works as control qubit
for the  .</p>
      <p>Quantum Classification works in two modalities, training and prediction. It is implemented
through two quantum circuits with a number of qubits determined by the number features
selected by Feature selection. During the training sessions, the classifier learns on the data block
currently built with labelled data instances. The classifier is instead used during the prediction
sessions to infer the class labels on the data block currently built with unlabelled data instances.</p>
      <p>The first quantum circuit takes the classical data and represents them as quantum states to be
assigned to the qubits. This circuit implements a feature mapping operation ℱ which encodes a
real-valued data instance  into quantum states spanning  qubits:</p>
      <p>where, |0⟩⊗ denotes the register with d-qubits at the state |0⟩ (|0⟩ ⊗ … ⊗ |0⟩). In this work, ℱ
has been implemented as follows
| ( X)⟩ = ℱ (X)|0⟩⊗</p>
      <p>⊗ (X) ⊗ |0⟩⊗
where, the parameter for each gate   is the normalized real-valued of of the feature
(corresponding to the qubit on which   works). The term  ⊗ denotes the tensor product H ⊗ … ⊗H
over  occurrences (that is, the number of selected features) of the gate H (the same holds for
  ).</p>
      <p>
        The second circuit is variational and manipulates the quantum states returned by the first
circuit. It implements a quantum neural network composed of layers of entangled rotation gates.
Entangled rotation gates are matrix operations which combine the gates Hadamard, CNOT
and Rotation under the quantum physics efect of the entanglement [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The second circuit
completes the structure of gates which builds the classifier 
|( X, )⟩ =  ()| (
      </p>
      <p>X)⟩
  ⊗</p>
      <p>⊗   ⊗
where,  is the variational circuit,  denotes the parameters of the parameterized gates that
being optimized. In this work,  has been implemented as follows</p>
      <p>where, each occurrence of the two-qubit gate CX takes one pair of qubits (over the d-qubit
register) composed by the consecutive qubits indexed as  and  + 1 .</p>
      <p>
        Finally, we perform measurements on the qubits and the measured state is recorded. So, we
can estimate the expectation value of the circuit on  and  , by measuring the state over multiple
runs, with the following
|ℰ (X, ) = ⟨( X, )|  ⊗ |( X, )⟩
(9)
where,   ⊗ is the tensor product of the single qubit gate   over  occurrences. The gate  
has the interesting property that if the measured quantum state has odd parity, it returns -1 (as
eigenvalue), while, if the measured quantum state has even parity, it returns 1. This implies that
the expectation value of the circuit will always be within the interval [
        <xref ref-type="bibr" rid="ref1">−1, 1</xref>
        ]. We can use this
property to relate the expectation value to the probability that a data instance X being assigned
to a class label  , that is:
      </p>
      <p>ℰ ( X, ) + 1
 ( | X) = (10)
2</p>
      <p>The probability P(y|X ) is exploited in the optimization process concerning the parameters  .
In particular, the optimizer iteratively updates the circuit parameters by minimizing a cost
function, which accounts for the negative log-likelihood of the probabilities P(y|X ) computed
on the current labelled data-blocks, that is:</p>
      <p>1 
−  ∑=1 ( (  |X )) (11)
where,  is the number of data instances of the data block, X is the i-th data instance of
the data block.</p>
      <p>
        The cost function is minimized by a classical computing optimizer based on gradient descent.
The derivative concerns the expectation value ℰ () with respect to the current values of  and is
computed by means of the parameter shift rule [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]:
ℰ = ℰ ( + ) − ℰ (  − ) (12)
 2
      </p>
      <p>The gradient value is the diference between the two output values of the circuit: the first
value is the output of the circuit with the parameter   increased by a value  , and the second
value is the parameter   decreased by  .</p>
      <p>
        Continual learning. Learning classification models on continuously incoming data can
be faced with time-windows models [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] in continual setting. Time-windows models allow
us to handle data instances by equally-sized blocks on which we train, update and apply the
classification model. During a training session, the training modality of Quantum classification
component is activated (Figure 1), which implies the execution the feature mapping ℱ on the
data instances of the current data block and optimization process of the parameters  of the
variational circuit  . Diferently, during a prediction session, the Quantum classification only
estimates the class labels on the current data block by using the classification model up there
updated.
      </p>
      <p>
        To keep the classifier updated, we have to deal with the catastrophic forgetting efect raising
when updating neural networks. By the way, one of the purposes of this work is investigating
whether this is what happens also on the quantum neural networks. In the literature, three
alternatives are mainly suggested, replay methods, regularization-based methods, parameter
isolation methods [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Considering that the replay methods represent the solution which asks
for less and leaves unchanged the number of hyper-parameters of the neural network, we lean
for this approach.
      </p>
      <p>The framework operates in three steps, namely initialization, update, prediction. Training
sessions are performed at the initialization and update. At the initialization step, the
classiifer is trained from scratch on the first data block

1 (Figure 1). The operation of Feature
selection is used only at the initialization step, so the other steps of the framework work on
the features before selected. Still at the initialization, two centroids, one for each class label,
are determined from the labelled data instances of the data block 
1. As new labelled data
blocks will be acquired, the centroids will be re-computed. However, the computation of the
centroids (Quantum centroid distance estimation) relies on the data representation based on
angle encoding to two dimensions which use two features that have not been selected for the
Quantum classification. Specifically, the feature value is added to 1 and multiplied by half
be an admissible value for the gate  3. This is a data encoding diferent from the feature mapping
used for the Quantum classification. The rationale behind is to use a diferent representational
space in order to capture a diferent characteristics of the data from those expressed by the
 , to
feature mapping of the Quantum classification.</p>
      <p>
        Next, the framework prepares the steps of update and prediction by collecting labelled data
instances in a data block 
 and unlabelled data instances in 
 . Both data blocks store the
data instances in the order they arrive. As soon as one of the two data blocks is being filled
(the number of collected data instances is equal to the predefined size), either updated step or
prediction step is performed. By supposing the data block 
 of labelled data instances has
been filled for first, the update step will be performed, otherwise it will be the turn of prediction
step working on 

. When the update step starts, it first checks for possible concept drifts
within the current data block, and, if any is present, it updates the classification model. To
check the presence of drifts, we rely on a classical computing technique, that is, Page-Hinkley
test [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which, in this work, detects the changes of a cluster that new data instances added
can imply. As indicator of the characteristics of a cluster, we use the sum of the squared errors
  ∶

∑
      </p>
      <p>(
ℎ∈ 
 , 
ℎ
)
(13)
where,</p>
      <p>is the current data block,   refers to one of the two centroids, the distance is the one
introduced in Quantum Centroid Distance Estimation component. This way, the Page-Hinkley
test spots the presence of drifts when the SSE computed on the new data instances greatly
difers from the one computed on the data instances previously processed. When this happens
the two centroids are recomputed considering the new data instances and the classification
model is updated. The training set used for the current learning round is composed of data
instances of the current data block (containing the new data instances) and those provided
by the component of Data sampling. As explained above, this is done to mitigate the efect of
catastrophic forgetting.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments on real-world datasets</title>
      <p>
        We implemented the proposed framework in IBM Qiskit [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and run experiments by using
simulators on two real-world datasets, more precisely Ozone level detection 1 (having 2536 data
instances, 73 features) and Spambase (having 4600 data instances, 57 features) 2. Data blocks
have been partitioned so as having a portion of 75% of the dataset as labelled data instances
(training sessions) and the remaining 25% as unlabelled data instances (prediction sessions and
testing sets of the evaluation). The classical computing components described in Section 2 are
those available in the toolkit Scikit-learn [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The number of runs of the classification model
to estimate the expectation values is 1024, while the number of iterations (epochs) to optimize
the parameters is 20. The number of layers for the variational quantum circuit is 3. The sample
size of Data sampling is 30% the data-block size.
      </p>
      <p>
        Experiments have been performed to emphasize the impact of the technical configuration
of the framework on the predictive capability, namely number of qubits (corresponding to the
features selected) and size of the data blocks (number of data instances in each training/prediction
session). In Table 1, we report the F1-score values of the proposed framework (named as    )
compared to i) a classical computing solution ( , originally designed for data stream learning)
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and ii) a baseline of the framework that works on the whole dataset (  ). Both CC and
FQC have been tested with the same features selected for HYQOL.
      </p>
      <p>The values illustrated have been computed as the average computed over the data blocks. As
we can see, except two trials, HYQOL does not never worst than CC, even when the number
of qubit is the higher (i.e., 8). Also, we note that the configurations of HYQOL with smallest
set of qubits (i.e., 2) are better than those with largest set (i.e., 8), without, however, particular
discrepancy between the two endpoints. The size of the data blocks seems not be determinant
for the accuracies, but, it is evident that continual learning can be beneficial for quantum-based
classifiers compared to the version that works on the whole dataset (   ). In Table 2, we
report the averaged values of the cost function taken at the last iteration (20-th iteration of the
optimizer of    ). As we can see, the lower costs are reached in correspondence of the
higher predictive performance, that is, data-block size at 200 with 2 and 5 qubits.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Related work</title>
      <p>
        Continual learning, also known as lifelong learning, incremental learning, or sequential learning,
is a field of research dedicated to finding solutions for acquiring knowledge from an
everchanging stream of data involving multiple tasks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        As our knowledge, the works of quantum machine learning focused on continual learning are
very few, quite recent and often in preliminary form. The paper [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] focuses on training sessions
of diferent classification tasks. Specifically, a sequence of quantum state classification tasks is
continually learned by a variational quantum classifier whose parameters are optimized by a
classical gradient-based optimizer. Inspired by the category replay methods (the same we follow
in this work), they propose to constrains the model updated by projecting the gradient direction
1https://archive.ics.uci.edu/ml/datasets/ozone+level+detection
2https://archive.ics.uci.edu/ml/datasets/Spambase
on the region outlined by previous task gradients. This is done also by storing a fraction of the
training data of previous tasks (gradient episodic memory, GEM) on which the gradient descent
is computed. A drawback is the necessity of computing gradients of previous tasks at each
training iteration. In [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], it has been observed that as quantum classifiers are exposed to new
classification tasks, their performance on previous tasks can deteriorate. To address this issue, a
method called elastic weight consolidation (EWC) has been introduced. EWC helps protect the
parameters deemed crucial for the previous tasks from undergoing drastic updates. Numerical
experiments have illustrated that a quantum classifier can continuously learn and adapt to three
distinct classification tasks without experiencing significant forgetting. However, this work
has been tested on classical data (e.g. MNIST images) which originally are not afected by drift
of data distribution and class labels, but characterized by new class labels (tasks). This may
make the changes of the characteristics of data of the same labels partially handled. The hybrid
quantum-classical method described in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] reports the use of classical convolutional neural
networks with quantum layers. To mitigate the performance drop they propose to inject the
information of the regions of prediction identified on visual saliency maps. All these works
deal with the update of the classification model through purely classical computing, while we
present a quantum distance-based method to identify data instances revealing the drift. By the
way, GEM uses previously processed data kept in the memory to modify the gradient of the
current data. Unlike EWC, the performance of the old data is more likely to increase because
the previous data are rehearsed in the parameter update. That is another argument behind our
decision of selecting old samples through quantum techniques.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this paper, we investigated the viability of quantum machine learning solutions to work on the
realistic scenarios of changeability of the statistical properties of the data, which often implies
the variability of the performances of the model. We conjecture this can be a machine learning
problem in which the quantum solutions can lead innovation. On simulated hardware, the
hybrid quantum-classical proposal ofers encouraging results, in terms of accuracy, often better
than a classical computing solution working on data stream and hybrid solution working in
batch mode (no continual learning). As our opinion, three take-home messages can be identified
from this paper. The first one is methodological, in that the continual learning opens to practical
applications able to combine quantum computing and classical computing techniques, which
is likely the only way to concretely use current quantum technologies. The second one is
experimental, in that it provides arguments on the fact that stable quantum devices could even
do better in terms of performances and quality of the results, when used in predictive tasks.
The third one tell us that, although the high-performance computation and tractability of hard
problems are the promises of quantum computing which, with the current devices, often are
not kept, the research on the lifelong computation can be a field in which quantum computing
can already bring interesting results.</p>
      <p>As future work, we plan two investigate two main research lines, one methodological, the
other one applicative. In the first case, we will focus on the technique of optimization of
parameters (component Quantum classification) and study machine/deep learning solutions to
predict the parameter values in order to reduce the computational cost of the classical optimizer.
In the second case, we will focus on more complex data, like social networks and interaction
graphs, in order to handle challenging tasks, like the one of link prediction between actors and
services.</p>
      <p>Corrado Loglisci acknowledges the financial support from the project ”PNRR MUR project
PE0000023-NQSTI” for this research. The paper also achieves the research activities of the project
MUR-Fondo Promozione e Sviluppo - DM 737/2021, CUP: H91I21001630006, INNOAGROECOS.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Preskill</surname>
          </string-name>
          ,
          <article-title>Quantum Computing in the NISQ era and beyond</article-title>
          ,
          <source>Quantum</source>
          <volume>2</volume>
          (
          <year>2018</year>
          )
          <article-title>79</article-title>
          . URL: https://doi.org/10.22331/q-2018
          <source>-08-06-79</source>
          . doi:
          <volume>10</volume>
          .22331/q-2018
          <source>-08-06-79.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Callison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chancellor</surname>
          </string-name>
          ,
          <article-title>Hybrid quantum-classical algorithms in the noisy intermediatescale quantum era and beyond</article-title>
          ,
          <source>Phys. Rev. A</source>
          <volume>106</volume>
          (
          <year>2022</year>
          )
          <article-title>010101</article-title>
          . URL: https://link.aps.org/ doi/10.1103/PhysRevA.106.010101. doi:
          <volume>10</volume>
          .1103/PhysRevA.106.010101.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Loglisci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Diliso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malerba</surname>
          </string-name>
          ,
          <article-title>A hybrid quantum-classical framework for binary classification in online learning</article-title>
          ,
          <source>in: SEBD</source>
          , volume
          <volume>3478</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Overcoming long-term catastrophic forgetting through adversarial neural pruning and synaptic consolidation</article-title>
          ,
          <source>IEEE Trans. Neural Networks Learn. Syst</source>
          .
          <volume>33</volume>
          (
          <year>2022</year>
          )
          <fpage>4243</fpage>
          -
          <lpage>4256</lpage>
          . URL: https://doi.org/10.1109/TNNLS.
          <year>2021</year>
          .
          <volume>3056201</volume>
          . doi:
          <volume>10</volume>
          .1109/TNNLS.
          <year>2021</year>
          .
          <volume>3056201</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>G. I. Parisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kemker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Part</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wermter</surname>
          </string-name>
          ,
          <article-title>Continual lifelong learning with neural networks: A review</article-title>
          ,
          <source>Neural Networks</source>
          <volume>113</volume>
          (
          <year>2019</year>
          )
          <fpage>54</fpage>
          -
          <lpage>71</lpage>
          . URL: https://doi.org/10. 1016/j.neunet.
          <year>2019</year>
          .
          <volume>01</volume>
          .012. doi:
          <volume>10</volume>
          .1016/j.neunet.
          <year>2019</year>
          .
          <volume>01</volume>
          .012.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Lange</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aljundi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Masana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Parisot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leonardis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Slabaugh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tuytelaars</surname>
          </string-name>
          ,
          <article-title>A continual learning survey: Defying forgetting in classification tasks</article-title>
          ,
          <source>IEEE Trans. Pattern Anal. Mach. Intell</source>
          .
          <volume>44</volume>
          (
          <year>2022</year>
          )
          <fpage>3366</fpage>
          -
          <lpage>3385</lpage>
          . URL: https://doi.org/10.1109/ TPAMI.
          <year>2021</year>
          .
          <volume>3057446</volume>
          . doi:
          <volume>10</volume>
          .1109/TPAMI.
          <year>2021</year>
          .
          <volume>3057446</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Arthur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Date</surname>
          </string-name>
          ,
          <article-title>Hybrid quantum-classical neural networks</article-title>
          ,
          <source>in: IEEE International Conference on Quantum Computing and Engineering</source>
          , QCE 2022,
          <article-title>Broomfield</article-title>
          , CO, USA, September
          <volume>18</volume>
          -
          <issue>23</issue>
          ,
          <year>2022</year>
          ,
          <year>2022</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>55</lpage>
          . doi:
          <volume>10</volume>
          .1109/QCE53715.
          <year>2022</year>
          .
          <volume>00023</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chalumuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kune</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. S.</given-names>
            <surname>Manoj</surname>
          </string-name>
          ,
          <article-title>A hybrid classical-quantum approach for multiclass classification</article-title>
          ,
          <source>Quantum Inf. Process</source>
          .
          <volume>20</volume>
          (
          <year>2021</year>
          )
          <article-title>119</article-title>
          . URL: https://doi.org/10.1007/ s11128-021-03029-9. doi:
          <volume>10</volume>
          .1007/s11128-021-03029-9.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Aïmeur</surname>
          </string-name>
          , G. Brassard,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gambs</surname>
          </string-name>
          ,
          <article-title>Machine learning in a quantum world</article-title>
          ,
          <source>in: Advances in Artificial Intelligence, 19th Conference of the Canadian Society for Computational Studies of Intelligence</source>
          ,
          <source>Canadian AI</source>
          <year>2006</year>
          ,
          <string-name>
            <given-names>Québec</given-names>
            <surname>City</surname>
          </string-name>
          , Québec, Canada, June 7-9,
          <year>2006</year>
          , Proceedings,
          <year>2006</year>
          , pp.
          <fpage>431</fpage>
          -
          <lpage>442</lpage>
          . URL: https://doi.org/10.1007/11766247_37. doi:
          <volume>10</volume>
          .1007/ 11766247\_
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kerenidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Landman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Luongo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Prakash, q-means: A quantum algorithm for unsupervised machine learning</article-title>
          , in: H.
          <string-name>
            <surname>M. Wallach</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>d'Alché-</article-title>
          <string-name>
            <surname>Buc</surname>
            ,
            <given-names>E. B.</given-names>
          </string-name>
          <string-name>
            <surname>Fox</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems</source>
          <year>2019</year>
          ,
          <article-title>NeurIPS 2019</article-title>
          , December 8-
          <issue>14</issue>
          ,
          <year>2019</year>
          , Vancouver, BC, Canada,
          <year>2019</year>
          , pp.
          <fpage>4136</fpage>
          -
          <lpage>4146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Nielsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. L.</given-names>
            <surname>Chuang</surname>
          </string-name>
          ,
          <source>Quantum Computation and Quantum Information</source>
          , Cambridge University Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wierichs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Izaac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          , C. Y.
          <string-name>
            <surname>-Y. Lin</surname>
          </string-name>
          ,
          <article-title>General parameter-shift rules for quantum gradients</article-title>
          ,
          <source>Quantum</source>
          <volume>6</volume>
          (
          <year>2022</year>
          )
          <article-title>677</article-title>
          . URL: https://doi.org/10.22331/q-2022
          <source>-03-30-677</source>
          . doi:
          <volume>10</volume>
          . 22331/q-2022
          <source>-03-30-677.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gama</surname>
          </string-name>
          , M. G. (Eds),
          <source>Learning from Data Streams - Processing techniques in Sensor Networks</source>
          , Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Page</surname>
          </string-name>
          , Continuous inspection schemes,
          <source>Biometrika</source>
          <volume>41</volume>
          (
          <year>1954</year>
          )
          <fpage>100</fpage>
          -
          <lpage>115</lpage>
          . URL: http: //www.jstor.org/stable/2333009.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Anis</surname>
          </string-name>
          , et al,
          <article-title>Qiskit: An open-source framework for quantum computing</article-title>
          ,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          . 5281/zenodo.2573505.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , E. Duchesnay,
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Montiel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Read</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bifet</surname>
          </string-name>
          , T. Abdessalem,
          <article-title>Scikit-multiflow: A multi-output streaming framework</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>19</volume>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . URL: http://jmlr.org/ papers/v19/
          <fpage>18</fpage>
          -
          <lpage>251</lpage>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <article-title>Quantum continual learning of quantum data realizing knowledge backward transfer, Physica A: Statistical Mechanics</article-title>
          and its
          <string-name>
            <surname>Applications</surname>
          </string-name>
          (
          <year>2023</year>
          )
          <article-title>128779</article-title>
          . doi:https://doi. org/10.1016/j.physa.
          <year>2023</year>
          .
          <volume>128779</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>W.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <article-title>Quantum continual learning overcoming catastrophic forgetting</article-title>
          ,
          <source>CoRR abs/2108</source>
          .02786 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2108.02786. arXiv:
          <volume>2108</volume>
          .
          <fpage>02786</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>Cqural: A novel CNN based hybrid architecture for quantum continual machine learning</article-title>
          ,
          <source>CoRR abs/2305</source>
          .09738 (
          <year>2023</year>
          ). URL: https://doi.org/10.48550/arXiv.2305.09738. doi:
          <volume>10</volume>
          .48550/arXiv.2305.09738. arXiv:
          <volume>2305</volume>
          .
          <fpage>09738</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>