<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>HC@AIxIA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Reinforcement Learning and Fuzzy Logic Modelling for Personalized Dynamic Treatment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Locatelli</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Clemens Cerioli</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Besozzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arjen Hommersom</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Stella</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bicocca Bioinformatics Biostatistics and Bioimaging Centre - B4, University of Milano-Bicocca</institution>
          ,
          <addr-line>Vedano al Lambro (MB)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, Open University</institution>
          ,
          <addr-line>P.O. Box 2960, 6401DL Heerlen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Informatics</institution>
          ,
          <addr-line>Systems and Communication</addr-line>
          ,
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>3</volume>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The treatment of patients sufering from chronic diseases is a dificult problem to be tackled. Its complexity mainly originates from the following sources: the patient-specific response to the prescribed therapy, the impact of the interplay between disease and therapy on the quality of life of the patient and relatives, and the economic costs incurred by the healthcare system. Recently, there has been considerable interest in developing, studying, and applying artificial intelligence methods to diagnosis, prognosis and treatment personalization. This paper combines two techniques from artificial intelligence, namely fuzzy logic and reinforcement learning, to develop optimal dynamic treatment for patients sufering from a chronic disease. In this paper, we focus on cancer as a chronic disease and leverage a biologically validated fuzzy logic model from the literature. Diferent problem settings, of increasing complexity, are taken into account, presented and analyzed. Results of an extensive numerical experimental plan confirm the potential of non-myopic decision-making when treating chronic disease patients.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Personalized dynamic treatment</kwd>
        <kwd>Fuzzy Logic</kwd>
        <kwd>Reinforcement Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Biomedical researchers, clinicians and healthcare practitioners are increasingly recognizing that the one
size fits all approach is no longer acceptable [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Indeed, biological interactions are known to happen at
an individual level while this is not true for statistical interactions, which instead occur at a population
level [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This heterogeneity stems from stochastic processes taking place at the molecular scale and
inducing biological noise [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which is not only well recognized as an indispensable trait in evolution
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] but also represents a critical feature that might be exploited in the identification of personalized
treatments [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. As highlighted in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], it is commonly agreed upon that drawing conclusions, diagnosing,
and making treatment decisions from population models may not be optimal for a specific patient.
Thus, the scientific community is developing new approaches based on the rationale that each patient
is unique.
      </p>
      <p>
        Personalized medicine ofers a tailored approach to healthcare by customizing medical treatments
according to each patient’s unique characteristics. A key component of personalized medicine are
dynamic treatment regimes (DTRs), which involve sequences of decision rules adjusted over time
depending on how the health condition of the patient evolves and the response to past treatments.
The development of DTRs [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8, 9, 10</xref>
        ] relies on statistical methods, such as reinforcement learning and
causal inference, to identify the optimal action, at each decision point, based on historical data. In
recent years, statistical methods for optimizing DTRs have made significant progress. Techniques like
Q-learning [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], A-learning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and structural nested models [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] allow researchers to estimate optimal
treatment strategies from both observational and experimental data. These methods are designed to
handle complex, multi-stage decision processes and account for time-varying confounders that can
afect both treatment decisions and outcomes. DTR algorithms have been successfully applied to develop
personalized treatment plans for chronic diseases, including cancer, diabetes, anaemia, HIV, and various
mental health disorders, as discussed in [12].
      </p>
      <p>From a diferent yet complementary perspective, personalized medicine also benefits from the
mathematical modeling and analysis of cellular processes, whose malfunctioning often lead to the
onset of diseases. Though, the definition of models able to provide an accurate description of biological
mechanisms still poses several challenges, which are mainly related to the efects of stochasticity, the
paucity of precise quantitative data, and the technical limits in the observation or measurement of an
organism as a whole. In this context, fuzzy logic [13] can be fruitfully exploited in the biomedical field to
describe and analyze complex biological systems as well as multifactorial diseases. The most important
features of fuzzy logic are the capacity to deal with uncertain data and loosely defined variables, not to
mention the ease of interpretation that represent a fundamental characteristic in clinical scenarios.</p>
      <p>To date, there are many applications of fuzzy logic as a means for clinical decision making within a
personalized medicine approach. By way of example, in [14] an automatically defined fuzzy rule base
was used to classify lung cancer patients based on transcriptomic data. In [15], a fuzzy classification
model was exploited to aid in vasopressor administration in intensive care unit (ICU) patients: similar
patients were grouped by a fuzzy clustering algorithm, and an ensemble fuzzy classifier was trained on
ICU data to estimate the necessity of vasopressor administration. Nevertheless, these approaches fail to
take into account the intricate dynamics of biological systems, which is due to a finely orchestrated
mechanism of positive and negative feedbacks, and their descriptive power is thus limited. This kind of
dynamics is usually analyzed using systems of ordinary diferential equations (ODEs), whose downside
is the need of well-defined crisp values for the model parameters. To bypass this issue, in [ 16], ODEs
were complemented with fuzzy logic to predict the disease course of HIV patients taking into account
the strength of their immune system. Fuzzy logic was also hybridized with other computational and
statistical methods to make the most of them. For instance, in [17], fuzzy logic was integrated into
causal inference to introduce fuzzy causal efect metrics able to take into account vague and imprecise
data and, as a practical example, the approach was used to measure the impact of age and sodium intake
on blood pressure.</p>
      <p>In this work, we address the DTR problem by leveraging a combined framework based on fuzzy logic
modeling and reinforcement learning. The paper is organized as follows. Section 2 gives the concept of
personalized dynamic treatment, while reviewing the main approaches from the specialized literature.
Dynamic fuzzy logic modelling, specifically tailored on biomedical processes, is introduced in Section 3,
together with definitions and main concepts on reinforcement learning. Section 4 presents a proof of
concept where these two methods are integrated to develop DTRs for optimizing the death of cancer
cells by apoptosis. The results of three sets of numerical experiments are reported and commented on
in Section 5. Conclusions and direction for future works are given in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Personalized Dynamic Treatment</title>
      <p>Managing chronic diseases, i.e., diseases which evolve over time with prolonged duration, requires
designing and developing methods for assisting physicians to make personalized treatment decisions.
In this work we informally define a personalized treatment decision as a treatment decision depending
on the unique characteristics of the patient, with the goal of optimizing their long term outcome while
reducing the risk of side-efects. Moreover, when treating chronic patients, the treatment decisions
unfold temporally depending on the progression of the patient’s illness state. In such a setting, dynamic
treatment regimes (DTRs) ofer a robust framework for making personalized treatment decisions over
time by leveraging patient’s characteristics and preferences, disease progression, as well as on the
patient’s response to past treatments.</p>
      <p>
        Given the inherently sequential nature of chronic disease progression and treatment, reinforcement
learning (RL) [18] ofers an efective option for developing DTRs. Indeed, the ability of RL to solve
sequential decision making problems naturally deals with the challenges of managing chronic diseases
with DTRs, as witnessed by several estimation methods developed for DTRs. As an example, Q-Learning
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is a model-free approach that estimates the optimal DTR by learning the action-value function. A
second approach is ofered by A-Learning [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], which directly estimates the advantage function, while
ofering eficiency gain when the baseline treatment is well-understood. Further advancements were
made in [19], where regret functions are incorporated into a regression model for the final observed
result, while in [20] a modification of the Q-learning approach was proposed, where both the observed
rewards and the estimated loss due to sub-optimal actions are used. Following this, several machine
learning methods were proposed in the specialized literature to approximate the Q-function.
      </p>
      <p>It is also worth mentioning the great interest received by artificial neural networks (ANNs) [ 21],
which resulted in the development of Deep Q-learning (DQL) [22], and its subsequent application to
DTRs by means of feed-forward neural networks (FNNs) [23] or via convolutional neural networks
(CNNs) to approximate the regret function [24]. More recently, adversarial networks were used to learn
a policy mimicking successful treatments while avoiding unsuccessful ones [25]. Later on, the concept
of pessimism and Bayesian machine learning (also Bayesian neural networks) methods were integrated
to find the best treatment strategy [ 26].</p>
      <p>An alternative approach aims to maximize the reward by learning the optimal policy or value directly.
Inverse probability of treatment weighting (IPTW) was used in [27] and [28] to estimate the causal
efect from observational data. This technique computes weights for each policy based on the value of
covariates, thus allowing the estimation of the optimal decision rule as a weighted classification problem.
Subsequently, modifications of this approach were proposed, mainly focused on the combination of
Q-function or A-function approximations with the IPTW estimator [29, 30]. Deep learning techniques
were also used to this extent: for instance, in [31] recurrent neural networks and multilayer perceptron
networks were used to estimate the value function of a DTR. Over the years, other machine learning
techniques have also been used, such as decision trees [32, 33] or methods focused on causality: [34]
based its model on causal trees, while [35] and [36] overcame the problem of non-identifiability through
causal-bounds and the use of instrumental variables, respectively.</p>
      <p>Up to this point, the focus has been on the methods used in the standard DTR setting, while the
problem of DTR has several facets. A main challenge is that of competing outcomes, i.e., the case when
multiple and competing outcomes exist, which requires balancing treatment decisions over time [37, 38].
Censored data, where the full information about the outcome is not observed, further complicates the
analysis, as does missing data, which lead to incomplete observations and potential biases [39]. An
infinite or indefinite time horizon adds another layer of dificulty, as it requires considering the long-term
efects of treatments beyond a finite period [ 40, 41]. Finally, challenges such as time-to-event analysis,
which measures the duration until an event occurs [42], and time-to-visit optimization, which estimates
the best timing for treatment interventions [43], further add complexity to developing efective DTRs.</p>
      <p>To the best of the authors’ knowledge, the only work using DTRs to optimize daily dosage adjustment
of radiotherapy within a clinical decision support system is the one proposed in [44]. Nevertheless,
many works validate their methods on data from real-world, though perhaps simplified, problems.
The field of DTRs has seen recent growth, while at the current state of the art there is a significant
research-practice gap, probably due to many of the presented methodologies being evaluated unevenly
and because they often consider simplified problems when compared to those that would be found in
real clinical settings. An additional limitation to the widespread adoption of these DTR optimization
methodologies is the potential lack of acceptance by domain experts. To gain their trust, these methods
need to provide interpretable results and clearly explain the decision-making process behind their
treatment recommendation.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <sec id="sec-3-1">
        <title>3.1. Fuzzy Logic Modelling of Biomedical Processes</title>
        <p>Dynamic fuzzy models (DFMs) represent a mathematical formalism—based on the concepts of fuzzy
logic [45, 46]—introduced in [47] to describe and analyze complex systems that consist of heterogeneous
components and are characterized by uncertainty, such as those related to biomedical contexts. DFMs can
be straightforwardly used to simulate the evolution over time of the system state, therefore predicting
its emergent behaviour in diferent conditions, such as physiological or pathological states of cellular
systems and individuals. In this section we provide a brief description of DFMs, and refer the reader to
[13, 48, 47, 49, 50] for further details.</p>
        <p>
          The definition of a DFM requires the identification of a directed graph representing the system
components 1, . . . ,  and their mutual interactions, which can correspond to either positive or
negative regulations. Each component is formalized by means of a linguistic variable, that is, a list of
linguistic terms together with the corresponding fuzzy sets and membership functions, defined over a
proper universe of discourse. Formally, for each  = 1, . . . , ,  is described by a linguistic variable
 = {, , Λ , }, where:
•  is the name of the linguistic variable;
•  is the universe of discourse, that is, the range of values in which  has meaning;
• Λ  = { ,1, . . . ,  , } is the set including the  linguistic terms that  can assume in ;
•  = {  ,1 , . . . ,   , } is the set of  fuzzy sets in which the universe of discourse is partitioned,
each one corresponding to a linguistic term appearing in Λ . Note that the symbol  can be used
to denote both the fuzzy set and its membership function: namely, a fuzzy set   , ,  = 1, . . . ,  ,
is uniquely characterized by a membership function   , :  → [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ], which maps every value
 ∈  to the so-called degree of membership of  to the fuzzy set   , .
        </p>
        <p>The interactions among the system components are formalized as fuzzy rules, which are conditional
statements consisting in an antecedent and a consequent written in the form “IF  IS  , THEN 
IS  ,”, where  , and  , are linguistic terms associated with variables  and  , respectively. In
general, the antecedent of a fuzzy rule can contain any logic expression involving two or more linguistic
variables connected by the fuzzy operators AND, OR, NOT. Figure 1 shows an example of fuzzy rules
(panel B) and fuzzy sets (panel C) associated with two linguistic variables,  and , which control a
third variable,  (panel A).</p>
        <p>
          The set of linguistic variables and fuzzy rules appearing in a DFM constitute a Fuzzy Inference System
(FIS). The simulation of DFMs can be performed by means of Simpful [50], a Python library designed to
easily define and analyse FISs based on Mamdani or Takagi–Sugeno fuzzy inference engines [ 46]. In this
work, we rely on the 0-order Sugeno fuzzy reasoning, as implemented in Simpful. To determine the next
state of any variable, this method calculates a weighted aggregation of the output values produced by all
fuzzy rules having that variable in their consequent and being satisfied at the current time step [ 51]. To
exemplify this step, in Figure 1 (panel D) we show how the next value  of variable  will be calculated
according to the fuzzy rules 1, 2 and 4, 5 that—given the current values of variables  and ,
respectively—are the only rules whose degree of satisfaction (denoted by 1, . . . , 5) is not equal to
zero. The degree of satisfaction of a fuzzy rule is a number in [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] to which the current values of the
(input) linguistic variables match the antecedent. In particular, for simple rules where no fuzzy operators
appear in the antecedent (as those shown in Figure 1, Panel B), the degree of satisfaction is equal to the
membership function value so that 1 =  Less-Functional() = 0.25, 2 =  Medium-Functional() = 0.75
and 3 =  More-Functional() = 0 for  = 4, while 4 =  Slow() = 0.1 and 5 =  Fast() = 0.9 for
 = 0.8 (see Figure 1, Panel D). The state of all variables is then updated, and the process is iterated
until a maximum time step is reached. Panel E in Figure 1 shows the overall process of calculating the
output of fuzzy rules, which includes fuzzification, fuzzy inference and defuzzification [ 46].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Reinforcement Learning</title>
        <p>
          Reinforcement learning (RL) is a machine learning methodology where an agent learns to make decisions
by trial-and-error interactions with its environment. In general, solving an RL task means learning an
optimal way to choose the set of actions, or learning an optimal policy, so as to maximise the cumulative
discounted outcome, denoted as . This cumulative reward is calculated as the sum of discounted
future rewards, which can be represented as
 =
 − − 1
∑︁   ++1,
=0
(1)
where  ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] is the discounted factor,  the outcome at , and  is the final time step.
        </p>
        <p>In the context of DTR, a policy is defined as a sequence of decision rules  = {1, 2, ...,  } where
each decision rule  corresponds to a treatment decision at time . A policy maps from the history space
ℋ to the set of possible treatment actions . An optimal policy, denoted as * , can be defined as a
policy that returns the largest (or equal) cumulative outcome in comparison to other policies. Formally:
* = arg max E

[︃  − − 1 ]︃
∑︁   ++1 .</p>
        <p>=0</p>
        <p>Another fundamental concept is the Q-function (or the action-value function), which provides the
expected cumulative reward starting from a given patient history ℎ at time , after taking an action 
and thereafter following the policy . It essentially quantifies the value of taking action  given the
patient history ℎ and following the policy  at time . Formally, the Q-function is defined as:
(ℎ, ) = E
[︃  − − 1 ]︃
∑︁   ++1| = ℎ,  =  .</p>
        <p>=0
Subsequently, the optimal Q-function can be defined as a function that maximizes the expected return
for each history-action pair under any policy , formally:
* (ℎ, ) = max (ℎ, ).</p>
        <p>It is possible to define Equation 4 recursively without making any reference to a particular policy. This
key property of RL is given by the Bellman optimality equation [52] and can be expressed formally as
follows:
* (ℎ, ) = E[+1 +  max *+1(ℎ+1, +1)|ℎ, ].</p>
        <p>+1</p>
        <p>
          Q-learning [
          <xref ref-type="bibr" rid="ref11">11, 53</xref>
          ] is a popular RL algorithm used to learn the optimal action-value function.
However, standard Q-learning can struggle in high-dimensional or continuous spaces, so approximation
techniques, such as linear regression, are often employed to estimate Q-values. Suppose we use a
Qfunction approximation (ℎ, ;  ), where   are the model parameters. By iteratively approximating
 for each time step, it is possible to determine the optimal treatment policy * .
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Optimizing the apoptotic death of cancer cells</title>
      <p>As a proof of concept of the efectiveness of combining fuzzy logic modelling with reinforcement
learning, we consider a DFM describing cell death processes in oncogenic K-ras cancer cells as a result of
progressive glucose depletion. Ras mutations are among the most frequent alterations in human cancers,
especially in solid tumors, and are linked to marked metabolic rewiring, which has been recognized
as one of the numerous hallmarks of cancer [54]. Oncogenic K-ras cancer cells have been considered
undruggable for long, and available therapeutic strategies—especially the single-agent targeted ones,
most often leading to acquired resistance—are still far from being fully curative [55]. Huge research
eforts are thus in progress to discover efective treatments and improve the outcome of patients afected
by Ras-mutant malignancies [56].</p>
      <p>The DFM of oncogenic K-ras cancer cells, defined in [ 47], takes into account the main cellular
components involved in energy production (e.g. glucose and glycolysis), diferent mitochondrial
processes (e.g. the mitochondrial potential variation DeltaPsi and the activity of the mitochondrial
Complex I (CI)), several processes and proteins involved in cellular adhesion, in the regulation of cell
death (e.g. unfolded protein response (UPR)) and in survival mechanisms (e.g. protein kinase A (PKA)),
as well as the phenotypes related to apoptosis and necrosis. Three variables in this DFM represent
the input of the system: (1) glucose, whose consumption is modeled by a custom update function that
drives the dynamical update of the system state; (2) Ras-GTP, which is always set to the High value to
mimic the hyperactivation of K-ras in cancer cells; (3) PKA, which mainly regulates cancer cells fate: it
is set to either the High or the Low state to mimic the ability of cancer cells to survive or die under
glucose starvation, respectively. We refer to [47] for further details on the relevance and role of all
model components.
(2)
(3)
(4)
(5)</p>
      <p>The model was validated through an extensive comparison with experimental data measured on
the MDA-MB-231 cancer cell line (human K-ras-mutated breast cancer) and on NIH3T3 K-ras cells
(mouse fibroblasts transformed by oncogenic K-ras expression) [ 47]. In previous works [47, 49], global
optimization algorithms were also combined with the simulation of this DFM to identify a set of
treatments able to maximize cancer cell death by apoptosis—and possibly minimize necrotic processes
too—while minimizing the total number of administered drugs. This approach allowed to identify
treatments that were already validated in the literature, as well as novel potential therapeutic strategies.</p>
      <p>The DFM of oncogenic K-ras cancer cell represents a valid case study for the novel methodology of
personalized dynamic treatment that we are introducing with this work. Indeed, despite considering
only processes occurring at the cellular level, it allows us to take into account the possible efects of
biological noise that, by inducing a high heterogeneity among cells, can act as a proxy to mimic the
status of patients at the individual level. In addition, it holds the added value of previous experimental
validation for various drug administrations, some of which have also been considered in this work.
Namely, here we focus on three diferent types of treatment:
• Complex I inhibition: disrupting Complex I in these cells generates harmful reactive oxygen
species, leading to oxidative stress and triggering apoptotic pathways;
• UPR activation: UPR is a cellular mechanism triggered by significant stress, such as nutrient
deprivation or oxidative damage. Prolonged activation of this process (i.e. "UPR is High") leads to
apoptotic cell death;
• Combination of inhibition and activation: this will simulate a condition characterized by a
high level of UPR activation (75%) and a lower level of concurrent Complex I inhibition (25%).</p>
      <p>This combination was not experimentally tested but might have the efect of increasing apoptosis.</p>
      <p>In this study we applied a DTR approach to optimize the treatment strategy for inducing cancer cell
death. In particular, a treatment strategy is represented as a policy, that is, a set of decision rules that
guide treatment choices over multiple stages based on the evolution of the disease. At each decision
point , a treatment  is selected from a set of possible treatments. These decisions are determined by
the policy, which maps the complete history of the disease progression to a treatment with the objective
of maximising the desired outcome. In this case, the desired outcome is the induction of apoptosis in
cancer cells. To achieve this, we first generated a dataset by simulating cancer cell behaviour using the
fuzzy cellular model over a 72-hour period, with treatment decisions made each 12 hours. The dataset
includes:
• the treatment assignment for each decision step. The choice of each treatment was based on
three covariates: PKA, CI and UPR. The static rules used to select the treatment are based on
previous results [47];
• the values of the three covariates at the beginning of the simulation and each decision step;
• the resulting state of the system at each decision step and at the end of the interval.
Using the generated dataset, we applied a DTR method to determine the optimal treatment strategy. The
DTR algorithm considers the three covariates at each decision point and recommends the best treatment
option based on the current state of the system. The recurring nature of the treatment assignment
allows for adaptive decision making, accounting for the evolving state of the system (Figure 2).</p>
      <p>To evaluate the efectiveness of the proposed approach we designed three numerical experiments
with diferent goals:
1. Maximizing apoptosis in cancer cells: In the first experiment the goal is to maximize cancer
cell death via apoptosis. We apply the DTR algorithm to our fuzzy cellular model, while using
the current apoptosis value as the primary outcome measure.
2. Balancing apoptosis and necrosis: The second experiment aims to find a balanced treatment
strategy that not only maximizes apoptosis but also minimizes uncontrolled cell death (necrosis),
which can lead to undesirable inflammation efects. To achieve this, we modify our DTR algorithm
to account for both apoptosis and necrosis in the outcome.
3. Impact of treatment intervals on DTR efectiveness : The goal of this experiment is to
investigate how diferent treatment intervals afect the performance of DTR strategies. We vary
the decision-making intervals from the original 12 hours, used in the first two types of experiments,
to include 6-hour and 24-hour intervals, generating a new dataset for each setting. For each
interval setting, we run the DTR algorithm to determine the optimal strategy for maximizing
apoptosis (as in the first experiment).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <sec id="sec-5-1">
        <title>5.1. Experimental setup</title>
        <p>This section presents and discusses the results of the numerical experiments performed under three
diferent experimental settings to evaluate the efectiveness of the proposed DTR approach 1. For
each experimental setting, we generated 30,000 samples and evaluated the results on a cohort of 50
samples, enabling direct comparison across diferent treatment strategies. The eficacy of the DTR
was specifically assessed against static strategies, including the consistent inhibition of Complex I, the
continuous activation of the UPR and the continuous usage of the combined treatment. The values of
apoptosis and necrosis are derived from the fuzzy model, where the outcome is expressed as a fuzzy
measure: for each individual sample, the value of apoptosis and necrosis will always range between 0
and 1. Thus, when considering a dataset of 50 samples, the maximum possible cumulative value for
both apoptosis and necrosis is 50.</p>
        <p>In particular, for each experimental setting, the available data consists of a set of finite horizon
trajectories:</p>
        <p>{PKA, 0, 0, 0, . . . , 5, 5, 5, 6},
where each  consists of {CI, UPR},  is the value of the apoptosis at each decision step and 6 is
the value of apoptosis after 72 hours. In order to maximize a sum of numerical rewards, we employed a
recursive form of Q-learning, with (, ) predicting:
ˆ = +1 +  max ˆ+1(ℎ+1, +1),
+1
(6)
where ℎ+1 = {pka, 0, 0, 0, . . . , , +1, +1} and ˆ+1 is the estimator of the Q-values. To
obtain the estimator ˆ we utilize the Random Forest Regressor [57] for fitting  backward and get
{ˆ5, ˆ4, . . . , ˆ0}. It is important to note that the performance of our model is afected by the choice
of hyperparameters:</p>
        <sec id="sec-5-1-1">
          <title>1The code is provided at: https://github.com/mLoca/FuzzyReinforcement.</title>
          <p>•  , the discount factor, quantifies how much importance we give to future rewards;
• maximum tree depth, a parameter that controls the complexity of individual decision trees and
avoids overfitting.</p>
          <p>The optimal values for the discount factor and the maximum tree depth were determined through a
process of hyperparameter optimization.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Experimental results</title>
        <p>5.2.1. Experiment 1: Maximizing Apoptosis in Cancer Cells
This experimental setting was designed to identify the most efective treatment strategy when the goal
is to maximize apoptosis in cancer cells, an indicator that the treatment is successful to target cancer
cells, a primary goal in the field of cancer treatment. The discount factor was set to  = 0.95, while
max depth was set to 5.</p>
        <p>The results show that the proposed DTR approach is more efective than static treatment strategies.
Indeed, DTR achieves significantly higher levels of apoptosis (Table 1) than all static treatment strategies.
Furthermore, DTR shows a controlled and progressive increase in apoptosis rates (Figure 3), while also
maintaining a high level of apoptosis throughout the 72-hour horizon. This indicates that DTR not only
achieves optimal results at the end of the treatment horizon but also carefully manages the intermediate
stages of the treatment, ensuring an efective therapeutic response.</p>
        <p>DTR
CI inhibition
UPR activation</p>
        <p>Combined</p>
        <sec id="sec-5-2-1">
          <title>Final</title>
          <p>Apoptosis
34.63
25.66
30.09
31.54</p>
        </sec>
        <sec id="sec-5-2-2">
          <title>Mean (SD)</title>
          <p>By progressively increasing the apoptotic response over the 72-hour period, DTR demonstrates its
capacity to balance short-term objectives with long-term goals, providing optimal results at every
stage and maintaining high levels of apoptosis throughout the treatment process. This continuous
progression highlights DTR’s capacity to optimize the overall outcome without any compromise in
eficacy at any point.
5.2.2. Experiment 2: Balancing Apoptosis and Necrosis
The experimental setting aims to identify the optimal strategy when the goal is to simultaneously
maximize apoptosis and minimize necrosis. In the context of cancer therapy, the objective of maximizing
apoptosis while minimizing necrosis is to improve outcomes by reducing the incidence of unwanted
side efects.</p>
          <p>In this context, the outcomes are  and , denoting apoptosis and necrosis at time , respectively.
Two sets of Q-value estimators {ˆ, ˆ} were used to estimate these outcomes. Specifically, the
(a)
(b)
predicted values for apoptosis and necrosis are:
ˆ = +1 +  max ˆ+1 (ℎ+1, +1) for apoptosis,
 +1
ˆ = +1 +  max ˆ+1 (ℎ+1, +1) for necrosis,</p>
          <p>+1
where ℎ+1 = {pka, 0, 0, 0, 0, . . . , , +1, +1 , +1 }. The optimal policies are identified
through the maximisation of the diference between apoptosis and necrosis:
ˆ(ℎ) = arg max ˆ(ℎ+1, +1) − ˆ(ℎ+1, +1).</p>
          <p>The results of this experiment, summarised in Figure 4 and Table 2 (with  = 0.9 and max depth set
to 5), show that DTR fails to identify a treatment combination which is better than the static UPR
activation strategy.</p>
          <p>However, CI inhibition leads to lower apoptosis and higher necrosis, while the combined treatment
produces comparable apoptosis but increased necrosis, indicating a less favorable balance. In conclusion,
these results demonstrate that DTR is not simply attempting combinations in an efort to improve the
outcomes. Instead, it identifies that the static strategy of UPR activation remains the optimal approach.
This underscores the algorithm’s capacity to make well-informed decisions based on the data, thereby
further confirming its potential for optimizing treatment strategies in cancer therapy.
5.2.3. Experiment 3: Impact of Treatment Intervals
The objective of this experimental setting is to evaluate the sensitivity of treatment strategies with
respect to treatment assignment frequency. This experimental setting provides valuable information to
optimize treatment regimens by investigating the efect of treatment frequency on their overall eficacy
and behavior throughout the treatment horizon.</p>
          <p>Specifically, we used  = 0.9 and max depth set to 4 for the 6-hourly strategy, while  = 0.95 and
max depth set to 5 were used for the daily (24-hourly) allocation strategy. As shown in Figure 5, the
outcomes of the DTR strategies at diferent frequencies are quite similar, with the 6-hourly treatment
assignment generating an average result of 0.693 ± 0.070 of apoptosis value, and the daily assignment
inducing 0.695 ± 0.065. The DTR strategy outperforms static treatment strategies in terms of eficacy
and consistency across both dosing schedules (6-hourly and 24-hourly). The highest levels of cumulative
apoptosis are achieved and maintained throughout the 72-hour. In general, all the strategies show
comparable overall eficacy to 6-hourly dosing by 72 hours.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and Future work</title>
      <p>In this study, we explored the integration of fuzzy logic modeling and reinforcement learning to enhance
personalized dynamic treatment regimes (DTRs), thereby ofering a tailored approach to therapeutic
treatments based on individual disease progression. This integration provides a robust framework
for addressing the complexities of chronic disease management: fuzzy logic captures uncertainty and
models complex biological systems, while reinforcement learning dynamically optimizes treatment
plans over time. Our findings highlight the potential of combining these advanced computational
techniques to improve treatment outcomes through non-myopic decision-making, leading to more
efective and adaptive treatment strategies.</p>
      <p>This work represents a significant advancement in the field of DTRs and a promising extension of the
current state of the art. In contrast to conventional DTR approaches, our methodology employs dynamic
fuzzy logic-based models (DFMs) to simulate how the biological processes involved in the onset and
progression of diseases change over time, thereby enhancing the interpretability of disease-related
mechanisms and their corresponding emergent behaviors at both local (cellular) and global (organism)
levels. Subsequently, the reinforcement learning component maximizes outcomes through treatment
decisions informed by the evolving system state. This novel combination of DFMs with reinforcement
learning supports more adaptive and biologically informed decision-making, thereby ofering a potential
pathway for future validation in diverse clinical settings.</p>
      <p>Despite these promising advancements, several limitations must be acknowledged. A key challenge
lies in the limited range of treatments that were explored in the context of programmed death of cancer
cells. Expanding this approach to include additional drugs from the literature would provide a more
comprehensive perspective. Even more important, it would be valuable to examine how the treatment
strategy evolves in response to diferent glucose depletion dynamics. Moreover, while our models
have shown eficacy in controlled experimental environments, their efectiveness in real-world clinical
settings remains unverified and requires further validation. It is also worth stating that, although the
framework demonstrated favorable outcomes with a synthetic dataset, the adequacy of the sample
size for practical applications remains uncertain. Future research should determine the optimal sample
size required to ensure the eficacy of the model in a healthcare environment, recognizing that small
datasets may not entirely capture the complexity needed for reliable treatment personalization.</p>
      <p>In conclusion, this research establishes a foundation for a more personalized approach to chronic
disease management, emphasizing the critical role of individual characteristics in guiding treatment
decisions. By addressing the current limitations and continuously refining these methodologies, we can
advance toward more efective and tailored healthcare strategies.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was supported by Fondazione Regionale per la Ricerca Biomedica (Regione Lombardia),
Project ERAPERMED2022-258, GA 779282. In particular the PhD Fellowship of M. Locatelli was funded
by the MG-PerMed research project: “Personalising myasthenia gravis medicine: from “one-fits-all” to
patient-specific immunosuppression” (ERAPERMED2022-258, GA 779282).</p>
      <p>This work was partially supported by the MUR under the grant “Dipartimenti di Eccellenza 2023-2027”
of the Department of Informatics, Systems and Communication of the University of Milano-Bicocca,
Italy, and by the National Plan for NRRP Complementary Investments (PNC, established with the
decree-law 6 May 2021, n. 59, converted by law n. 101 of 2021) in the call for the funding of research
initiatives for technologies and innovative trajectories in the health and care sectors (Directorial Decree
n. 931 of 06-06-2022) - project n. PNC0000003 - AdvaNced Technologies for Human-centrEd Medicine
(ANTHEM). This work reflects only the authors’ views and opinions; neither the Ministry for University
and Research nor the European Commission can be considered responsible for them.
[12] C. Yu, J. Liu, S. Nemati, G. Yin, Reinforcement learning in healthcare: A survey, ACM Computing</p>
      <p>Surveys (CSUR) 55 (2021) 1–36.
[13] P. Cazzaniga, S. Spolaor, C. Fuchs, M. S. Nobile, D. Besozzi, et al., Fuzzy logic for knowledge-driven
and data-driven modeling in biomedical sciences, in: B. Carpentieri, P. Lecca (Eds.), Big Data
Analysis and Artificial Intelligence for Medical Sciences, John Wiley &amp; Sons, 2024, p. 17.
[14] N. Potie, S. Giannoukakos, M. Hackenberg, A. Fernandez, On the need of interpretability for
biomedical applications: Using fuzzy models for lung cancer prediction with liquid biopsy, in:
2019 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), IEEE, 2019, pp. 1–6.
[15] C. M. Salgado, S. M. Vieira, L. F. Mendonça, S. Finkelstein, J. M. Sousa, Ensemble fuzzy models in
personalized medicine: Application to vasopressors administration, Engineering Applications of
Artificial Intelligence 49 (2016) 141–148.
[16] H. Zarei, A. V. Kamyad, A. A. Heydari, Fuzzy modeling and control of HIV infection, Computational
and Mathematical Methods in Medicine 2012 (2012) 893474.
[17] A. Saki, U. Faghihi, Integrating fuzzy logic with causal inference: Enhancing the Pearl and
Neyman</p>
      <p>Rubin methodologies, 2024. arXiv:2406.13731.
[18] R. S. Sutton, Reinforcement learning: an introduction, A Bradford Book (2018).
[19] R. Henderson, P. Ansell, D. Alshibani, Regret-regression for optimal dynamic treatment regimes,</p>
      <p>Biometrics 66 (2010) 1192–1201.
[20] X. Huang, S. Choi, L. Wang, P. F. Thall, Optimization of multi-stage dynamic treatment regimes
utilizing accumulated data, Statistics in medicine 34 (2015) 3424–3443.
[21] Y. Bengio, I. Goodfellow, A. Courville, Deep learning, volume 1, MIT press Cambridge, MA, USA,
2017.
[22] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller,
A. K. Fidjeland, G. Ostrovski, et al., Human-level control through deep reinforcement learning,
nature 518 (2015) 529–533.
[23] Y. Liu, B. Logan, N. Liu, Z. Xu, J. Tang, Y. Wang, Deep reinforcement learning for dynamic
treatment regimes on medical registry data, in: 2017 IEEE international conference on healthcare
informatics (ICHI), IEEE, 2017, pp. 380–385.
[24] S. Liang, W. Lu, R. Song, Deep advantage learning for optimal dynamic treatment regime, Statistical
theory and related fields 2 (2018) 80–88.
[25] L. Wang, W. Yu, X. He, W. Cheng, M. R. Ren, W. Wang, B. Zong, H. Chen, H. Zha, Adversarial
cooperative imitation learning for dynamic treatment regimes, in: Proceedings of The Web
Conference 2020, 2020, pp. 1785–1795.
[26] Y. Zhou, Z. Qi, C. Shi, L. Li, Optimizing pessimism in dynamic treatment regimes: A bayesian
learning approach, in: International Conference on Artificial Intelligence and Statistics, PMLR,
2023, pp. 6704–6721.
[27] B. Zhang, A. A. Tsiatis, M. Davidian, M. Zhang, E. Laber, Estimating optimal treatment regimes
from a classification perspective, Stat 1 (2012) 103–114.
[28] Y. Zhao, D. Zeng, A. J. Rush, M. R. Kosorok, Estimating individualized treatment rules using
outcome weighted learning, Journal of the American Statistical Association 107 (2012) 1106–1118.
[29] B. Zhang, A. A. Tsiatis, E. B. Laber, M. Davidian, A robust method for estimating optimal treatment
regimes, Biometrics 68 (2012) 1010–1018.
[30] B. Zhang, M. Zhang, C-learning: A new classification framework to estimate optimal dynamic
treatment regimes, Biometrics 74 (2018) 891–899.
[31] Y. Zhang, M. van der Schaar, Gradient regularized v-learning for dynamic treatment regimes,</p>
      <p>Advances in Neural Information Processing Systems 33 (2020) 2245–2256.
[32] Y. Tao, L. Wang, D. Almirall, Tree-based reinforcement learning for estimating optimal dynamic
treatment regimes, The Annals of Applied Statistics 12 (2018) 1914.
[33] Y. Sun, L. Wang, Stochastic tree search for estimating optimal dynamic treatment regimes, Journal
of the American Statistical Association 116 (2021) 421–432.
[34] T. Blumlein, J. Persson, S. Feuerriegel, Learning optimal dynamic treatment regimes using causal
tree methods in medicine, in: Machine Learning for Healthcare Conference, PMLR, 2022, pp.
146–171.
[35] J. Zhang, Designing optimal dynamic treatment regimes: A causal reinforcement learning approach,
in: International Conference on Machine Learning, PMLR, 2020, pp. 11012–11022.
[36] S. Chen, B. Zhang, Estimating and improving dynamic treatment regimes with a time-varying
instrumental variable, Journal of the Royal Statistical Society Series B: Statistical Methodology 85
(2023) 427–453.
[37] D. J. Lizotte, M. H. Bowling, S. A. Murphy, Eficient reinforcement learning with multiple reward
functions for randomized controlled trial analysis., in: ICML, volume 10, 2010, pp. 695–702.
[38] E. B. Laber, D. J. Lizotte, B. Ferguson, Set-valued dynamic treatment regimes for competing
outcomes, Biometrics 70 (2014) 53–61.
[39] Y.-Q. Zhao, R. Zhu, G. Chen, Y. Zheng, Constructing dynamic treatment regimes with shared
parameters for censored data, Statistics in Medicine 39 (2020) 1250–1263.
[40] A. Ertefaie, R. L. Strawderman, Constructing dynamic treatment regimes over indefinite time
horizons, Biometrika 105 (2018) 963–977.
[41] W. Zhou, R. Zhu, A. Qu, Estimating optimal infinite horizon dynamic treatment regimes via
pt-learning, Journal of the American Statistical Association 119 (2024) 625–638.
[42] X. Wang, H. Lee, B. Haaland, K. Kerrigan, S. Puri, W. Akerley, J. Shen, A matching-based machine
learning approach to estimating optimal dynamic treatment regimes with time-to-event outcomes,
Statistical Methods in Medical Research 33 (2024) 794–806.
[43] W. Hua, H. Mei, S. Zohar, M. Giral, Y. Xu, Personalized dynamic treatment regimes in continuous
time: a bayesian approach for optimizing clinical decisions with timing, Bayesian Analysis 17
(2022) 849–878.
[44] D. Niraula, W. Sun, J. Jin, I. D. Dinov, K. Cuneo, J. Jamaluddin, M. M. Matuszak, Y. Luo, T. S.</p>
      <p>Lawrence, S. Jolly, et al., A clinical decision support system for AI-assisted decision-making in
response-adaptive radiotherapy (ARCliDS), Scientific Reports 13 (2023) 5279.
[45] G. J. Klir, B. Yuan, Fuzzy Sets and Fuzzy Logic: Theory and Applications, Prentice Hall, Upper</p>
      <p>Saddle River, New Jersey, USA, 1995.
[46] J. Yen, R. Langari, Fuzzy Logic: Intelligence, Control, and Information, Prentice Hall, Upper Saddle</p>
      <p>River, New Jersey, USA, 1998.
[47] M. S. Nobile, G. Votta, R. Palorini, S. Spolaor, H. De Vitto, P. Cazzaniga, F. Ricciardiello, G. Mauri,
L. Alberghina, F. Chiaradonna, et al., Fuzzy modeling and global optimization to predict novel
therapeutic targets in cancer cells, Bioinformatics 36 (2020) 2181–2188.
[48] D. Chicco, S. Spolaor, M. S. Nobile, Ten quick tips for fuzzy logic modeling of biomedical systems,</p>
      <p>PLOS Computational Biology 19 (2023) e1011700.
[49] S. Spolaor, M. Scheve, M. Firat, P. Cazzaniga, D. Besozzi, M. S. Nobile, Screening for combination
cancer therapies with dynamic fuzzy modeling and multi-objective optimization, Frontiers in
Genetics 12 (2021) 617935.
[50] S. Spolaor, C. Fuchs, P. Cazzaniga, U. Kaymak, D. Besozzi, M. S. Nobile, Simpful: a user-friendly
Python library for fuzzy logic, International Journal of Computational Intelligence Systems 13
(2020) 1687–1698.
[51] M. Sugeno, Industrial Applications of Fuzzy Control, Elsevier Science Inc., New York, NY, 1985.
[52] R. Bellman, Dynamic programming, Science 153 (1966) 34–37.
[53] J. Clifton, E. Laber, Q-learning: Theory and applications, Annual Review of Statistics and Its</p>
      <p>Application 7 (2020) 279–301.
[54] D. Hanahan, Hallmarks of Cancer: New Dimensions, Cancer Discovery 12 (2022) 31–46.
[55] A. D. Cox, S. W. Fesik, A. C. Kimmelman, J. Luo, C. J. Der, Drugging the undruggable RAS: Mission
possible?, Nature Reviews Drug Discovery 13 (2014) 828–851.
[56] S. R. Punekar, V. Velcheti, B. G. Neel, K.-K. Wong, The current state of the art and future trends in</p>
      <p>RAS-targeted cancer therapies, Nature Reviews Clinical Oncology 19 (2022) 637–655.
[57] L. Breiman, Random forests, Machine learning 45 (2001) 5–32.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Kosorok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Laber</surname>
          </string-name>
          , Precision medicine,
          <source>Annu Rev Stat Appl</source>
          <volume>6</volume>
          (
          <year>2019</year>
          )
          <fpage>263</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bzdok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Steyerberg</surname>
          </string-name>
          ,
          <article-title>Prediction, not association, paves the road to precision medicine</article-title>
          ,
          <source>JAMA Psychiatry 78(2)</source>
          (
          <year>2021</year>
          )
          <fpage>127</fpage>
          -
          <lpage>128</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <article-title>Traversing the conceptual divide between biological and statistical epistasis: systems biology and a more modern synthesis</article-title>
          ,
          <source>Bioessays</source>
          <volume>27</volume>
          (
          <issue>6</issue>
          ) (
          <year>2005</year>
          )
          <fpage>637</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Eldar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Elowitz</surname>
          </string-name>
          ,
          <article-title>Functional roles for noise in genetic circuits</article-title>
          ,
          <source>Nature</source>
          <volume>467</volume>
          (
          <year>2010</year>
          )
          <fpage>167</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kitano</surname>
          </string-name>
          , Biological robustness,
          <source>Nature Reviews Genetics</source>
          <volume>5</volume>
          (
          <year>2004</year>
          )
          <fpage>826</fpage>
          -
          <lpage>837</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sejdić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Lipsitz</surname>
          </string-name>
          ,
          <article-title>Necessity of noise in physiology and medicine</article-title>
          ,
          <source>Computer Methods and Programs in Biomedicine</source>
          <volume>111</volume>
          (
          <year>2013</year>
          )
          <fpage>459</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Melograna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Galazzo</surname>
          </string-name>
          , N. van
          <string-name>
            <surname>Best</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Mommers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Penders</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Stella</surname>
            ,
            <given-names>K. Van Steen</given-names>
          </string-name>
          ,
          <article-title>Edge and modular significance assessment in individual-specific networks</article-title>
          .,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P. W.</given-names>
            <surname>Lavori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dawson</surname>
          </string-name>
          ,
          <article-title>A design for testing clinical strategies: biased adaptive within-subject randomization</article-title>
          ,
          <source>Journal of the Royal Statistical Society: Series A (Statistics in Society)</source>
          <volume>163</volume>
          (
          <year>2000</year>
          )
          <fpage>29</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <article-title>Optimal dynamic treatment regimes</article-title>
          ,
          <source>Journal of the Royal Statistical Society Series B: Statistical Methodology</source>
          <volume>65</volume>
          (
          <year>2003</year>
          )
          <fpage>331</fpage>
          -
          <lpage>355</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J. M. Robins</surname>
          </string-name>
          ,
          <article-title>Optimal structural nested models for optimal sequential decisions</article-title>
          ,
          <source>in: Proceedings of the Second Seattle Symposium in Biostatistics: analysis of correlated data</source>
          , Springer,
          <year>2004</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>326</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Watkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dayan</surname>
          </string-name>
          ,
          <article-title>Q-learning</article-title>
          ,
          <source>Machine learning 8</source>
          (
          <year>1992</year>
          )
          <fpage>279</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>