<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Improving the Prediction of Individual Engagement in Recommendations using Cognitive Models⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roderick Seow</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yunfan Zhao</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duncan Wood</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milind Tambe</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cleotilde Gonzalez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Mellon University</institution>
          ,
          <addr-line>5000 Forbes Ave, Pittsburgh, Pennsylvania</addr-line>
          ,
          <country country="US">United States of America</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Harvard University</institution>
          ,
          <addr-line>Massachusetts Hall, Cambridge, Massachusetts</addr-line>
          ,
          <country country="US">United States of America</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>For public health programs with limited resources, the ability to predict how behaviors change over time and in response to interventions is crucial for deciding when and to whom interventions should be allocated. Using data from a real-world maternal health program, we demonstrate how a cognitive model based on InstanceBased Learning (IBL) Theory can augment existing purely computational approaches. Our findings show that, compared to general time-series forecasters (e.g., LSTMs), IBL models, which reflect human decision-making processes, better predict how individuals' behaviors change over time (transition-consistency) and in response to receiving an intervention (intervention-sensitivity). We further show that IBL parameters capture the individual diferences in transition-consistency and intervention-sensitivity and that other time series models can use these individual-level IBL parameters to improve their training eficiency.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Limited resource allocation</kwd>
        <kwd>Restless multi-armed bandit</kwd>
        <kwd>Time-series forecasting</kwd>
        <kwd>Cognitive modeling</kwd>
        <kwd>Instancebased learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Public health programs play an essential role in improving the health outcomes of individuals and
communities, often through education and subsequent behavioral change. Some health programs
interact with their intended beneficiaries in a broad and infrequent manner. For example, a campaign
about the health risks of smoking may address a general population of smokers through scattered
advertisements in the media [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Others rely on repeated direct interactions with their intended
beneficiaries—for example, a maternal health program that sends automated messages about exercise
and nutrition to enrolled expectant mothers [2]. In this case, it is crucial that mothers remain engaged for
the duration of the program or as long as possible to receive the maximum benefit. Unfortunately, many
such programs face high levels of disengagement and dropout, which severely limit their efectiveness
[3].
      </p>
      <p>To reduce dropout, programs can provide interventions designed to increase beneficiary engagement.
For example, program staf can make personalized service calls to beneficiaries at risk of dropping out
to address concerns about declining participation. However, programs with limited staf and resources
cannot provide interventions to all their beneficiaries. Instead, they must choose a subset of beneficiaries
to receive the intervention.</p>
      <p>Recent approaches have adopted the Restless Multi-Armed Bandit (RMAB) framework [4]. Classically,
the RMAB framework models the engagement dynamics of each beneficiary (i.e., the transition between
discrete engagement and disengagement states) as a Markov Decision Process (MDP). A centralized
planner then recommends which beneficiaries to intervene on at each time point based on the dynamics
learned from the MDPs. Importantly, beneficiaries can change their engagement state even without an
intervention. The best-performing approach to this framework is the Whittle index [5], which ranks
beneficiaries by their likelihood of becoming and remaining engaged after getting an intervention.</p>
      <p>This framework assumes that beneficiaries’ dynamics—how their level of engagement changes over
time—can be modeled as a MDP. However, recent work has shown that these dynamics are often
non-Markovian; that is, a beneficiary’s transition between being engaged and disengaged (their state)
depends on their individual history [6]. As a substitute, Danassis and colleagues (2023) developed
the Time-series Arm Ranking Index (TARI), which allows non-Markovian dynamics and continuous
engagement levels within the RMAB framework. The TARI approach models beneficiaries using
timeseries forecasters such as LSTMs [7] or transformers [8] instead of MDPs. Using these learned time-series
models, TARI computes an index that captures the relative engagement benefit of intervening on a
beneficiary by comparing their expected future engagement with and without an intervention at the
current time point and then ranks beneficiaries by the computed indices.</p>
      <p>Within the TARI approach to the RMAB framework, we propose modeling individual beneficiaries
with computational cognitive models instead of LSTMs. Specifically, we use models based on
InstanceBased Learning Theory (IBLT) [9] to represent each beneficiary’s time series activities. We expect using
an IBL model will convey two advantages over general time-series forecasters (TSFs).</p>
      <p>First, IBL models are personalized to the individual instead of trained across the entire dataset,
allowing the models to capture individual diferences in behavioral dynamics (i.e., individual-level
diferences in the response to interventions and general engagement patterns). Because TSFs need
large amounts of data to fine-tune their many parameters and there is relatively sparse data available
per beneficiary, existing approaches train a single model using the combined data from all available
beneficiaries. For instance, Danassis and colleagues (2023) create a set of fixed-length vectors from
each beneficiary’s trajectory. Each vector represents the intervention history and engagement levels of
seven consecutive timesteps. Then an LSTM trains on these vectors to predict the engagement level
at the next timestep given the previous seven [6]. In contrast, an IBL model has enough cognitively
grounded structure that it can be applied to an individual with sparse data [10]. This is because IBL
theory makes strong assumptions about the cognition of human decision making derived from prior
research. In particular, IBL theory accounts for memory efects (i.e., the more recent a data point is, the
more it contributes to the prediction) and similarity efects (i.e., engagement levels that are observed
under similar contexts should be similar to each other) that general TSFs do not.</p>
      <p>Second, the parameters of IBL cognitive models are grounded in psychological constructs such as
memory and attention, making them more interpretable than the parameters of general LSTMs. This
advantage allows cognitive models to shed light on the individual diferences that drive beneficiaries’
behaviors. We demonstrate how these individual diferences can be used to augment existing approaches
that rely primarily on general LSTMs.
1.1. Our Contributions
1. Using real-world data from a maternal health program, we show that cognitive models more
accurately capture the temporal dynamics of individual behaviors and predict the behaviors of
individual beneficiaries than LSTMs.
2. We show that personalized cognitive models reveal individual characteristics through weights
iftting. We illustrate how to cluster individuals using learned characteristics.
3. We show clustering based on learned individual characteristics could guide and improve the
performance of time series models. Specifically, we show that LSTMs trained in clusters better
predict behaviors within those clusters compared to LSTMs trained on random samples of
beneficiaries.</p>
      <sec id="sec-1-1">
        <title>1.2. Background</title>
        <sec id="sec-1-1-1">
          <title>1.2.1. Time-series Arm Ranking Index (TARI)</title>
          <p>In the context of multiple potential beneficiaries treated as arms in an RMAB, each arm is considered
an independent time series [6]. We train a model to predict the next state +1 based on three pieces of
information:
• A historical record (length h) of the past states and actions of the arm (denoted by {(, )}ℎ&lt;).
• The current state (e.g., engagement level) of the arm .</p>
          <p>• A potential action , where 1 represents an intervention and 0 represents no intervention.
This training takes place ofline. During test time, the model’s parameters are fixed, and we use the
trained model with iterated multi-step forecasting to generate a long-term forecast of future states
+1, +2, ..., + [11].</p>
          <p>For each arm, we use the TSF model to estimate two values:
• Time to disengagement with intervention : The predicted number of timesteps until the arm
becomes non-engaged if we intervene at the current timestep (t) and never again.
• Time to disengagement without intervention : The predicted number of timesteps until the
arm becomes non-engaged with no interventions at any point.</p>
          <p>The TARI index for an arm is the ratio of these two times. A higher TARI index suggests that intervening
now would be more beneficial. Finally, similar to the Whittle index, we choose to act on the  arm with
the highest TARI index at each timestep:</p>
          <p>TARI() = 

.</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>1.2.2. Instance-Based Learning Theory</title>
          <p>Instance-Based Learning Theory (IBLT) is a cognitive theory of dynamic decision making, grounded
in human learning and memory mechanisms. It has successfully accounted for human behavior in a
variety of contexts, ranging from abstract repeated-choice tasks to more real-world search and choice
[12] and continuous control tasks [9]. In most of these contexts, human participants are required to
make a series of decisions while exposed to changing environments.</p>
          <p>The essence of IBLT is storing past decisions in the form of an instance. An instance has three parts:
the context, the choice made, and the associated utility. For example, these would correspond to the
cloudiness of the sky, the decision to bring an umbrella, and the inconvenience of bringing an umbrella
when it does not rain (supposing it did not rain in this instance).</p>
          <p>IBL agents compute the expected utility of a choice by considering past instances where they made
that choice, weighted by the contextual similarity to the present. For example, an IBL agent would
compute the expected utility of bringing an umbrella by considering its past experiences with an
umbrella, most strongly considering the experiences with the most similar cloudiness to the current
moment.</p>
          <p>More precisely, the expected value of a choice is the weighted average of the utilities from all the
past instances where that choice was made weighted by the retrieval probability (i.e., the probability
of remembering that instance) to generate estimates of the choice’s expected utility called a blended
value. The agent will deterministically make the available choice with the highest computed blended
value (expected utility). The utilities of instances in memory are the actual outcome the agent observed
when it chooses, regardless of what it predicted. After choosing, a new instance is added to the agent’s
memory, reflecting what happened. If the exact combination of context, choice, and utility has appeared
before, that instance instead has its frequency and recency information updated.</p>
          <p>The contribution of an instance’s utility to a choice’s expected value depends on the instance’s memory
activation (the salience of that instance given the current context). The activation of an instance ()
reflects how readily it comes to mind, which in turn depends on the frequency, recency, and similarity
factors according to the following equation as proposed by the ACT-R cognitive architecture [13].
 = (∑︁( − ′, )− ) +  ∑︁(((,, ,) − 1) + 
 
(1)</p>
          <p>Three expressions additively determine an instance’s activation (salience in memory). The first
expression is the contribution of frequency and recency where  is the current timestep, , is the
timestep of the th appearance of instance , and  is a decay parameter. The second expression is
the contribution of similarity where  is a scaling parameter on the overall impact of similarity to
activation,  is the importance (or weight) of an attribute in the similarity calculations,  is a
similarity function that returns the degree of similarity for a particular attribute  between instance 
(,) and the current context (,). The third expression adds some randomness to the memory retrieval
process where  is a scaled noise distribution. The activation of an instance is then converted into a
retrieval probability using the Boltzmann softmax function, where  is the temperature parameter:
/
 = ∑︀ /
 () = ∑︁</p>
          <p />
          <p>The IBL model computes the expected utility of a choice () from the utility  of each retrieved
instance  weighted by their respective retrieval probability  [9, 14]:</p>
          <p>The agent then makes the choice with the highest expected utility,  ().</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Proposed Approach</title>
      <p>We model each beneficiary with their own individual-level IBL model. We do not have the IBL model
make a choice. Instead, we use the IBL’s expected utility calculation to predict the engagement level
in the next period (see Figure 1a). Because we only use the model’s blended value, the choice part of
the instance is always held constant. The utility of an instance is a beneficiary’s engagement level
(thus, the expected utilities are weighted averages of past engagement levels then used to predict future
engagement levels). The context has two components derived from a beneficiary’s history: (1) the
engagement level of the previous period and (2) the number of timesteps since the last intervention. The
IBL model determines the similarity between the current context and each instance in memory according
to these two attributes, and these similarities are factored into an instance’s activation according to
Equation 1.</p>
      <p>The IBL models are trained through model tracing: recording each timestep of an individual
beneficiary as an instance with the relevant context (previous engagement level and time since the last
intervention) and utility (engagement level). Thus, the model “traces” the beneficiary it is modeling by
giving itself memories as if it were the beneficiary that produced the data.</p>
      <p>As shown in Equation 1, the similarity of two instances depends not only on the similarity between
attribute values, but also on the weight or influence of the attribute (  in Equation 1). Psychologically, an
attribute weight represents how relevant the agent considers an attribute to make similarity judgments
between two contexts. To capture potential individual diferences between beneficiaries, we personalized
an IBL model for a specific beneficiary by finding the combination of attribute weight values that resulted
in the smallest training prediction error.</p>
      <p>Given the predictions of engagement level, we follow TARI and generate a ranked list of beneficiaries
according to the predicted relative benefit in the improvement in engagement each beneficiary receives
from an intervention versus not receiving one.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Simulation Setup</title>
      <p>3.1. Data
We tested our approach on data from a maternal healthcare program operated by an NGO called
ARMMAN [15]. This program sends automated messages about maternal and infant care to expectant
and recent mothers in vulnerable communities. However, over the many weeks of pregnancy and
postnatal care, many mothers stop engaging with the automated messages. ARMMAN has a limited
(2)
(3)
capacity to intervene: Program staf can call mothers individually, allowing them to hear from a real
person and ask questions, hopefully increasing their future engagement.</p>
      <p>
        ARMMAN collected the data in 2022 from 12,000 mothers over 40 weeks. The data have two recorded
values per mother per week: the amount of time she spent listening to the automated health message
(engagement level, recorded as a number in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]) and if she received a call from ARMMAN staf (if
she received an intervention, recorded as a binary value). The data also contain mothers’ demographic
information; however, previous work found little or no prediction benefit in incorporating demographic
variables as additional features for time-series forecasters [6]. Thus, we do not consider them in this
approach.
      </p>
      <p>For an IBL model to predict an intervention’s efect on a mother’s engagement level, it needs to be
trained on a trajectory that includes at least one intervention. However, of the 12,000 mothers, only
5,400 mothers received at least one intervention early on in their program tenure. We also assume that
mothers receive an intervention (or the equivalent of an intervention) upon their enrollment in the
program. Because the time needed to train all the IBL models scales with the number of mothers, we
explored our approach with a subset of 210 mothers out of the 5,400 mothers who received at least one
intervention in their actual trajectories.</p>
      <p>=t3
=t0
=t2
=t1
a
b
SLTM</p>
      <p>BIL
7.s=0
9.s=0
7.s=0
6.s=0
=a0
=a0
=a1
.</p>
      <p>=t2
=t1
ts-7
BIL</p>
      <p>SLTM
aringT
ta-7
negamt
no
setingT
negamt
no
the Within and Outside methods.</p>
      <sec id="sec-3-1">
        <title>3.2. Model Training</title>
        <p>We designated the first 25 weeks of each mother’s trajectory as training data. Following the approach
of previous studies, we reconstructed the training data into sliding windows of 7 consecutive timesteps
to train the LSTM model. Then, we trained an LSTM model on the entire training dataset (see Figure
1b for a visual description of the diferences between the IBL and LSTM training / testing procedure).</p>
        <p>Each IBL model (one for each mother) traced a mother’s trajectory through each timestep up to
week 25. To find the best-fitting profile of attribute weights per mother, a grid search looked for the
combination that minimized the weighted sum of next-step prediction errors (loss). The grid iterates all
possible combinations of parameter values in the range (0, 0) to (5, 5) in intervals of 0.5.</p>
        <p>An IBL model should improve its predictive accuracy as it accumulates more instances and becomes
more “familiar” with the individual it’s modeling. So, we weight earlier prediction errors (e.g., at weeks
1 and 2) less than later ones (e.g., at weeks 24 and 25) according to the following equations:
 = ∑︁  *</p>
        <p>/10
 = ∑︀ /10
(4)
(5)
where  is the squared prediction error and  is the normalized weight factor for the timestep .</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Next-step prediction task</title>
        <p>For the next-step prediction task, the IBL and LSTM models generated predictions for the next 14 weeks,
starting from week 26. The models iteratively generated predictions for the next timestep (i.e., for week
26, then for week 27, then for week 28, etc.). In our data, ARMANN did not provide interventions
to mothers past 14 weeks, so there are no interventions during the testing period for the next-step
prediction task. This allows us to compare the models’ engagement predictions with the ground truth
data without the need to generate synthetic counterfactuals.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Model comparison with simulated counterfactuals</title>
        <p>In a separate analysis, we compared the IBL-TARI and LSTM-TARI policies with three other baseline
policies: (i) myopic uniformly random allocation (ii) round-robin allocation, which iteratively cycles
through the population, and (iii) no interventions. Consistent with ARMMAN’s limited resources, all
policies are allowed a 3% budget within a timestep (except the no intervention policy), i.e., 6 beneficiaries
per timestep.</p>
        <p>To simulate the efect of these policies’, we followed previous studies and trained a separate LSTM
model on trajectories from the entire dataset (approximately 12000 beneficiaries)—far larger than the
policy LSTM’s training set—to generate mothers’ counterfactual engagement. This counterfactual
generator only operates when a mother’s simulated trajectory deviates from her ground-truth trajectory
(i.e., received a simulated intervention when she did not receive one in reality or not receiving a
simulated intervention when she did receive one in reality).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Predicting next-step engagement levels</title>
        <p>Figure 2 shows the prediction error of the two types of models during the 14 week the testing period of
the next-step prediction task. The IBL models consistently achieve lower prediction errors than the
LSTM—about 10% less mean error (0.23 on average) compared to the LSTM model (0.32 on average).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Using predictions to inform intervention allocation</title>
        <p>Having demonstrated that IBL models yield more accurate next-step predictions of engagement levels
compared to the LSTM model, we employed simulations to examine if integrating IBL models with
the TARI index would result in higher overall levels of engagement. Figure 7 (in Appendix) shows
the percentage of mothers who have an engagement level of 0.25 and above during each test timestep
following the various intervention allocation policies. The LSTM-TARI policy outperforms all other
policies, particularly later on in the testing period. The average percentage of engaged mothers across
testing timesteps for the various policies are as follows: IBL-TARI (55.58%); Round-robin (55.65%);
Random (54.29%); LSTM-TARI (61.05%); Control (46.22%). Generating counterfactuals using an
LSTM model may explain why the IBL’s advantage in next-step prediction does not translate to higher
counterfactual engagement in the simulated task.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Describing individual diferences with IBL parameters</title>
        <p>Another advantage of the proposed IBL approach over more general TSFs (as exemplified by the LSTM
approach) is identifying individual diferences in beneficiaries’ behavioral characteristics. Because
we model each beneficiary with a separate IBL model, the models’ parameter values capture how
beneficiaries difer from each other in their engagement and intervention dynamics.</p>
        <p>Specifically, the attribute weights of each mother’s IBL model trained on data from the first 25 weeks
tracks (1) how reliably her engagement changes in response to an intervention (intervention-sensitivity)
and (2) how consistently her engagement changes between timesteps (transition-consistency), i.e., how
similarity between the engagement levels in diferent timesteps corresponds to similarity between the
engagement levels in their following timesteps.</p>
        <p>Figure 3 plots mothers according to their attributes weights. Visually, there appear to be three
types of mothers. We confirmed that the optimal number of clusters is between 3 to 4 clusters using
multiple measures of cluster validity indices (silhouette: 4, within-sum-of-squares: 4, gap statistic: 3).
For interpretability, subsequent analyses assume three clusters.</p>
        <p>Each cluster was labeled according to its defining characteristics. Intervention-sensitive mothers ( 
= 62) have a high intervention lag weight; i.e., a mother’s engagement level  timesteps after her most
recent intervention will be similar to other times she was  timesteps after the most recent intervention.
Transition-consistent mothers ( = 66) have a high previous engagement weight; i.e., if engagement
levels at timesteps  − 1 and  − 1 (− 1 and − 1) are similar to each other, then we should expect
engagement levels  and  to also be similar to each other. State-stable mothers ( = 82) have
low values in both dimensions; i.e., these mothers will tend to engage at the level they engage most
frequently and most recently.</p>
        <sec id="sec-4-3-1">
          <title>4.3.1. Testing the generalizability of individual diferences in IBL attribute weights</title>
          <p>The previous results showed that individual mothers have diferent optimal weight profiles for the IBL
instance similarity attributes. However, we still must show these weights capture individual diferences
that predict behavioral dynamics beyond an IBL modeling context. If these attribute weights reflect
underlying individual diferences in engagement and intervention dynamics, then we would expect
other modeling approaches that account for these attribute weights to perform better than those that
do not.</p>
          <p>We tested this hypothesis by training and testing LSTM models with cluster-related samples. Using
the cluster assignments described in the previous section, we randomly split each cluster in half to
construct a training and a testing set. For all methods, the task is to predict the next-state engagement
of the testing set during the testing period (i.e., the final 14 weeks). We examined the following methods
(see 1c for a visual description):
• Entire: An LSTM trained on all training mothers and predicting for all test mothers.
• Within-cluster: An LSTM trained on each cluster and predicting for test mothers in their
respective training cluster. We hypothesize that this method should yield the most accurate
predictions given the informativeness of the identified IBL clusters.
• Outside-cluster: An LSTM trained on each cluster and predicting for test mothers from a
diferent cluster. Because there are 3 clusters, we average prediction error from running the
model on both out-of-cluster testing sets.
• Random: An LSTM trained on a random subset of one-third of the training set ( = 35) and
predicting for all test mothers</p>
          <p>Figure 4 shows the average prediction error for the various methods. Supporting the generalizability
of the IBL attribute weight clusters, LSTMs trained and tested on diferent clusters (Outside-cluster)
and the LSTM trained on a random sample (Random) performed worse than LSTMs trained and tested
on the same clusters (Within).</p>
          <p>The Within-cluster and Entire methods yield similar average prediction accuracies (despite each
Within-cluster LSTM having less training data). However, as shown in Figure 5, the Within-cluster
method produces more accurate predictions than the Entire method for two-thirds of the mothers
(Within 70 vs. Entire: 35), which suggests that the clusters efectively categorize the majority of the
population.</p>
          <p>Follow-up exploratory analyses suggest that mothers closer to their cluster’s centroid are better
predicted by the Within-cluster method as compared to the Entire method. Using a distinct binomial
regression model for each cluster, we used the Euclidean distance between a mother’s attribute weights
and her centroid to predict the method (Within-cluster or Entire) with lower prediction error. The efects
were statistically insignificant, plausibly due to the imbalanced and small classes. Still, for State-stable
(  = − 0.8637) and Transition-consistent mothers (  = − 1.235), the closer they were to the
cluster centroid, the more likely the Within-cluster method gave better predictions than the Entire
method (see Figure 6).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future work</title>
      <p>Public health programs with limited resources to provide interventions must face the dificult problem
of determining which beneficiaries to give interventions. Prior work demonstrated that the TARI
algorithm—which relies on the predictions of time-series forecasters—outperforms existing
recommendation algorithms [6]. To improve the efectiveness of TARI’s recommendations through more
accurate predictions, we developed a novel approach that uses computational cognitive models derived
from Instance-Based Learning Theory [9, 14] to model the human transition dynamics of program
engagement. Given the non-Markovian nature of a beneficiary’s transitions between states, a
cognitive algorithm that reflects the efects of time, attention, and context similarity on human memory
and decisions provides more accurate predictions of beneficiaries’ engagement dynamics and more
meaningfully describes relevant individual diferences that other prediction methods can leverage.</p>
      <p>We tested our method on real-world engagement data from mothers enrolled in a maternal health
program, finding it results in higher next-step prediction accuracy compared to existing time-series
methods using LSTM models. The personalized IBL models revealed that mothers clustered into three
types based on their IBL attribute weights— specifically, how their engagement patterns could be
attributed to intervention-sensitivity and transition-consistency. These clusters generalized to other
time-series models: LSTMs trained and tested on within-cluster data outperformed all other LSTM
training methods for most mothers.</p>
      <p>Future work could explore recommending personalized interventions and intervention schedules
based on these identified individual diferences in behavioral characteristics. For example,
interventionsensitive beneficiaries could receive more frequent but shorter calls as the model predicts that each call
would provide a substantial boost to their engagement levels. In comparison, state-stable beneficiaries
who are likely to maintain their engagement levels over time could receive fewer but more intensive
calls that maximize their engagement levels by the end of each call. These beneficiaries would then be
able to sustain a high engagement level until the next time they receive a call.</p>
      <p>One limitation of this work is the counterfactual generator used to simulate the efect of recommended
interventions—an LSTM model trained on a much larger dataset. The higher next-step prediction
accuracy of the IBL model suggests that the LSTM architecture may systematically fail to pick up certain
cognitively-determined regularities in the human data. Thus, the LSTM-generated counterfactuals may
also fail to produce simulated data with those regularities. This might explain why IBL’s next-step
prediction advantage over the LSTM does not translate to better performance in the model comparison
with simulated counterfactuals. Future work should prioritize testing our method with other datasets,
particularly with a paradigm where the efects of recommended interventions on participants’ task
engagement can be experimentally measured, rather than relying on plausibly biased counterfactuals.</p>
      <p>Another limitation is that our method uses the first 25 weeks of the program for model training, only
making predictions from the 26th week of a mother’s pregnancy onward. Possible extensions of this
work could explore hybrid methods, such as initially relying on predictions from other trained IBL
models and gradually transitioning to the fully personalized model.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
      <p>This research was supported by the AI Research Institutes Program funded by the National Science
Foundation under AI Institute for Societal Decision Making (AI-SDM), Award No. 2229881
[2] A. Lalan, S. Verma, P. R. Diaz, P. Danassis, A. Mahale, K. M. Sudan, A. Hegde, M. Tambe, A. Taneja,
Improving health information access in the world’s largest maternal mobile health program via
bandit algorithms, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38,
2024, pp. 22913–22919. doi:https://doi.org/10.1609/aaai.v38i21.30329.
[3] S. Amagai, S. Pila, A. J. Kaat, C. J. Nowinski, R. C. Gershon, Challenges in participant engagement
and retention using mobile health apps: literature review, Journal of medical Internet research 24
(2022) e35120. doi:10.2196/35120.
[4] J. A. Killian, A. Biswas, L. Xu, S. Verma, V. Nair, A. Taneja, A. Hegde, N. Madhiwalla, P. R.</p>
      <p>Diaz, S. Johnson-Yu, et al., Robust planning over restless groups: engagement interventions
for a large-scale maternal telehealth program, in: Proceedings of the AAAI Conference on
Artificial Intelligence, volume 37, 2023, pp. 14295–14303. doi: https://doi.org/10.1609/aaai.
v37i12.26672.
[5] P. Whittle, Restless bandits: Activity allocation in a changing world, Journal of applied probability
25 (1988) 287–298. doi:10.2307/3214163.
[6] P. Danassis, S. Verma, J. A. Killian, A. Taneja, M. Tambe, Limited resource allocation in a
nonmarkovian world: the case of maternal and child healthcare, arXiv preprint arXiv:2305.12640
(2023). doi:https://doi.org/10.48550/arXiv.2305.12640.
[7] S. Hochreiter, J. Schmidhuber, Long short-term memory, Neural computation 9 (1997) 1735–1780.
[8] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin,</p>
      <p>Attention is all you need, Advances in neural information processing systems 30 (2017).
[9] C. Gonzalez, J. F. Lerch, C. Lebiere, Instance-based learning in dynamic decision making, Cognitive</p>
      <p>Science 27 (2003) 591–635. doi:https://doi.org/10.1207/s15516709cog2704_2.
[10] E. H. Bugbee, C. Gonzalez, Making predictions without data: How an instance-based learning
model predicts sequential decisions in the balloon analog risk task, in: Proceedings of the annual
meeting of the cognitive science society, volume 44, 2022.
[11] S. B. Taieb, R. J. Hyndman, et al., Recursive and direct multi-step forecasting: the best of both
worlds, volume 19, Department of Econometrics and Business Statistics, Monash Univ., 2012.
[12] T. N. Nguyen, C. Gonzalez, Theory of mind from observation in cognitive models and humans,</p>
      <p>Topics in cognitive science 14 (2022) 665–686. doi:https://doi.org/10.1111/tops.12553.
[13] J. R. Anderson, C. J. Lebiere, The atomic components of thought, Psychology Press, 2014. doi:https:
//doi.org/10.4324/9781315805696.
[14] C. Gonzalez, V. Dutt, Instance-based learning: integrating sampling and repeated decisions from
experience., Psychological review 118 (2011) 523. doi:https://doi.org/10.1037/a0024558.
[15] ARMMAN, mmitra - armman - helping mothers and children, 2024. URL: https://armman.org/
mmitra/.
[16] A. Mate, L. Madaan, A. Taneja, N. Madhiwalla, S. Verma, G. Singh, A. Hegde, P. Varakantham,
M. Tambe, Field study in deploying restless multi-armed bandits: Assisting non-profits in improving
maternal and child health, in: Proceedings of the AAAI Conference on Artificial Intelligence,
volume 36, 2022, pp. 12017–12025. doi:https://doi.org/10.1609/aaai.v36i11.21460.
[17] Y. Zhao, T. Wang, D. Nagaraj, A. Taneja, M. Tambe, The bandit whisperer: Communication
learning for restless bandits, arXiv preprint arXiv:2408.05686 (2024). doi:https://doi.org/10.
48550/arXiv.2408.05686.
[18] G. Xiong, J. Li, Finite-time analysis of whittle index based q-learning for restless multi-armed
bandits with neural network function approximation, in: Thirty-seventh Conference on Neural
Information Processing Systems, 2023.
[19] N. Jaques, A. Lazaridou, E. Hughes, C. Gulcehre, P. Ortega, D. Strouse, J. Z. Leibo, N. De Freitas,
Social influence as intrinsic motivation for multi-agent deep reinforcement learning, in: International
conference on machine learning, PMLR, 2019, pp. 3040–3049.
[20] N. Modi, P. Mary, C. Moy, Transfer restless multi-armed bandit policy for energy-eficient
heterogeneous cellular network, EURASIP Journal on Advances in Signal Processing 2019 (2019)
1–19. doi:https://doi.org/10.1186/s13634-019-0637-1.
[21] Y. Zhao, N. Behari, E. Hughes, E. Zhang, D. Nagaraj, K. Tuyls, A. Taneja, M. Tambe, Towards</p>
    </sec>
    <sec id="sec-7">
      <title>A. Real-world ARMMAN Dataset</title>
      <sec id="sec-7-1">
        <title>A.1. Secondary Analysis</title>
        <p>Our experiment falls into the category of secondary analysis of the data shared by ARMMAN. This
paper does not involve deploying the proposed algorithm or any other baselines to the service call
program. As noted earlier, the experiments are secondary analyses with approval from the ARMMAN
ethics board.</p>
      </sec>
      <sec id="sec-7-2">
        <title>A.2. Consent and Data Usage</title>
        <p>Consent is obtained from every beneficiary enrolling in the NGO’s mobile health program. The NGO
owns the data collected through the program and only the NGO is allowed to share data. In our
experiments, we use anonymized call listenership logs to calculate empirical transition probabilities. No
personally identifiable information (PII) is available to us. The data exchange and usage were regulated by
clearly defined exchange protocols, including anonymization, read-access only to researchers, restricted
use of the data for research purposes only, and approval by ARMMAN’s ethics review committee.</p>
      </sec>
      <sec id="sec-7-3">
        <title>A.3. Interventions in ARMMAN data</title>
        <p>
          The interventions in ARMMAN data are chosen based on restless multi-arm bandit (RMAB) algorithms
[16, 17]. RMABs are a model for sequentially distributing scarce resources to a set of agents [5, 18, 19].
Concretely, we have a set of arms and a limited budget and face the question of deciding which arms
to pull in each round. The state of arms evolves according to a Markov Decision Process where
transition probabilities depend on whether the arm is pulled in this step. RMABs have a broad range of
applications, including resource allocation in anti-poaching, machine maintenance, cellular networks
[20, 21]. RMABs have had extensive use in healthcare settings such as call scheduling in a maternal
and child care program [
          <xref ref-type="bibr" rid="ref2 ref3">22, 23</xref>
          ], screening patients at risk of cancer [
          <xref ref-type="bibr" rid="ref4">24</xref>
          ], and allocating hepatitis C
treatment [
          <xref ref-type="bibr" rid="ref5">25</xref>
          ].
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. O.</given-names>
            <surname>Nogueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McNeill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. N.</given-names>
            <surname>Kyriakos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Mons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fernández</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Zatoński</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Trofor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Demjen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tountas</surname>
          </string-name>
          , et al.,
          <article-title>Impact of anti-smoking advertising on health-risk knowledge and quit attempts across 6 european countries from the eurest-plus itc europe survey</article-title>
          ,
          <source>Tobacco Induced Diseases</source>
          <volume>16</volume>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .18332/tid/96251.
          <article-title>zero shot learning in restless multi-armed bandits</article-title>
          ,
          <source>AAMAS</source>
          (
          <year>2023</year>
          ). doi:https://doi.org/10. 48550/arXiv.2310.14526.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>N.</given-names>
            <surname>Behari</surname>
          </string-name>
          , E. Zhang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Taneja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nagaraj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tambe</surname>
          </string-name>
          ,
          <article-title>A decision-language model (dlm) for dynamic restless multi-armed bandit tasks in public health</article-title>
          ,
          <source>arXiv preprint arXiv:2402.14807</source>
          (
          <year>2024</year>
          ). doi:https://doi.org/10.48550/arXiv.2402.14807.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Boehmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Taneja</surname>
          </string-name>
          , M. Tambe,
          <article-title>Group fairness in predict-thenoptimize settings for restless bandits</article-title>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>E.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Lavieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Volk</surname>
          </string-name>
          ,
          <article-title>Optimal screening for hepatocellular carcinoma: A restless bandit model</article-title>
          ,
          <source>Manufacturing &amp; Service Operations Management</source>
          <volume>21</volume>
          (
          <year>2019</year>
          )
          <fpage>198</fpage>
          -
          <lpage>212</lpage>
          . doi:https://doi. org/10.1287/msom.
          <year>2017</year>
          .
          <volume>0697</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ayer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifonte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Spaulding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chhatwal</surname>
          </string-name>
          ,
          <article-title>Prioritizing hepatitis c treatment in us prisons</article-title>
          ,
          <source>Operations Research</source>
          <volume>67</volume>
          (
          <year>2019</year>
          )
          <fpage>853</fpage>
          -
          <lpage>873</lpage>
          . doi:https://doi.org/10.1287/opre.
          <year>2018</year>
          .
          <year>1812</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>