<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applying Psychology of Persuasion to Conversational Agents through Reinforcement Learning: an Exploratory Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francesca Di Massimo</string-name>
          <email>francesca.dimassimo01@universitadipavia.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Carfora</string-name>
          <email>valentina.carfora@unicatt.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrizia Catellani</string-name>
          <email>patrizia.catellani@unicatt.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Piastra</string-name>
          <email>marco.piastra@unipv.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Vision and Multimedia Lab, Università degli Studi di Pavia</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dipartimento di Psicologia, Università Cattolica di Milano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This study is set in the framework of taskoriented conversational agents in which dialogue management is obtained via Reinforcement Learning. The aim is to explore the possibility to overcome the typical end-to-end training approach through the integration of a quantitative model developed in the field of persuasion psychology. Such integration is expected to accelerate the training phase and improve the quality of the dialogue obtained. In this way, the resulting agent would take advantage of some subtle psychological aspects of the interaction that would be difficult to elicit via end-to-end training. We propose a theoretical architecture in which the psychological model above is translated into a probabilistic predictor and then integrated in the reinforcement learning process, intended in its partially observable variant. The experimental validation of the architecture proposed is currently ongoing.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        A typical conversational agent has a multi-stage
architecture: spoken language, written language
and dialogue management, see
        <xref ref-type="bibr" rid="ref1">Allen et al. (2001)</xref>
        .
This study focuses on dialogue management for
task-oriented conversational agents. In particular,
we focus on the creation of a dialogue manager
aimed at inducing healthier nutritional habits in
the interactant.
      </p>
      <p>Given that the task considered involves
psychosocial aspects that are difficult to program
directly, the idea of achieving an effective dialogue</p>
      <p>
        Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
manager via machine learning techniques,
reinforcement learning (RL) in particular, may seem
attractive. At present, many RL-based approaches
involve training an agent end-to-end from a dataset
of recorded dialogues, see for instance
        <xref ref-type="bibr" rid="ref22">Liu (2018)</xref>
        .
However, the chance of obtaining significant
results in this way entails substantial efforts in both
collecting sample data and performing
experiments. Worse yet, such efforts ought to rely on the
even stronger hypothesis that the RL agent would
be able to elicit psychosocial aspects on its own.
As an alternative, in this study we envisage the
possibility to enhance the RL process by
harnessing a model developed and accepted in the field
of social psychology to provide a more reliable
learning ground and a substantial accelerator for
the process itself.
      </p>
      <p>
        Our study relies on a quantitative, causal model
of human behavior being studied in the field of
social psychology
        <xref ref-type="bibr" rid="ref7 ref8">(see Carfora et al., 2019)</xref>
        aimed at
assessing the effectiveness of message framing to
induce healthier nutritional habits. The goal of the
model is to assess whether messages with different
frames can be differentially persuasive according
to the users’ psychosocial characteristics.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Psychological model: Structural</title>
    </sec>
    <sec id="sec-3">
      <title>Equation Model</title>
      <p>Three relevant psychosocial antecedents of
behaviour change are the following: Self-Efficacy
(the individual perception of being able to eat
healthy), Attitude (the individual evaluation of the
pros and cons) and Intention Change (the
individual willingness of adhering to a healthy diet).
These psychosocial dimensions cannot be directly
observed and need to be measured as latent
variables. To this purpose, questionnaires are used,
each composed by a set of questions or items
(i.e. observed variables). Self-Efficacy is
measured with 8 items, each associated to a set of
answers ranging from "not at all confident" (1)
to "extremely confident" (7). Attitude is assessed
through 8 items associated to a differential scale
ranging from 1 to 7 (the higher the score, the more
positive the attitude). Intention Change is
measured with three items on a Likert scale, ranging
from 1 (“definitely do not”) to 7 (“definitely do”).
See Carfora et el. (2019).</p>
      <p>
        In our study, the psychosocial model was
assessed experimentally on a group of volunteers.
Each participant was first proposed a
questionnaire (Time 1 – T1) for measuring Self-Efficacy,
Attitude and Intention Change. In a subsequent
phase (i.e. message intervention), participants
were randomly assigned to one of four groups,
each receiving a different type of persuasive
message: gain (i.e. positive behavior leads to
positive outcomes), non-gain (negative behavior
prevents positive outcomes), loss (negative behavior
leads to negative outcomes) and non-loss
(positive behavior prevents negative outcomes)
        <xref ref-type="bibr" rid="ref17 ref9">(Higgins, 1997; Cesario et al., 2013)</xref>
        . In a last phase
(Time 2 - T2), the effectiveness of the message
intervention was then evaluated with a second
questionnaire, to detect changes in participants’
Attitude and Intention Change in relation to healthy
eating.
      </p>
      <p>
        The overall model is described by the
Structural Equation Model
        <xref ref-type="bibr" rid="ref29">(SEM, see Wright, 1921)</xref>
        in Figure 1. For simplicity, only three items are
shown for each latent variable. Besides
allowing the description of latent variables, SEMs are
causal models in the sense that they allow a
statistical analysis of the strength of causal relations
among the latents themselves, as represented by
the arrows in figure. SEMs are linear models, and
thus all causal relations underpin linear equations.
      </p>
      <p>Note that latent variables in a SEM have
different roles: in this case
gain/non-gain/loss/nonloss messages are independent variables, Intention
Change is a dependent variable, Attitude is a
mediator of the relationship between the independent
and the dependent variables, and Self-Efficacy is a
moderator, namely, it explains the intensity ot the
relation it points at. Intention Change was
measures at both T1 and T2, Attitude was measured at
both T1 and T2, and Self-Efficacy was measured at
T1 only. Note that the time transversality (i.e. T1
! T2) is implicit in the SEM depiction above.</p>
    </sec>
    <sec id="sec-4">
      <title>Probabilistic model: Bayesian Network</title>
      <p>
        Once the SEM is defined, we aim to translate
it into a probabilistic model, so as to obtain the
probability distributions needed for the learning
process. We resort to a graphical model, and in
particular to a Bayesian Network
        <xref ref-type="bibr" rid="ref6">(BN, see Ben
Gal, 2007)</xref>
        , namely a graph-based description of
both the observable and latent random variables in
the model and their conditional dependencies. In
BNs, nodes represent the variables and edges
represent dependencies between them, whereas the
lack of edges implies their independence, hence
a simplification in the model. As a general rule,
the joint probability of a BN can be inferred as
follows:
      </p>
      <p>N
P (X1; : : : ; XN ) = Y P (Xi j parents(Xi));
i=1
where X1; : : : ; XN are the random variables in
the model and parents(Xi) indicate all the nodes
having an edge oriented towards Xi.</p>
      <p>
        In the case at hand, a temporal description of
the model, accounting for the time steps T1 and
T2, is necessary as well. For this purpose, we use
a Dynamic Bayesian Network
        <xref ref-type="bibr" rid="ref11">(DBN, see Dagum
et al., 1992)</xref>
        . The DBN thus obtained is shown in
Figure 2.
      </p>
      <p>Notice that the messages are only significant at
T2, as they have not been sent yet at T1. We
gathered message in the one node Message Type,
assuming it can take four, mutually exclusive values.
The mediator Attitude is measured at both time
steps while the moderator Self-Efficacy is constant
over time, as suggested in Section 2. Intention
Change has relevance at T2 only since, as we will
mention in Section 5, it will be used to estimate a
reward function once the final time step is reached.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Learning the BN</title>
      <p>The collected data are as follows. The analysis
was conducted on 442 interactants, divided in four
groups, each one receiving a different type of
messages1. The answers to the items of the
questionnaire always had a range of 7 values.
However, this induces a combinatory esplosion,
making it impossible to cover all the subspaces (78 =
5:764:801 different combinations for Attitude, for
instance). We thus decide to aggregate: low :=
1The original study included also a control group, which
we do not consider here.
(1 to 2); medium := (3 to 5); high := (6 to 7).</p>
      <p>Our aim is to learn the Joint Probability
Distribution (JPD) of our model, as that would make us
able to answer, through marginalizations and
conditional probabilities, any query about the model
itself. The conditional probability distributions to
be learnt in the case in point are then the
following:</p>
      <p>P (Item Ai), for i = 1; : : : ; 8;
P (Item SEi), for i = 1; : : : ; 8;
P (Message Type);
P (Attitude T 1 j Item Ai; i = 1; : : : ; 8);
P (Self-Efficacy j Item SEi; i = 1; : : : ; 8);
P (Attitude T 2 j Item Ai; i = 1; : : : ; 8;
Message Type; Self-Efficacy);</p>
      <p>P (Intention Change j Attitude T 2;</p>
      <p>Self-Efficacy).</p>
      <p>
        The first three can be easily inferred from the raw
data as relative frequencies. As for the following
four, even aggregating the 7 values as mentioned,
a huge amount of data would still be necessary
(38 24 3 = 314:928 subspaces for Attitude T2, for
instance). As conducting a psychological study on
that amount of people would not be feasible, we
address the issue with an appropriate choice of the
method. To allow using Maximum Likelihood
Estimation (MLE) to learn the BN, we resort to the
Noisy-OR approximation
        <xref ref-type="bibr" rid="ref25">(see Onis´ko, 2001)</xref>
        .
According to this, through a few appropriate changes
(not shown) to the graphical model, the number of
subspaces can be greatly reduced (e.g. 3 2 3 = 18
for Attitude T2).
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Reinforcement Learning: Markov</title>
    </sec>
    <sec id="sec-7">
      <title>Decision Problems</title>
      <p>
        The translation into a tool to be used for
reinforcement learning is obtained in the terms of Markov
Decision Processes (MDPs), see
        <xref ref-type="bibr" rid="ref14">Fabiani et al.
(2010)</xref>
        .
      </p>
      <p>
        Roughly speaking, in a MDP there is a finite
number of situations or states of the environment,
at each of which the agent is supposed to select an
action to take, thus inducing a state transition and
obtaining a reward. The objective is to find a
policy determining the sequence of actions that
generates the maximum possible cumulative reward,
over time. However, due to the presence of latents,
in our case the agent is not able to have complete
knowledge about the state of the environment. In
such a situation, the agent must build its own
estimate about the current state based on the memory
of past actions and observations. This entails using
a variant of the MDPs, that is Partially Observable
Markov Decision Processes
        <xref ref-type="bibr" rid="ref19">(POMDPs, see
Kaelbling 1998)</xref>
        . We then define the following, with
reference to the variables mentioned in Figure 2:
S := fstatesg = fAttitude T 2; Self-Efficacyg;
      </p>
      <p>A := factionsg = fask A1; : : : ; ask A8g [
fask SE1; : : : ; ask SE8g [ fG; N G; L; N Lg,
where Ai denotes the question for Item Ai,
SEi denotes the question for Item SEi and
G; N G; L; N L denote the action of sending Gain,
Non-gain, Loss and Non-loss messages
respectively;</p>
      <p>:= fobservationsg =
fItem A1; : : : ; Item A8; Item SE1; : : : ; Item SE8g.</p>
      <p>Starting from an unknown initial state s0 (often
taken to be uniform over S, as no information is
available), the agent takes an action a0, that brings
it, at time step 1, to state s1, unknown as well.
There, an observation o1 is made.</p>
      <p>The process is then repeated over time, until a
goal state of some kind has been reached. Hence,
we can define the history as an ordered succession
of actions and observations:</p>
      <p>ht := fa0; o1; : : : ; at 1; otg ; h0 = ;:</p>
      <p>As at all steps there is uncertainty about the
actual state, a crucial role is played by the agent’s
estimate about the state of the environment, i.e. by
the belief state. The agent’s belief at time step t,
denoted as bt, is driven by its previous belief bt 1
and by the new information acquired, i.e. the
action taken at 1 and observation made ot. We then
have:
bt+1(st+1) = P (st+1 j bt; at; ot+1):
In the POMDP framework, the agent’s choices
about how to behave are influenced by its belief
state and by the history. Thus, we define the
agent’s policy:</p>
      <p>= (bt; ht);
that we aim to optimize. To complete the picture,
we define the following functions to describe the
model evolution in time (the notation 0 indicates a
reference to the subsequent time step):</p>
      <p>state-transition function:
T : (s; a) 7! P (s0 j s; a) := T (s0; s; a);</p>
      <p>observation function:
O : (s; a) 7! P (o0 j a; s0) := O(o0; a; s0);</p>
      <p>reward function:</p>
      <p>R : (s; a) 7! E [r0 j s; a] := R(s; a).</p>
      <p>These functions can be easily adapted to the
specifics of the case at hand. It can be seen that,
once the JPD derived from the DBN is completely
specified, the reward is deterministic. In
particular, it is computed by evaluating the changes in the
values for the latent Intention Change.</p>
      <p>As we are interested in finding an optimal
policy, we now need to evaluate the goodness of each
state when following a given policy. As there is
no certainty about the states, we define the value
function as a weighted average over the possible
belief states:</p>
      <p>V (bt; ht) :=</p>
      <p>X bt(st)V (st; bt; ht);
st
where V (st; bt; ht) is the state value function.
The latter depends on the expected reward (and on
a discount factor 2 [0; 1] stating the preference
for fast solutions):</p>
      <p>V (st; bt; ht) :=R(st; (bt; ht)) +</p>
      <p>X T (st+1;st; (bt; ht))
st+1
X O(ot+1; (bt; ht);st+1)V (st+1; bt+1; ht+1):
ot+1
Finally, we define the target of our seek, namely
the optimal value function and the related optimal
policy, as:
(V (bt; ht) := max V (bt; ht);</p>
      <p>(bt; ht) := argmax V (bt; ht):
It can be shown that the optimal value function in a
POMDP is always piecewise linear and convex, as
exemplified in Figure 3. In other words, the
optimal policy (in bold in Figure 3) combines different
policies depending on their belief state values.</p>
      <p>The next step is to use the POMDP to detect the
optimal policy, that is the sequence of questions to
ask to the interactant, in order to draw her/his
profile, hence the message to send, which maximizes
the effectiveness of the interaction. To this end, the
contribution of the DBN is fundamental. From the
JPD associated, in fact, we construct the
probability distributions necessary to define the functions
T , O, R that compose the value function.
6</p>
    </sec>
    <sec id="sec-8">
      <title>Policy from Monte Carlo Tree Search</title>
      <p>It is evident from Figure 4, describing the full
expansion of the policy tree for the case in point,
that the computational effort and power required
for a brute-force exploration of all possible
combinations is unaffordable.</p>
      <p>Among all the policies that can be considered,
we want to select the optimal ones, thus
avoiding coinsidering policies that are always
underperforming. In other words, with reference to
Figure 3, we want to find Vp1 , Vp2 , Vp3 among those
of all possible policies, and use them to identify
the optimal policy V .</p>
      <p>
        To accomplish this, we select the Monte Carlo
Tree Search (MCTS) approach, see
        <xref ref-type="bibr" rid="ref10">Chaslot et al.
(2008)</xref>
        , due to its reliability and its applicability to
computationally complex practical problems. We
adopt the variant including an Upper Confidence
Bound formula, see
        <xref ref-type="bibr" rid="ref20">Kocsis et al. (2006)</xref>
        . This
method combines exploitation of the previously
computed results, allowing to select the game
action leading to better results, with exploration of
different choices, to cope with the uncertainty of
the evaluation. Thus, using V (st; bt; ht) as
defined before to guide the exploration, the MCTS
method reliably converges (in probability) to
optimal policies. These latter will be applied by the
conversational agent in the interaction with each
specific user, to adapt both the sequence and the
amount of questions to her/his personality profile
and selecting the message which is most likely to
be effective.
7
      </p>
    </sec>
    <sec id="sec-9">
      <title>Conclusions and future work</title>
      <p>In this work we explored the possibility of
harnessing a complete and experimentally assessed
SEM, developed in the field of persuasion
psychology, as the basis for the reinforcement
learning of a dialogue manager that drives a
conversational agent whose task is inducing healthier
nutritional habits in the interactant. The
fundamental component of the method proposed is a DBN,
which is derived from the SEM above and acts like
a predictor for the belief state value in a POMDP.</p>
      <p>The main expected advantage is that, by doing
so, the RL agent will not need a time-consuming
period of training, possibly requiring the
involvement of human interactants, but can be trained ‘in
house’ – at least at the beginning – and be released
in production at a later stage, once a first
effective strategy has been achieved through the DBN.
Such method still requires an experimental
validation, which is the current objective of our working
group.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>The authors are grateful to Cristiano Chesi of
IUSS Pavia for his revision of an earlier version
of the paper and his precious remarks. We also
acknowledge the fundamental help given by
Rebecca Rastelli, during her collaboration to this
research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferguson</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Stent</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>An architecture for more realistic conversational systems</article-title>
          .
          <source>In Proceedings of the 6th international conference on Intelligent user interfaces</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>Ronald D.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Vastag</surname>
          </string-name>
          , Gyula.
          <year>2004</year>
          .
          <article-title>Causal modeling alternatives in operations research: Overview and application</article-title>
          .
          <source>European Journal of Operational Research</source>
          .
          <volume>156</volume>
          .
          <fpage>92</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Auer</surname>
          </string-name>
          , Peter &amp;
          <string-name>
            <surname>Cesa-Bianchi</surname>
            ,
            <given-names>Nicolò</given-names>
          </string-name>
          &amp; Fischer, Paul. 2002Kocsis,
          <string-name>
            <surname>Levente</surname>
          </string-name>
          &amp; Szepesvári, Csaba.
          <year>2006</year>
          .
          <article-title>Bandit Based Monte-Carlo Planning. Finite-time Analysis of the Multiarmed Bandit Problem</article-title>
          .
          <source>Machine Learning</source>
          .
          <volume>47</volume>
          .
          <fpage>235</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bandura</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>1982</year>
          .
          <article-title>Self-efficacy mechanism in human agency</article-title>
          .
          <source>American Psychologist</source>
          ,
          <volume>37</volume>
          ,
          <fpage>122</fpage>
          -
          <lpage>147</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Baron</surname>
            ,
            <given-names>Robert A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Byrne</surname>
            ,
            <given-names>Donn</given-names>
          </string-name>
          <string-name>
            <surname>Erwin</surname>
          </string-name>
          &amp; Suls,
          <string-name>
            <surname>Jerry</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>1989</year>
          .
          <article-title>Exploring social psychology</article-title>
          , 3rd ed. Boston, Mass.:
          <source>Allyn and Bacon</source>
          .
          <volume>0205119085</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Ben Gal</surname>
            <given-names>I.</given-names>
          </string-name>
          <year>2007</year>
          .
          <string-name>
            <given-names>Bayesian</given-names>
            <surname>Networks</surname>
          </string-name>
          .
          <source>Encyclopedia of Statistics in Quality and Reliability</source>
          . John Wiley &amp; Sons.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Bertolotti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carfora</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Catellani</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Different frames to reduce red meat intake: The moderating role of self-efficacy</article-title>
          .
          <source>Health Communication</source>
          , in press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Carfora</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertolotti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Catellani</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Informational and emotional daily messages to reduce red and processed meat consumption</article-title>
          .
          <source>Appetite</source>
          ,
          <volume>141</volume>
          ,
          <fpage>104331</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Cesario</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corker</surname>
            ,
            <given-names>K. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jelinek</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>A selfregulatory framework for message framing</article-title>
          .
          <source>Journal of Experimental Social Psychology</source>
          ,
          <volume>49</volume>
          ,
          <fpage>238</fpage>
          -
          <lpage>249</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Chaslot</surname>
          </string-name>
          , Guillaume &amp; Bakkes, Sander &amp; Szita, Istvan &amp; Spronck, Pieter.
          <year>2008</year>
          .
          <article-title>Monte-Carlo Tree Search: A New Framework for Game AI</article-title>
          . Bijdragen.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Dagum</surname>
          </string-name>
          , Paul and Galper, Adam and Horvitz, Eric.
          <year>1992</year>
          .
          <article-title>Dynamic Network Models for Forecasting</article-title>
          .
          <source>Proceedings of the Eighth Conference on Uncertainty in Artificial Intelligence.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Dagum</surname>
          </string-name>
          , Paul and Galper, Adam and Horvitz, Eric and Seiver, Adam.
          <year>1999</year>
          .
          <article-title>Uncertain reasoning and forecasting</article-title>
          .
          <source>International Journal of Forecasting.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>De Waal</surname>
          </string-name>
          , Alta &amp; Yoo, Keunyoung.
          <year>2018</year>
          .
          <article-title>Latent Variable Bayesian Networks Constructed Using Structural Equation Modelling</article-title>
          .
          <source>2018 21st International Conference on Information Fusion</source>
          .
          <fpage>688</fpage>
          -
          <lpage>695</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Fabiani</surname>
          </string-name>
          , Patrick &amp;
          <string-name>
            <surname>Teichteil-Königsbuch</surname>
          </string-name>
          ,
          <year>Florent</year>
          .
          <year>2010</year>
          .
          <article-title>Markov Decision Processes in Artificial Intelligence</article-title>
          . Wiley-ISTE.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Gupta</surname>
          </string-name>
          , Sumeet &amp; W. Kim, Hee.
          <year>2008</year>
          .
          <article-title>Linking structural equation modeling to Bayesian networks: Decision support for customer retention in virtual communities</article-title>
          .
          <source>European Journal of Operational Research</source>
          .
          <volume>190</volume>
          .
          <fpage>818</fpage>
          -
          <lpage>833</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Heckerman</surname>
          </string-name>
          , David.
          <year>1995</year>
          .
          <article-title>A Bayesian Approach to Learning Causal Networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Higgins</surname>
            ,
            <given-names>E.T.</given-names>
          </string-name>
          <year>1997</year>
          .
          <article-title>Beyond pleasure and pain</article-title>
          .
          <source>American Psychologist</source>
          ,
          <volume>52</volume>
          ,
          <fpage>1280</fpage>
          -
          <lpage>1300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Howard</surname>
          </string-name>
          , Ronald.
          <year>1972</year>
          .
          <article-title>Dynamic Programming and Markov Process</article-title>
          .
          <source>The Mathematical Gazette</source>
          .
          <volume>46</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Pack</given-names>
            <surname>Kaelbling</surname>
          </string-name>
          , Leslie &amp; Littman,
          <string-name>
            <given-names>Michael &amp; R.</given-names>
            <surname>Cassandra</surname>
          </string-name>
          , Anthony.
          <year>1998</year>
          .
          <article-title>Planning and Acting in Partially Observable Stochastic Domains</article-title>
          .
          <source>Artificial Intelligence</source>
          .
          <volume>101</volume>
          .
          <fpage>99</fpage>
          -
          <lpage>134</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Kocsis</surname>
          </string-name>
          , Levente &amp; Szepesvári, Csaba.
          <year>2006</year>
          .
          <article-title>Bandit Based Monte-Carlo Planning</article-title>
          .
          <source>Machine Learning: ECML 2006</source>
          . Springer Berlin Heidelberg. 282-
          <fpage>293</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>T.L</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Robbins</surname>
          </string-name>
          , Herbert.
          <year>1985</year>
          .
          <article-title>Asymptotically Efficient Adaptive Allocation Rules</article-title>
          .
          <source>Advances in Applied Mathematics. 6</source>
          .
          <fpage>4</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , Bing.
          <year>2018</year>
          .
          <article-title>Learning Task-Oriented Dialog with Neural Network Methods</article-title>
          .
          <source>PhD thesis.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Murphy</surname>
          </string-name>
          , Kevin.
          <year>2012</year>
          .
          <article-title>Machine Learning: A Probabilistic Perspective</article-title>
          . The MIT Press.
          <volume>58</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Pearl</given-names>
            <surname>Judea</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference</article-title>
          .
          <source>Representation and Reasoning Series (2nd printing ed.)</source>
          . San Francisco, California: Morgan Kaufmann.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>Onis´ko</article-title>
          , Agnieszka &amp; Druzdzel,
          <string-name>
            <given-names>Marek J.</given-names>
            &amp;
            <surname>Wasyluk</surname>
          </string-name>
          , Hanna.
          <year>2001</year>
          .
          <article-title>Learning Bayesian network parameters from small data sets: application of Noisy-OR gates</article-title>
          .
          <source>International Journal of Approximate Reasoning</source>
          .
          <volume>27</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Silver</surname>
          </string-name>
          , David &amp; Veness, Joel.
          <year>2010</year>
          .
          <article-title>Monte-Carlo Planning in Large POMDPs</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          .
          <volume>23</volume>
          .
          <fpage>2164</fpage>
          -
          <lpage>2172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Matthijs</surname>
            <given-names>T. J.</given-names>
          </string-name>
          <string-name>
            <surname>Spaan</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Partially Observable Markov Decision Processes. In: Reinforcement Learning: State of the Art</article-title>
          . Springer Verlag.
          <volume>387</volume>
          -
          <fpage>414</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Sutton</surname>
          </string-name>
          , Richard &amp; G. Barto, Andrew.
          <year>1998</year>
          .
          <article-title>Reinforcement Learning: An Introduction</article-title>
          .
          <source>IEEE transactions on neural networks / a publication of the IEEE Neural Networks Council. 9</source>
          . 1054.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Wright</surname>
          </string-name>
          , Sewall.
          <year>1921</year>
          .
          <article-title>Correlation and causation</article-title>
          .
          <source>Journal of Agricultural Research</source>
          .
          <volume>20</volume>
          .
          <fpage>557</fpage>
          -
          <lpage>585</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Young</surname>
          </string-name>
          , Steve &amp; Gasic, Milica &amp; Thomson, Blaise &amp; Williams, Jason.
          <year>2013</year>
          .
          <article-title>POMDP-based statistical spoken dialog systems: A review</article-title>
          .
          <source>Proceedings of the IEEE</source>
          ,
          <volume>101</volume>
          .
          <fpage>1160</fpage>
          -
          <lpage>1179</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>