<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Large Language Models and RAG Approach for Conversational Coaching - Experiments for Enhancing e-VITA Virtual Coach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kristiina Jokinen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kristina Deryagina</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulio Napolitano</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abrar Hyder</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AI Research Center, National Institute of Advanced Industrial Science and Technology</institution>
          ,
          <addr-line>2-4-7 Aomi, Koto, Tokyo 135-0064</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institut für Angewandte Informatik (InfAI)</institution>
          ,
          <addr-line>Goerdelerring 9, 04109 Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Siegen</institution>
          ,
          <addr-line>Adolf-Reichwein-Straße 2a, 57076 Siegen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We give an overview of the dialogue modelling research conducted for the e-VITA Virtual Coach, developed in the EU-Japan collaboration project e-VITA with the aim of supporting personalised interaction with older adults, on topics related to active healthy living. The system integrates Rasa Conversational AI and RAG-based LLMs and addresses two main requirements for AI-based applications: how to provide reliable information, and how to maintain smooth personalised conversation. We focus on motivational coaching dialogues and qualitative user evaluation of the prototype system, which tailors dialogue interaction to develop user interest and to support their daily activities.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;e-VITA Virtual Coach</kwd>
        <kwd>Dialogue system</kwd>
        <kwd>Large Language Models</kwd>
        <kwd>RAG</kwd>
        <kwd>LangChain</kwd>
        <kwd>reliable interaction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>With the release of OpenAI’s ChatGPT in November 2022, Large Language Models (LLMs) have made
generative AI known to a large audience and excelled in several language-based tasks ranging from code
writing to translation and from question-answering to chatty communication. LLMs also ofer attractive
solutions to build social robotics applications, which have so far sufered from clumsy and cumbersome
interaction. However, much research has been conducted to alleviate the notorious “hallucinations” of
the LLMs in order to improve their performance for various interactive tasks. This is especially important
for practical social robotics applications which not only need smooth communication capabilities but
also provide reliable information for users in a personalised manner.</p>
      <p>
        In this article, we describe our work on one such practical application, the e-VITA Virtual Coach,
and focus especially on the e-VITA Dialogue Manager. The system has been developed in the EU-Japan
collaboration project e-VITA (www.e-vita.coach/homepage) to provide interactive coaching for active
and healthy ageing. The project aims at improving subjective well-being of older adults in Europe and
Japan, and supporting quality of life in daily activities through integrating advanced living technologies,
robots, AI, and conversational interaction in the context of smart homes, personalised coaching, and
companionship. The main characteristics of the project are described in [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ].
      </p>
      <p>The article is structured as follows. Section 2 describes the framework for our dialogue modelling,
the e-VITA Coaching system, Section 3 describes the motivational coaching approach with personalised
coaching cycle, and Section 4 discusses the three dialogue management approaches used in the e-VITA
Dialogue Model. Section 5 provides an evaluation of the system, and Section 6 discusses ethical aspects
related to reliable and sustainable coaching models. Finally, Section 7 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Overview of the e-VITA Coaching System</title>
      <p>
        Previous work on coaching systems within healthcare and eldercare domains has focussed on technical
aspects of the systems, or on behaviour change models and intervention methods using scripted
dialogues. Only a few explicitly address dialogue aspects [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. Relational agents [6] engage users in a
natural interaction by relating to their emotional state, while Zhou et al. [7] present a neurosymbolic
goal summarizer to help health coaches keep track of the goals. An overview of virtual coaching for
older adults is presented e.g. in [8].
      </p>
      <sec id="sec-2-1">
        <title>2.1. Specification of the system</title>
        <p>The main goal of the e-VITA Virtual Coach development has been trustworthiness, i.e. to design and
develop a coaching system that can sustain older adults’ well-being in everyday activities. For this, it is
important that the coach can interact in natural language, and facilitate interaction through intuitive
conversational coaching strategies. The Virtual Coach should also converse with the older adults in
their native language, and the multilingual project setting of the EU and Japan is thus challenging: the
Coach needs to manage five diferent languages (English, German, French, Italian, and Japanese) and
also take into account intercultural diferences such as presentation styles, politeness, and coaching
strategies. Finally, the system is designed to be used in smart living environments so the integration
of various devices and sensors is pertinent within the e-VITA platform, including their impact on the
interaction. This is another challenge for the technical management of the e-VITA Coach which aims
to support interoperability and standardization, and protect user privacy issues.</p>
        <p>In this context, coaching dialogues have been designed to support and stimulate the daily activities
of older adults, as well as provide information and advice in several domains, including exercising,
sleeping, healthy eating, mental health, and spirituality. When interacting with a particular user, the
Virtual Coach is required to take into account the personal profile of the user, including their cultural
background, age, gender, health situation as well as their readiness level for motivational activity
changes, which the content experts considered relevant features for personalising the Coach responses.
The physical environment surrounding the user is exposed to the dialogue system from sensor readings,
while emotion detection informs the system about the likely psychological status of the interlocutor.
The challenges related to multilinguality are addressed by translation services and language-specific
ifne-tuning: user utterances and system responses are translated between the user’s language and
English, used as lingua-franca in system-internal processing and domain-specific databases, whereas
language-specific aspects of greetings and database answers are manually fine-tuned and coded in the
Dialogue Manager rules.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. The e-VITA Dialogue Manager</title>
        <p>The e-VITA Dialogue Manager implements dialogue management of the e-VITA Coach, taking care of
the interpretation of the user’s intent and the generation of an appropriate system response. It uses
the Rasa Open-Source Framework (https://rasa.com) to build the conversational agent. To capture a
wider array of user intents and to make system responses more fluent, the e-VITA Dialogue Manager
takes advantage of the GPT models, and incorporates ChatGPT in the system via an extension of the
Retrieval Augmented Generation (RAG) approach.</p>
        <p>Figure 1 shows the integration of the Rasa dialogue model in the e-VITA coach framework. The coach
platform, called the Digital Enabler, coordinates the information flow and also contains an API hub
which includes information from the sensors, the emotion recognition system, motion trackers, and
environmental sensors. It also contains the user’s private information (such as BMI, exercise, fitness,
hobbies, eating habits), which is used to drive the dialogues through diferent individualized pathways.</p>
        <p>When a user starts a conversation with the e-VITA Coach, the audio signal is captured by the
microphone of the interface agent (Nao, Gatebox, Android, Daruma) and sent to the Google speech
recognizer (https://cloud.google.com/speech-to-text), except for the Gatebox interface, which relies on
the Microsoft Azure speech engines (https://portal.azure.com/). The transcribed text is subsequently
managed by Digital Enabler, which uses DeepL translator (https://www.deepl.com/en/translator) to
translate the text into English (if not originally in English), and sends the text to the Rasa dialogue
service. The NLU component of the Rasa system processes the input and uses the DIET (Dual Intent
and Entity Transformer) to recognize dialogue intents and entities from the user input. Rasa’s Dialogue
Manager component produces the next system response with the help of stories (model dialogues)
and exploiting the TED (Transformer Embedding Dialogue) policy. The response text is returned to
Digital Enabler, which translates it into the user’s native language and sends this further to the selected
coaching device. The robot agent then speaks the text to the user using its built-in TTS technology.</p>
        <p>The system also has access to external services (Wikipedia, News, and Weather services), and it
includes domain-specific knowledge bases (KB), i.e. json-type databases and collections of relevant
documents for the RAG approach, compiled by the project’s domain experts. Digital Enabler also
provides a notification management system which enables external triggers to start dialogues based
on sensor status or timing. For instance, the user can be prompted at 8 am every morning to exercise
outside, if the weather is not too hot.</p>
        <p>An example of a coaching dialogue is given in Table 1. It shows the recognized user intent and the
AI agent’s response intent (to be executed as a system action like knowledge-base search, request to
external services, request to the robot utterancece). As the dialogues are tailored according to the
user’s specific characteristics, Table 1 also shows some user features and their values that will guide the
dialogue. The features are provided by the content experts and include interests, family relations, age,
etc., and can also include more personal data depending on the privacy constraints.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Coaching Model</title>
      <sec id="sec-3-1">
        <title>3.1. Motivational Coaching</title>
        <p>The e-VITA coaching model is based on Motivational Coaching Theory [9, 10], which incorporates
an understanding of the user’s preparedness to change their life style. The user’s preparedness is
categorized into four readiness levels: non-interested, interested, ready (prepared), and maintenance, and
the e-VITA Coaching Model interprets the first three levels as diferent user types, for whom the Virtual
Dialogue Coach should tailor the coaching dialogues accordingly, in order to motivate behavioural
change to the next level. These are:
Absolutely, I’d be happy to help. Do you
have specific goals in mind, like improving
strength, flexibility, or overall fitness? Let’s
work together to create a plan that suits
your interests and lifestyle.</p>
        <p>I have problems with motivation. I cannot
make myself do even a little bit of sport
every day. Can you help me?
Yes, setting small realistic goals and
tracking your progress can help with motivation.</p>
        <p>You can also try doing activities you enjoy
or joining a group for accountability and
camaraderie.</p>
        <p>agree_coachingexercise
request_findingmotivation
agree_findingmotivation
interest-level=0,
hobby=cooking,
gender=female
• Not interested person: Mainly "working on the idea”. The goal is to make the user aware of the
need for behavior change. Increase the user’s knowledge of health behaviors and help them
understand the benefits of behavior change and the risks of not doing so. Also ask users to express
their feelings.
• Interested person: The main focus is on “giving the idea”. The goal is to motivate the user and give
the user more confidence in their ability to change their behavior. Identify obstacles to behavior
change. Continue to increase knowledge of health behaviors.
• Prepared person: Main focus is on “working on behavior”. The goal is to clarify the action plan, i.e.
have the user make a concrete and achievable plan and be determined to implement the behavior.
Once the user has started even a little, follow up with them so that their determination does not
waver. To encourage the use of self-monitoring, rewards such as points, and social support.</p>
        <p>For the current implementation of the coaching dialogues, particular attention is given to the first two
stages, non-interested and interested, operating under the assumption that the user intends to modify
their behavior. As for the "maintenance" level users, it was assumed that they would benefit more from
a system that monitors their performance and accurately tells what is the current situation, encouraging
them to continue or notifying them of any anomalies. Since the project focussed on behavioural change
and personalised coaching, dialogue management for the "maintenance" level users was not included in
the system development.</p>
        <p>The coach provides customized advice for older adults on domains such as active healthy living,
nutrition, and exercises. Motivational coaching is delivered through a structured process known as
the coaching cycle, a conceptual framework which guides users through steps designed to foster
selfawareness, encourage self-reflection, and facilitate the adoption of healthier lifestyle choices. For each
of the three focussed user types, diferent “stories”, i.e. sample dialogues, are manually defined by
the content experts. The stories represent prototypical dialogues that a user with a given profile and
readiness level can conduct with the coach, and they are included in the document pool from which data
is retrieved in order to generate the system response. For instance, if the user is “interested”, then the
sample stories for an interested user are selected and processed together with the content documents.
A sample story with diferent options for the user utterances is given in Table 2.</p>
        <p>The interested/non-interested user types were extensively studied in the experiments exploring
usability of ChatGPT and the RAG approach in the project-specific tasks such as generating personal
iftness and exercise plans [ 11]. While ChatGPT is capable of generating exercise plans, its use as a</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Personalised Coaching Mode</title>
        <p>The primary distinction between a personalised coaching mode and a general, non-personalised coaching
mode lies in the level of customization and specificity of interactions based on the user data. In the
personalised coaching mode, the system adapts its responses using detailed information about the user,
while in the general mode, the system has no access to the user’s private information and thus cannot
tailor its responses according to the individual’s unique context and needs.</p>
        <p>The user’s private information is categorized into several subdomains in the e-VITA Coach as follows:
• Personal Information: the user’s login ID, name, and location
• Medical Information: the user’s height, weight, age, and medical history (e.g., any illnesses or
allergies)
• Daily Routine: details about the user preference for activities like grocery shopping (morning,
evening, or afternoon) and whether they are independent or dependent in doing these tasks
• Eating Habits: the number of meals the user consumes per day and daily calorie intake
• Exercise: the user’s preferred exercises and their readiness level for exercising
• Social Habits: details about the user living alone or with family, their social preferences (e.g.,
preference for socializing or solitude), pet ownership, religious afiliation.</p>
        <p>Users can update their private information via the Evita dashboard (an interface part of the Dialogue
Enabler framework). In the beginning of a dialogue session between a user and the Virtual Coach, the
user has to login to the system using their unique ID, which is detected and stored in the Dialogue
Manager. The user ID is then used to collect user-specific information from the Digital Enabler, thus
enabling e-VITA Dialogue Manager to access the user’s private information in order to personalise its
responses appropriately.</p>
        <p>By default, the system is designed to run fixed, tailor-made dialogues based on RASA’s stories. It
is assumed that the personalized coaching cycle enhances user engagement by leveraging detailed
personal information, whereas the normal coaching cycle provides broader, less tailored responses
due to the lack of user-specific data. Tables 3 and 4 illustrate the diferences in responses between
personalized and normal coaching cycles, highlighting the impact of user-specific data on the quality
and relevance of interactions.</p>
        <p>Dialogue
What do you think about my living condition?
Hello, test_user99. Living alone and feeling lonely can be challenging,
especially when you don’t have frequent contact with friends and relatives. It’s
important to find ways to stay connected and engaged with others. You might
consider joining community activities or groups that interest you, which can
help reduce feelings of loneliness. Additionally, maintaining a routine that
includes physical activity and hobbies can improve your overall well-being.</p>
        <p>How do you feel about exploring some new activities or groups in your area?</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. RASA and RAG Dialogue Management</title>
      <p>As discussed in Section 2, dialogue management for the e-VITA Coaching system is operationalised
with the help of the Rasa Conversational AI framework, but it also utilises LLMs to produce smooth
system responses, and experiments with ChatGPT in a special coaching mode. Dialogue management
in the e-VITA Coach thus deals with three types of dialogues:
• Story-driven Dialogues, using typical stories designed in RASA
• LLM-assisted Dialogues, based on the RAG approach and motivational Coaching cycles
• Open-ended Conversations with OpenAI’s ChatGPT.</p>
      <p>The conversation flow is given in Figure 2. After recognition of the user intent, i.e. processing the
utterance through the NLP module and classifying it into one of the available user intents, a relevant
story is triggered when possible. If the confidence for triggering a relevant story is below the threshold,
the system will start the LLM-assisted dialogue management, i.e. it utilises the RAG-based approach to
generate an appropriate response with the reliable data source.</p>
      <p>Due to the popularity of OpenAI’s ChatGPT, the e-VITA system also gives the user an opportunity to
experiment with it. If the user wishes to engage in direct ChatGPT conversations through the e-VITA
Coach, they can activate the AI-mode by a particular keyword phrase "Activate artificial intelligence".
Similarly, to end the interaction with ChatGPT, the user must say "Deactivate artificial intelligence".
This is to ensure that the users are fully aware of starting an interaction with Open AI’s ChatGPT,
and that the responses may be inaccurate or false so they need to assess responses carefully. Since the
LLM-assisted dialogues with the e-VITA Coach use a curated document set for response generation, the
responses are assumed to be reliable given the expert knowledge available in the project. However, the
direct conversations with ChatGPT were only available in the Japanese version of the e-VITA Coach.
This was due to difering views between the ethical committees in the EU and Japan of the acceptability
of the direct use of ChatGPT in user experiments. In the EU system, the command "Activate artificial
intelligence" always uses the path for LLM-assisted dialogue management, i.e. the RAG approach,
whereas in the Japanese system, it evokes direct conversation with ChatGPT. In the EU-system, the
LLM-assisted Dialogues and the Open-ended dialogues thus end up using the same RAG-based approach,
as seen in Figure 2 where the this path branches out to LLM-assisted dialogue management.</p>
      <p>Retrieval Augmented Generation (RAG) [12] is utilised in the e-VITA Coach to address the issue of
reliable and truthful response generation. Custom-selected content is used as a resource to guide the
truthfulness of the responses and enable the dialogue manager to answer queries based on the material
which is pertinent to the application domain. The dialogue manager thus adopts a form of guided text
generation shown in Figure 3. The retrieved text, together with the original user query, is input to the
LLM, which generates a suitable response to the user.</p>
      <p>The documents are selected by the content experts. The texts are encoded in a vector format using
LangChain (https://www.langchain.com/) and the FAISS (Facebook AI Similarity Search) vector store.
Since there is a limit for tokens in QA tasks, token size is reduced by extracting information to create
chunks of information.</p>
      <p>Figure 4 shows a prompt designed to produce an answer with 3-4 sentences, in a given context. The
token limit is set to 256 and the temperature parameter to 0, favoring tokens with the highest probability
in the output string. The prompt template strictly instructs the system to generate responses from the
given documents and avoid any made-up information. This helps to ensure accuracy and eliminate the
risk of hallucinations or fabricated answers.</p>
      <p>The input variables in the prompt template include the context and the user question, along with the
additional parameter: "user info". This parameter contains a comprehensive list of the user’s private
data, collected by the Dialogue Manager across the subdomains of user information (see Section 3.2). It
is used to direct the system to produce highly personalised responses. When generating the response,
the information in all three parameters in the prompt template is taken into account, thus enabling the
coach to adapt responses according to the user data. On the other hand, the prompt for the general
coaching cycle lacks the "user info" parameter, and thus the system does not have access to the specific
user data. Without knowledge of the user’s personal details, the system cannot tailor its responses to
the individual’s needs, and response generation results in more generalized utterances.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation</title>
      <p>
        The present article details the user evaluation results of the earlier versions of the dialogue system
[
        <xref ref-type="bibr" rid="ref2">14, 2</xref>
        ], and provides a comparison with the version incorporating the coaching cycle. At the time
of writing, the results of user tests for the latest version of the system were still under analysis and
will be reported in forthcoming papers and project documentation. On the technical side, the e-Vita
platform evaluation is described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], while the initial evaluations of ChatGPT and the RAG approach
are reported in [11].
      </p>
      <sec id="sec-5-1">
        <title>5.1. Evaluation methodology</title>
        <p>The evaluation methodology comprised mostly qualitative approaches, ensuring user satisfaction and a
thorough understanding of the system’s performance. The project included a long evaluation period
through series of user studies when the coaching system was placed in an elder person’s home and the
user was asked to test and evaluate the system usage at their convenient time and technical assessment.
Each session was designed to simulate real-life interactions and included tasks such as daily exercise
recommendations, nutritional advice, and mental well-being support.</p>
        <p>Participants were recruited from various demographics, focusing on older adults who would benefit
from the e-VITA coaching system. The recruitment process involved collaboration with community
centers and healthcare providers to ensure a diverse participant pool.</p>
        <p>Data were collected through questionnaires, interviews, and system logs. Pre-study and post-study
questionnaires measured changes in participants’ health behavior and attitudes. Interviews provided
qualitative insights into user experiences, while system logs recorded interaction patterns and system
performance.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Evaluation results</title>
        <p>In one of the studies, primarily aimed at assessing participants’ interaction with the robot and their
acceptance of it, participants were engaged in various activities, such as motivational talks, memory
training, cognitive challenges, and physical exercises. In interviews, participants mentioned they
appreciated the opportunity to engage in conversations with the robot on topics such as health and
family, which provided a sense of interaction and companionship. Users expressed that having the robot
equipped with a dialogue system at home could motivate them to engage in exercises, even outside of
formal exercise classes. Some participants found the conversations with the robot to be intellectually
stimulating. Thus, it can be concluded that one of the practical values of the dialogue system inside the
robot for the participants was facilitating interactions and providing a sense of connection and support.
However, the restricted scope of dialogue topics and technical constraints such as limitations in the
robot’s ability to engage in spontaneous and natural conversations were noted as drawbacks.</p>
        <p>In the final version of the system, OpenAI’s text generation models were utilized. The advanced
capabilities of Large Language Models enable more dynamic and contextually relevant responses by
leveraging a vast amount of pre-trained knowledge, allowing the robot to engage in more spontaneous
and natural conversations. This method reduces the occurrence of repetitive or irrelevant responses
and enhances the system’s ability to handle a broader range of topics and user intents.</p>
        <p>Another study aimed to enhance user experience by ofering a means of communication with the
robot for daily living tasks. Participants especially appreciated the system’s capability to engage in
conversations about everyday routine. Users found the system useful for tasks such as checking the
weather forecast, expressing emotions, and planning activities. The functional advantage of the dialogue
system inside the robot for the participants was to provide them with a conversational interface that
could assist them in various daily living activities. However, issues with timing and the speed of
responses were mentioned, with some feeling that conversations were rushed or one-sided, and delays
between conversations. Certain users found it challenging to maintain a smooth conversation with the
robot, requiring them to repeat themselves or speak loudly.</p>
        <p>These issues were addressed by optimizing the response generation process. The integration of the
Retrieval-Augmented Generation model and improvements in the dialogue management flow reduced
response times and enhanced the fluidity of conversations. Additionally, the use of advanced Automatic
Speech Recognition and Text-to-Speech technologies improved the system’s ability to understand and
respond promptly and accurately, thereby minimizing delays and ensuring smoother interactions. In
[11], RAG’s promising start with increased reliability of the responses ia confirmed, but the need to
validate the initial data to ground responses in real-world representation and to thoroughly evaluate
the assistant’s responses are also pointed out.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Trustworthy Coaching</title>
      <p>The basic principles of Trustworthy AI underpin the research and are an important part of the
standardization work [15], which continues after the project end. There are working groups and organizations
like MPAI (https://mpai.community/standards/mpai-mmc/) which aim to unify AI tools and interfaces
to give recommendations on use cases and functional requirements. As for dialogue modeling research,
there are standardization eforts related to ISO-standards of dialogue acts [ 16]. We also draw on risk
management research to mitigate epistemic (the process of producing knowledge) as well as interactive
(usability) risks, especially in the context of healthy aging in the digital world (AAL guidelines).</p>
      <p>Ethical aspects of speech-based AI agents also need to be considered carefully, and the main points
are discussed in [17]. Advances in LLM technology enable chatbots to generate and analyze the content
of conversations and store personal data on the web. As generative AI agents can predict responses
and produce coherent sounding but inaccurate or fabricated content, this can cause both mistrust
in the service (even in cases where the information is true) and over-trust (even in cases where the
information is clearly wrong). In coaching scenarios, mistrust and over-trust are not desired, so the
design of the coaching dialogues needs to pay attention to the accuracy of the information as well as
how to support the facts and explain their validity in context. The AI agents do not ‘know’ the meaning
of the utterances they generate, so technically their understanding and generation capability needs to
be extended with such information and be grounded in the true facts of the context and the multimodal
environment, besides in the dialogue history and context of the conversation (see discussion in [18]).</p>
      <p>Moreover, other issues appear as important aspects in the quest for balanced trustworthy interaction.
As is briefly noted in [ 11], interactive learning and adversarial learning may be useful to adapt the
system to the user’s behavior and to control the LLM agent’s behavior, but also have a downside in the
vulnerability of the development which can include malicious intents. It is thus necessary to support
ethical and sustainable practices in designing and using practical AI agents such as virtual coaches and
friendly assistants.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>This article has provided an overview of the Dialogue Manager in the e-VITA Virtual Coach, focussing
on trustworthy and reliable interaction for older adults in order to support their well-being and active
healthy life. The work comprises of the necessary tools to acquire, represent, integrate, and analyze a
variety of data generated in the interactions between the human user and the robotic interface devices,
and the implementation of the e-VITA Dialogue Manager, which integrates advanced AI technology
and various natural language processing components into an interactive system aiming to provide
useful and flexible coaching for older users.</p>
      <p>The Dialogue Manager is capable of personalised coaching conversations with the users, and its
design emphasises reliable and responsible development of LLM-based systems. With the three diferent
dialogue management models, i.e. regular Story-driven, LLM-assisted, and Open-ended dialogues,
the e-VITA Dialogue Manager exemplifies a hybrid combination of rule-based and AI-based dialogue
management approaches, which together support various functionalities of the Virtual Coach, and
contribute to the active research area of LLMs and interactive AI systems. Furthermore, the experience
gained through co-creating requirements for technical development together with the end-users will be
valuable when designing and developing future applications in the context of smart home techology,
generative AI, and robot interfaces, with the purpose of supporting well-being and quality of life in a
responsible and sustainable manner.</p>
      <p>The combination of the dialogue management approaches is to be evaluated further in diferent
realworld situations and application scenarios, so as to reach full understanding of its (combined) benefits
and drawbacks for practical applications. For instance, integration of data from sensors and smart living
environments can be extended and elaborated further in order to supply the coaching system with a
model of how to include dynamic data in the triggering system, to enable system-initiated interaction
management. The current system deals with environmental sensors and time-scheduling, and interesting
aspects concern extensions to multimodal data from smart home scenarios. An important extension is
also integration of external knowledge (environment, world, situation, user) into dialogue modelling, and
in particular, how to best use knowledge graphs in LLM-based interaction management. Much research is
currently ongoing and published in this respect, here we only refer to two events where these challenges
have been discussed: the e-VITA final workshop (https://www.e-vita.coach/homepage/finalconference/)
and the SIGDial Special Session (https://2024.sigdial.org/call-for-papers-ss/).</p>
      <p>Finally, the technical aspects related to the deployment of the components of the platform in the
e-VITA project are fundamental issues connected to key technical and privacy-related issues. It may
not be possible to define one single methodology for the deployment and provisioning of coaching
and NLP services because this is strictly related to the technical requirements of the components (e.g.
base development technology, sensor specifications), typology of data processing (e.g. algorithms),
and use cases (e.g. performance and privacy restrictions). In order to choose the best way to realize a
concrete components deployment, it is necessary to adopt multiple strategies that cover both cloud and
edge scenarios to fulfill specific requirements related to these features. In the hybrid approach, a single
application or service can use components that are on the edge or cloud side, or a combination of the
two depending on the unique requirements of the application.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>We thank all the project partners for useful and fruitful discussions and enjoyable cooperation. Special
thanks go to the colleagues in technical workpackages cooperating in the integration of the dialogue
module in the e-VITA platform and to the content experts on the example dialogues for coaching.
[6] T. W. Bickmore, Relational Agents in Health Applications: Leveraging Afective
Computing to Promote Healing and Wellness, in: The Oxford Handbook of Afective Computing,
Oxford University Press, 2015. URL: https://doi.org/10.1093/oxfordhb/9780199942237.013.010.
doi:10.1093/oxfordhb/9780199942237.013.010.
[7] Y. Zhou, B. D. Eugenio, B. Ziebart, L. Sharp, B. Liu, N. Agadakos, Modeling low-resource health
coaching dialogues via neuro-symbolic goal summarization and text-units-text generation, 2024.</p>
      <p>URL: https://arxiv.org/abs/2404.10268. arXiv:2404.10268.
[8] N. Lete, A. Beristain, A. García-Alonso, Survey on virtual coaching for older adults, Health
Informatics Journal 26 (2020) 3231–3249. URL: https://doi.org/10.1177/1460458220926894. doi:10.
1177/1460458220926894, pMID: 32744137.</p>
      <p>[9] W. R. Miller, S. Rollnick, Motivational interviewing: Helping people change, Guilford Press, 2012.
[10] M. Droppa, H. Lee, Motivational interviewing: A journey to improve health, Nursing 44 (2014).</p>
      <p>doi:10.1097/01.NURSE.0000443312.58360.82.
[11] K. Jokinen, Is the plan ready yet? - Exploring LLMs when talking about well-being and health, in:
14th International Workshop on Spoken Dialogue Systems Technology (IWSDS 2024), Sapporo,
Japan, 2024.
[12] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih,
T. Rocktäschel, S. Riedel, D. Kiela, Retrieval-augmented generation for knowledge-intensive NLP
tasks, in: Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS
2020), Vancouver, Canada, 2020, pp. 9459–9474.
[13] Welcome to World of LLMs - Session 2 - Knowing RAGS, Youtube video https://www.youtube.</p>
      <p>com/watch?v=Q6shHtNFiuQ, 2024.
[14] R. Browne, T. Ogawa, J. Vizcarra, K. Jokinen, Experiments on the Nao robot as a virtual coach, in:
Annual Conference of the Japanese Society for Artificial Intelligence (JSAI 2022), Kyoto, Japan,
2022.
[15] ITU, Definition of "Open Standards", https://www.itu.int/en/ITU-T/ipr/Pages/open.aspx, 2005.
[16] H. Bunt, J. Alexandersson, J.-W. Choe, A. C. Fang, K. Hasida, V. Petukhova, A. Popescu-Belis,
D. Traum, ISO 24617-2: A semantically-based standard for dialogue annotation, in: N. Calzolari,
K. Choukri, T. Declerck, M. U. Doğan, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, S. Piperidis
(Eds.), Proceedings of the Eighth International Conference on Language Resources and Evaluation
(LREC’12), European Language Resources Association (ELRA), Istanbul, Turkey, 2012, pp. 430–437.</p>
      <p>URL: http://www.lrec-conf.org/proceedings/lrec2012/pdf/530_Paper.pdf.
[17] T. Williams, C. Matuszek, K. Jokinen, R. Korpan, J. Pustejovsky, B. Scassellati, Voice in the Machine:
Ethical Considerations for Language-Capable Robots, Communications of the ACM 66 (2023)
20–23.
[18] G. Wilcock, K. Jokinen, To err is robotic; to earn trust, divine: Comparing ChatGPT and
knowledge graphs for HRI, in: 32nd IEEE International Conference on Robot and Human Interactive
Communication (RO-MAN 2023), Busan, Korea, 2023, pp. 1396–1401.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Jokinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Homma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsumoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Fukuda</surname>
          </string-name>
          ,
          <article-title>Integration and interaction of trustworthy AI in a virtual coach: An overview of EU-Japan collaboration on eldercare</article-title>
          , in: Y.
          <string-name>
            <surname>Takama</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Matsumura</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Yada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Matsushita</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Katagami</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Abe</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Kashima</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hiraoka</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Uchiya</surname>
          </string-name>
          , R. Rzepka (Eds.),
          <source>Advances in Artificial Intelligence: Selected papers from the Annual Conference of the Japanese Society for Artificial Intelligence (JSAI</source>
          <year>2021</year>
          ), Springer,
          <year>2022</year>
          , pp.
          <fpage>190</fpage>
          -
          <lpage>200</lpage>
          . doi: https: //doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -96451-1_
          <fpage>17</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>McTear</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jokinen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Alam</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Saleem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Napolitano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Szczepaniak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hariz</surname>
            , G. Chollet,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lohr</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boudy</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Azimi</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          <string-name>
            <surname>Roelen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Wieching</surname>
          </string-name>
          ,
          <article-title>Interaction with a virtual coach for active and healthy ageing</article-title>
          ,
          <source>Sensors</source>
          <volume>23</volume>
          (
          <year>2023</year>
          ). URL: https://www.mdpi.com/1424-8220/23/5/2748. doi:
          <volume>10</volume>
          .3390/s23052748.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Naccarelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. D'Agresti</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          <string-name>
            <surname>Roelen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Jokinen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Casaccia</surname>
            ,
            <given-names>G. M.</given-names>
          </string-name>
          <string-name>
            <surname>Revel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Maggio</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Azimi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Alam</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Saleem</surname>
            ,
            <given-names>A. H.</given-names>
          </string-name>
          <string-name>
            <surname>Mohammed</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Napolitano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Szczepaniak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hariz</surname>
            , G. Chollet,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lohr</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boudy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Wieching</surname>
          </string-name>
          , T. Ogawa,
          <article-title>Empowering smart aging: Insights into the technical architecture of the e-VITA virtual coaching system for older adults</article-title>
          ,
          <source>Sensors</source>
          <volume>24</volume>
          (
          <year>2024</year>
          ). URL: https://www.mdpi.com/1424-8220/24/2/638. doi:
          <volume>10</volume>
          .3390/s24020638.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Torres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tenorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cordasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gonzalez-Fraile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Esposito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Kornes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chollet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fernandez-Ruinova</surname>
          </string-name>
          ,
          <article-title>Ageing in the 21st century in europe: social challenges and innovation opportunities to support elderly independency and wellbeing, 2021</article-title>
          . URL: https://addi.ehu.es/ bitstream/handle/10810/50944/EMPATHIC_book_Chapter.pdf,
          <source>technical Report EHU-ADDI.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Callejas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Griol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Benghazi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Noguera</surname>
          </string-name>
          , G. Chollet,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Torres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Esposito</surname>
          </string-name>
          ,
          <article-title>Measuring and fostering engagement with mental health e-coaches</article-title>
          ,
          <source>in: Companion Publication of the 2020 International Conference on Multimodal Interaction, ICMI '20 Companion</source>
          , Association for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>275</fpage>
          -
          <lpage>279</lpage>
          . URL: https://doi.org/10.1145/3395035. 3425316. doi:
          <volume>10</volume>
          .1145/3395035.3425316.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>