<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Robotic Dietitian with Adaptive Linguistic Style</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Hannes Ritschel, Kathrin Janowski, Andreas Seiderer and Elisabeth AndrØ Human-Centered Multimedia, Augsburg University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>9638</volume>
      <fpage>5</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>This work outlines a concept and the necessary building blocks for creating a persuasive and personalized robotic dietitian for everyday health-related support based on existing technology and recent research insights. Key components include natural language generation for the social robot's linguistic style, mobile sensing hardware for tracking nutrition, and machine learning for adaptation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Figure 1 illustrates the general idea. Information about the human’s food consumption and activity throughout
the day can be acquired with mobile and stationary hardware automatically, such as a smartphone, smartwatch
or tness tracker. When the user is on the go, he might use a traditional smartphone application’s GUI or a
virtual agent (e.g. an animated 3D model of the robot) for entering information and getting advice. Interaction
in the user’s domestic environment benets from technology which can be installed stationary. This includes the
social robot and additional sensors, such as a smartscale [SFA17, RSJ +18]. Both the problem of identifying the
meal and estimating its amount needs to be solved before additional information from a nutrition database can
be used to calculate the nutritional value and match the user’s food consumption with the diet plan.</p>
      <p>In order to provide personalized diet support we propose a machine learning approach: the learning agent’s
goal is that the user adheres to the diet plan, interpreting deviations as failure. Since we focus on the robot’s
linguistic style in this work, the robot can e.g. explore dierent politeness strategies to generate the most
persuasive message for the individual user. RL is of special interest for this task since it allows to explore the
robot’s most ecient behavior autonomously. Based on the diet plan, acquired data from sensors, the user’s
activity and the meals’ properties, a reward needs to be calculated. This positive or negative scalar indicates
whether the robot’s last action was expedient or not, so that the robot’s linguistic style can be personalized to
the user’s reactions over time. Finally, the adaptation approach decides how to present the information. The
NLG component generates the corresponding utterances and sends them to the robot. Additionally, multimodal
cues, such as corresponding gaze behavior or facial expression can be added to emphasize the spoken language.
2.1</p>
      <p>Nutrition Logging
Gathering information about the type and amount of consumed food is essential for the robot’s advice and
adaptation process. Other data, such as the calorie amount or the intake of specic nutrients can roughly be
derived from this information with food databases. For behavioral analysis, the context in which food is consumed
might be helpful (e.g. in the evening while watching TV), which might be extractable from smart home technology
to a certain degree. The specic user’s requirements based on demographic and health information (gender, age,
weight, height, illnesses, medication) must be encoded by the diet plan. Additionally, the calorie consumption
should be estimated e.g. by using data of a tness tracker.</p>
      <p>Several technical possibilities exist to sense the type of food. Most of them just work in specic use cases and
usually a combination is required to allow a mostly complete automatic process. For example, image recognition
is able to detect many types of food [MBO +18] as long as they are not pureed. In such cases gas sensors [DSA18]
might be a better choice, nevertheless they can easily be disturbed by other odors. For both recognition methods
the detection of self prepared food is a problem hard to solve if the preparation of the food was not logged. In
such uncertain cases it might be a good choice to ask the user in the currently most convenient way as long as
there is no perfect solution for unobtrusive, fast, mobile, automatic chemical food analysis available.</p>
      <p>If the type of food is known, optical systems can roughly estimate the amount (weight), taking into account
the vessel in which the food is located. One challenge is that usually not all parts of the food are visible to
the camera depending on the perspective. A mobile scale [SFA17] can be a solution for this problem if higher
precision is required, which however involves more eort than simply taking a photo.</p>
      <p>One of the biggest problems is to detect the context in which a person eats or drinks. For behavioral analysis
it might be sucient to know how consciously a person eats or drinks the food as this is a major problem. Eye
tracking is one technical option to give hints in this regard.
disagreeable
cold
dominant
extraverted</p>
      <p>warm
agreeable
introverted</p>
      <p>submissive</p>
      <p>After automatically sensing the user’s nutrition, generating persuasive and personalized advice is the next task
of the robotic dietitian. Apart from the actual message content, the way in which it is formulated and presented
to the human plays an important role. Its expressed personality can be reected in its multimodal output.</p>
      <p>Interaction behavior is typically classied using the Interpersonal Circumplex [DWQP13]. It is dened by the
two dimensions status and aliation , with the former ranging from submissive to dominant and the latter from
cold to warm. Alternatively, the same relationships can be expressed through the personality traits extraversion
and agreeableness, which can be found at approximately 20 to 45 degrees relative to the other pair of axes
[DWQP13]. Extraversion thus corresponds to a combination of high status and high aliation.</p>
      <p>Oakman et al. [OGC03] suggest that the Interpersonal Circumplex dimensions are also related to the politeness
theory by Brown and Levinson. The so-called negative face is a person’s desire to have autonomy with regards to
their actions, while positive face is the desire to have others approve of one’s own goals. Positive politeness , which
minimizes threats to somebody’s positive face, can be mapped to the aliation dimension while the presence or
absence of negative politeness roughly corresponds to status.</p>
      <p>When looking at the robot’s linguistic style, these relationships imply that extraverted persons are less
concerned with threats to another person’s negative face, but more inclined to apply positive politeness strategies
such as treating the other person as a member of the same group. Conversely, introverted persons are more distant
and submissive, and therefore avoid threats to the other party’s autonomy while being less likely to use positive
politeness. Figure 3 compares dierent phrasings for a simple example suggestion with regard to the expressed
status and aliation. With the exibility of NLG the robot’s dietary advice can be tweaked and formulated to
increase its persuasiveness. Adapting the robot’s politeness has recently been explored for a domestic robotic
companion in the context of health-related recommendations based on template-based utterances [RSJ +19]. In
contrast, NLG is a promising option due to the complexity of the diet context at hand.
2.3</p>
      <p>Adaptation Process
The machine learning approach uses insights about the user’s actual nutrition in comparison to the diet plan to
improve the robot’s behavior. RL can be used as a framework for optimizing details in the robot’s linguistic style.
For example, the robot’s expressed politeness can be modeled as a nonstationary k-armed bandit problem [SB18],
which is a reduced form of RL. The goal is to nd the most eective of k actions A (politeness strategies) by
estimating each action’s value Q, which is calculated based on a scalar feedback, the so-called reward R. In each
time step t the agent selects an action At 2 A, executes it, receives a reward Rt and updates the action’s new value
Qt+1 based on Rt, the old value Qt and constant learning rate 2 [0; 1]: Qt+1(At) = Qt(At) + [Rt Qt(At)].</p>
      <p>In order to react to changes in the user’s preferences, Upper Condence Bound (UCB) action selection [SB18]
can be used for balancing exploitation and exploration, i.e., the agent’s choice of the greedy (best) action with the
highest Q-value versus exploring another supposedly suboptimal one. Therefore, Nt(a) is the number of times
action a already has been executed while c &gt; 0 is a constant for controlling exploration. Based on this information
the agent selects actions not only depending on their estimated values Q but also with regard to its uncertainty
about the fact that their value might have changed in the meantime: At = arg max hQt(a) + cq Nltn(ta) i.
a</p>
      <p>By calculating the reward R based on the user’s actual nutrition and the diet plan, the learning approach can
optimize the robot’s generated behavior over time by expressing itself in the most persuasive manner.
Our concept illustrates an approach for building a robotic dietitian, which personalizes its linguistic style to the
individual user. With the ultimate goal of supporting the human’s diet, persuasive messages are produced by a
Natural Language Generation component, which enriches the robot’s advice with personality-derived
characteristics. Building on recent research, the reward for a reinforcement learning component is calculated based on the
user’s diet plan and its actual nutrition, making it possible to optimize the robot’s messages for the individual
user. We outlined necessary technologies to track the user’s nutrition based on mobile and stationary sensor
technology in an intelligent environment. All in all, we expect the robot to become more persuasive over time
and thus foster a healthy lifestyle.</p>
    </sec>
    <sec id="sec-2">
      <title>Acknowledgments</title>
      <p>This research was funded by the Bavarian State Ministry for Education, Science and the Arts (StMWFK) as
part of the ForGenderCare research association.
[AM18]</p>
      <p>Luca Anselma and Alessandro Mazzei. Designing and testing the messages produced by a virtual
dietitian. In Proceedings of the 11th International Conference on Natural Language Generation ,
pages 244253. Association for Computational Linguistics, 2018.</p>
      <p>Chi Tai Dang, Andreas Seiderer, and Elisabeth AndrØ. Theodor: A step towards smart home
applications with electronic noses. In Proceedings of the 5th international Workshop on Sensor-based
Activity Recognition and Interaction, iWOAR 2018 , pages 11:111:7. ACM, 2018.
[MBO+18] Javier Marn, Aritro Biswas, Ferda Oi, Nicholas Hynes, Amaia Salvador, Yusuf Aytar, Ingmar
Weber, and Antonio Torralba. Recipe1m: A dataset for learning cross-modal embeddings for cooking
recipes and food images. CoRR, abs/1810.06553, 2018.
[OGC03]</p>
      <p>Jonathan Oakman, Shannon Giord, and Natasha Chlebowsky. A multilevel analysis of the
interpersonal behavior of socially anxious people. Journal of Personality , 71(3):397434, 2003.
Hannes Ritschel, Tobias Baur, and Elisabeth AndrØ. Adapting a robot’s linguistic style based on
socially-aware reinforcement learning. In 26th IEEE International Symposium on Robot and Human
Interactive Communication, RO-MAN 2017 , pages 378384. IEEE, 2017.</p>
      <p>Hannes Ritschel, Andreas Seiderer, Kathrin Janowski, Ilhan Aslan, and Elisabeth AndrØ.
Drink-omender: An adaptive robotic drink adviser. In Proceedings of the 3rd International Workshop on
Multisensory Approaches to Human-Food Interaction , pages 3:13:8. ACM, 2018.</p>
      <p>Hannes Ritschel, Andreas Seiderer, Kathrin Janowski, Stefan Wagner, and Elisabeth AndrØ.
Adaptive linguistic style for an assistive robotic health companion based on explicit human feedback. In
Proceedings of the 12th ACM International Conference on PErvasive Technologies Related to
Assistive Environments , pages 247255. ACM, 2019.</p>
      <p>Richard S. Sutton and Andrew G. Barto. Reinforcement Learning - An Introduction (Second Edition) .
Adaptive Computation and Machine Learning. MIT Press, 2018.</p>
      <p>Andreas Seiderer, Simon Flutura, and Elisabeth AndrØ. Development of a mobile multi-device
nutrition logger. In Proceedings of the 2nd ACM SIGCHI International Workshop on Multisensory
Approaches to Human-Food Interaction, MHFI@ICMI 2017 , pages 512. ACM, 2017.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>[DWQP13] Colin G. DeYoung</surname>
          </string-name>
          , Yanna J.
          <string-name>
            <surname>Weisberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Lena C. Quilty</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jordan</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Peterson</surname>
          </string-name>
          .
          <article-title>Unifying the Aspects of the Big Five, the Interpersonal Circumplex, and Trait Aliation</article-title>
          .
          <source>Journal of Personality</source>
          ,
          <volume>81</volume>
          (
          <issue>5</issue>
          ):
          <fpage>465475</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>