<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IDRE: AI Generated Dataset for Enhancing Empathetic Chatbot Interactions in Italian language.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simone Manai</string-name>
          <email>simone.manai@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Gemme</string-name>
          <email>l.gemme@softjam.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Zanoli</string-name>
          <email>zanoli@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Lavelli</string-name>
          <email>lavelli@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLiC-it 2024: Tenth Italian Conference on Computational Linguistics</institution>
          ,
          <addr-line>Dec 04 - 06, 2024, Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>38123 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Lutech-Softjam</institution>
          ,
          <addr-line>16148 Genova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Trento</institution>
          ,
          <addr-line>38123 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper introduces IDRE (Italian Dataset for Rephrasing with Empathy), a novel automatically generated Italian linguistic dataset. IDRE comprises typical chatbot user utterances in the healthcare domain, corresponding chatbot responses, and empathetically enhanced chatbot responses. The dataset was generated using the Llama2 language model and evaluated by human raters based on predefined metrics. The IDRE dataset offers a comprehensive and realistic collection of Italian chatbot-user interactions suitable for training and refining chatbot models in the healthcare domain. This facilitates the development of chatbots capable of natural and productive conversations with healthcare users. Notably, the dataset incorporates empathetically enhanced chatbot responses, enabling researchers to investigate the effects of empathetic language on fostering more positive and engaging human-machine interactions within healthcare settings. The methodology employed for the construction of the IDRE dataset can be extended to generate sentences in additional languages and domains, thereby expanding its applicability and utility. The IDRE dataset is publicly available for research purposes.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Empathy</kwd>
        <kwd>LLMs</kwd>
        <kwd>Llama2</kwd>
        <kwd>Dataset</kwd>
        <kwd>Chatbot</kwd>
        <kwd>Healthcare1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Emotional intelligence has been widely recognized
as a crucial factor influencing human communication,
impacting aspects such as behavioral choices and the
interpretation of information [1]. Consequently, there
has been a growing interest in developing chatbots
capable of exhibiting empathetic responses [2] [3] [4].
While significant strides have been made in this
direction, the integration of empathy into commercial
chatbots remains challenging due to the rigid
constraints imposed by business rules such as the
response must not lose the original meaning and the
dialogue must maintain structure.</p>
      <p>To address this limitation, one possible approach is
to build a layer that rephrases the bot's response by
increasing empathy without altering the structure or
meaning of the underlying dialogue. This strategy offers
the potential to enhance user experience and create a
foundation for more sophisticated empathetic dialogue
systems.</p>
      <p>To facilitate the development of such systems, a
robust dataset containing empathetic responses is
essential. Despite the increasing body of research on
emotion recognition and generation in
humancomputer interaction, there is a notable absence of
publicly available datasets specifically focused on
empathy in chatbot interactions.</p>
      <p>This paper introduces the IDRE dataset, a new
Italian language resource comprising human-bot
interactions within the healthcare domain. The dataset
is available publicly, and the address is provided in the
Online Resource section. The dataset includes the user
questions, original bot responses and corresponding
empathetic reformulations for a total of 480 sentences,
providing a valuable foundation for research and
0000-0002-7175-6804 (A. Lavelli); 0000-0003-0870-0872 (R.
Zanoli)
© 2024 Copyright for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
development in empathetic chatbot technology, see
Table 1 for an example. The paper also elaborates on the
methodology employed for dataset generation,
highlighting its applicability to diverse domains and
languages.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>The development of empathetic chatbots capable of
understanding and responding to human emotions
represents a research area of growing interest [5].
However, building such systems requires high-quality
datasets that include examples of human-machine
interactions with empathic components.</p>
      <p>Despite the growing availability of datasets for
machine learning and natural language processing, the
lack of resources dedicated specifically to empathetic
Italian-language chatbots represents a significant
challenge.</p>
      <p>There are datasets that contain emotional
information, such as [6] [7] [8] [9]. However, these
resources focus primarily on labelling words or
sentences with generic emotions and do not provide the
context for complex, nuanced conversational
interactions like those required for developing
empathetic chatbots.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>This chapter details the methodology employed for
the construction of the IDRE dataset and outlines the
evaluation process implemented.</p>
      <p>The dataset created consists of 480 sentences and
roughly 18k total tokens divided as follows: 2k for the
question, 7k for the bot's response and 9k for the
response with empathy.
3.1. Dataset Creation</p>
      <p>The IDRE dataset comprises triplets of sentences, the
first sentence represents a user query, the second
sentence is the corresponding response generated by a
chatbot, and the third sentence is a transformed version
of the second sentence intended to enhance its
empathetic tone.</p>
      <p>The sentence generation process was done by the
Llama2 13B language model [11], operating on an Azure
Virtual Machine equipped with four NVIDIA Tesla V100
GPUs. The choice of Llama 2 was motivated by its
opensource nature, which allowed flexible and
providerindependent access.</p>
      <p>The dataset generation process consists of two
phases as illustrated in Figure 1:</p>
      <p>QnA Sentence Generation: To ensure the
generation of empathetic and compassionate responses,
the healthcare domain was selected as the focus for the
initial set of bot-human sentence pairs. This domain,
characterized by sensitive topics, is well-suited for
evaluating the model’s ability to generate empathetic
responses.</p>
      <p>The thirteen specific topics chosen for the sentence
pairs were invented for the purposes of the experiment:
'information on breast cancer', 'breast cancer
prevention', 'therapies for breast cancer', 'psychological
support after a cancer diagnosis', 'life expectancy after a
cancer diagnosis', 'psychological support after surgery',
'hospital admissions', 'post-operative care', 'information
on leukemia', 'psychological support', 'anti-cancer
therapies', 'information on stroke', and 'preparation for
surgeries.'</p>
      <p>An initial set of bot-human sentence pairs was
generated using the Llama2 model. These pairs
simulated a typical chatbot interaction concerning a
specific health issue or domain. For instance, a human
query such as "What are the symptoms of COVID-19?"
would elicit a corresponding chatbot response like "The
most common symptoms of COVID-19 are fever, dry
cough, and tiredness".</p>
      <p>Empathy Enhancement: After the generation of
the initial sentence pairs, an empathy enhancement
process was undertaken. Leveraging the Llama2 model
once more, the chatbot responses were modified to
convey a more empathetic tone. This was achieved by
prepending expressions of concern or appreciation, and
by substituting specific words to engender a supportive
demeanor. To illustrate, the aforementioned chatbot
response could be transformed into "I understand that
you’re concerned about COVID-19. Some common
symptoms include fever, dry cough, and fatigue".</p>
      <p>Both prompts are included in the Appendix.</p>
      <sec id="sec-3-1">
        <title>3.2. Evaluation Methodology</title>
        <p>To ensure the quality of the generated sentences, a
rigorous evaluation process was implemented. Twelve
volunteer annotators from Lutech-Softjam, experienced
IT developers and project managers with a solid
understanding of chatbot domain, participated. Despite
lacking prior experience in linguistic annotation, their
familiarity with chatbots significantly accelerated the
evaluation process. Before start, they underwent
comprehensive training on the evaluation task.</p>
        <p>Each evaluator was assigned 70 sentences for
assessment. To ensure diverse evaluations, 40 sentences
were unique to each evaluator and used for dataset
creation, while 30 common sentences were evaluated by
all evaluators, solely for measuring agreement and will
not be part of the dataset. This approach ensured that
each sentence received focused evaluation while also
providing a consistent assessment across evaluators.</p>
        <p>The evaluation process involves the administration
of a metric-specific question, which requires a response
on a scale of 1 to 5.</p>
        <sec id="sec-3-1-1">
          <title>The rating scale used is the following:</title>
          <p>Bot sentence correctness: measures the
absence of spelling, grammatical, or
punctuation errors in the question and the bot’s
answer. The question used is: “Il testo della
risposta con empatia è corretto sia dal punto di
vista grammaticale che semantico.”
Absence of English words in bot sentences:
checks if there are any words or sentences in
English within the sentences generated by the
model. The question used is: “Nel testo della
domanda dell’utente e della risposta del bot
(colonne QUESTION e ANSWER) non sono
presenti parole o frasi in lingua inglese, a meno
che non siano di uso comune in italiano (ad
esempio “badge”, “sport”, ecc.)”
Empathic answer correctness: measures the
absence of spelling, grammatical, or
punctuation errors in the bot’s answer with the
insertion of empathy. The question used is: “Il
testo della risposta con empatia è corretto sia
dal punto di vista grammaticale che semantico.”
Absence of English words in empathic
sentences: checks if there are any words or
sentences in English within the sentences with
empathy generated by the model. The question
used is: “Nel testo della risposta con empatia
non sono presenti parole o frasi in lingua
inglese, a meno che non siano di uso comune in
italiano (ad esempio “badge”, “sport”, ecc.)”
Semantic coherence: measures if the bot’s
answer and the bot’s answer with empathy are
semantically similar. The question used is: “La
risposta con empatia ha lo stesso significato
semantico della risposta del chatbot. Non ci
sono concetti mancanti o contraddittori”
Empathy increase: measures if the bot’s
answer with empathy has an effective increase
of empathy compared to the bot’s answer. The
question used is: “La frase nella colonna
ANSWER WITH EMPATHY esprime più
empatia rispetto alla frase nella colonna
ANSWER”</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Dataset Analysis</title>
      <p>This section analyses data quality by examining both
the distribution of agreement scores and the level of
inter-annotator agreement (IAA).</p>
      <p>Due to a limited pool of available evaluators, the
dataset was constrained to 480 annotated sentences.
These sentences were evenly distributed among 12
volunteers, each assessing 40 sentences (excluding the
30 sentences used for measuring agreement). This
approach was made to ensure the quality of the
annotations while preventing evaluator fatigue.
Nevertheless, a more in-depth analysis reveals that 223
sentences, equal to 46.5% of the total, have the score
grater or equal to 3 on all the metrics considered. This
means that these sentences were judged to be of high
quality in every aspect analysed. This subset of data can
be used to finetune language models.</p>
      <p>To obtain a more robust analysis and less subject to
small variations, the annotation categories were
grouped into three macro-categories: scores 1 and 2,
score 3 (neutral) and score 4 and 5.</p>
      <p>The analysis of sentences with lower score (1 and 2)
revealed three key factors: grammatical errors, the
presence of non-Italian words and lack of a significant
increase in empathy as shown in Figure 2.</p>
      <p>Grammatical Errors: A substantial portion of
sentences with lower score exhibited grammatical errors
(words in red). This highlights the importance of
incorporating robust grammar checks during the
generation process. Example: “Ohimini, cara/o utente, è
comprensibile che durante il trattamento del tumore possa
esserti difficile gestire i sintomi. Sono qui per aiutarti a
trovare soluzioni e supporti per farcela insieme”.
“Ohimini” is a made-up word and “supporti” contains a
typo.</p>
      <p>Non-Italian Words: the lower score sentences
frequently included non-Italian words (words in red),
primarily English. This deviation from the dataset’s
focus on Italian-language interactions can be attributed
to the underlying multilingual language model, which
was predominantly trained on English text. This
highlights the need for improved language model
training to prioritize Italian vocabulary. Example: “Per
prevenire le infezioni after surgery, è importante seguire le
istruzioni del medico e del personale ospedaliero, come ad
esempio lavare le mani frequentemente, evitare di toccare
la ferita e utilizzare dispositivi di protezione individuali.”</p>
      <p>Lack of a significant increase in empathy:
Among the lower score sentences (173, representing
36%), the transformed responses (indicated by the blue
and orange columns) did not exhibit a significant rise in
empathy or indecision compared to the original chatbot
responses. This suggests that further refinement of the
empathy-enhancing techniques might be necessary.
English words in bot sentences. In contrast, the
distribution for “Empathy increase” or “Bot sentence
correctness” is more dispersed across the entire range of
possible scores, suggesting a greater degree of variability
in annotator assessments of bot empathy increase.</p>
      <p>The observed disparity in distribution patterns
between the metrics can be attributed to the inherent
nature of the annotation tasks. The task of identifying
the absence of English words in bot sentences is
relatively straightforward and objective, leading to a
higher degree of agreement among annotators. On the
other hand, assessing bot empathy increase involves a
more subjective judgment of factors such as
grammatical accuracy, coherence, and relevance,
resulting in a wider range of annotations.</p>
      <p>The same behaviour can be noticed with metric
“Empathic answer correctness”.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Conclusion</title>
      <p>In this work, we have presented the creation of a
dataset of sentences representing typical interactions
with a healthcare chatbot. The dataset includes both user
input sentences and empathetic responses generated by
the chatbot. Human validation has confirmed the quality
and usefulness of the dataset for developing and
evaluating empathetic chatbots in the healthcare
domain.</p>
      <p>This work presents a two-pronged contribution to
the field of empathetic chatbots, specifically focusing on
the Italian language.</p>
      <p>Firstly, it addresses the critical issue of data scarcity
by providing a high-quality, annotated dataset for
training and evaluating empathetic chatbots within a
healthcare context. This dataset can be employed to
finetune large language models (LLMs) such as Llama2,
enabling them to generate responses with demonstrably
enhanced empathetic qualities. The limitations of
nonfine-tuned models are exemplified through the
observation that they can produce factually incorrect or
unempathetic sentences (e.g., " Il tuo corpo è vulnerabile
al rischio del tumore al seno a causa della tua età
avanzata, nonostante la tua vitalità e forza interiori. La
storia familiare di tumori al seno nella tua famiglia e la
tua condizione di obesità possono aumentare il rischio,
come pure l'abuso di tabacco e alcool. Inoltre, la tua scelta
di non avere figli o di averli dopo l'età di 35 anni può
aggiungere ulteriore rischio al tuo corpo."). By leveraging
the proposed dataset and selecting sentences with
demonstrably high empathy scores, a targeted training
set can be constructed specifically for this purpose. This,
in turn, allows for the fine-tuning of the LLM,
significantly improving its ability to generate
empathetic responses in a healthcare setting.</p>
      <p>Secondly, the work contributes a rigorous human
validation methodology for evaluating the effectiveness
of empathy expression in chatbots. This methodology
provides a valuable tool for researchers and developers
working in this domain.</p>
      <sec id="sec-5-1">
        <title>5.1. Future Work</title>
        <p>In the future, we intend to expand the work in two
main directions:</p>
        <p>Domain expansion: We will explore the creation
of similar datasets for other domains, such as customer
service or education, to assess the applicability of our
approach in different contexts.</p>
        <p>Comparison of language models: We will
conduct a comparative study to evaluate the
performance of different language models in generating
empathetic chatbot responses. This study will allow us
to identify the most suitable language model for this
specific task.</p>
        <p>We believe that this work represents an important
step towards the development of empathetic chatbots
capable of offering a more natural and engaging user
experience, especially in sensitive contexts such as
healthcare.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The authors would like to express their sincere
gratitude to Lutech-Softjam colleagues for their
invaluable contributions to the evaluation of the dataset
used in this study. Their expertise and meticulous work
in assessing the dataset’s quality and relevance were
instrumental in ensuring the robustness of our research
findings.</p>
      <p>downloaded
at</p>
    </sec>
    <sec id="sec-7">
      <title>A. Online Resource</title>
      <sec id="sec-7-1">
        <title>The dataset can be https://github.com/smanai/idre</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>B. Appendix</title>
      <p>Below the prompts used for both steps of dataset
creation are shown.</p>
      <p>Prompt for QnA Sentence Generation: """genera
{} coppie di domande utente e risposta di un assistente
virtuale.</p>
      <p>Le domande devono essere in lingua italiana e
rappresentano frasi tipiche una persona che vuole
informazioni nel dominio "{}".</p>
      <p>Le risposte sono quelle di un tipico chatbot di un call
center di un'azienda ospedaliera.</p>
      <p>Le risposte devono solo esporre dei fatti oggettivi e
scientifici ma prive di empatia.</p>
      <p>la struttura del output deve essere:
#
utente:
assistente:"""</p>
      <p>Prompt for Empathy Enhancement: """La
seguente frase è la risposta di un chatbot di un call center
di un ospedale ad una persona che richiede informazioni.
La frase è informativa, ma non trasmette empatia per la
situazione della persona che chiama. Puoi modificare la
seguente frase aggiungendo l'empatia mancante?</p>
      <p>Puoi modificare la frase aggiungendo testo o
modificandolo ma deve mantenere lo stesso significato
semantico.</p>
      <p>la frase modificata deve essere scritta in lingua
italiana.</p>
      <p>Non devi scrivere altro testo oltre alla frase
trasformata.</p>
      <p>inizia la modifica della frase con il carattere "-" come
in un elenco puntato.</p>
      <p>"""</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Fellous</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jean-Marc</surname>
            and
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Arbib</surname>
          </string-name>
          ,
          <article-title>"Who needs emotions?: The brain meets the robot</article-title>
          .," Oxford University Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Z.</given-names>
            <surname>Emmanouil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Paraskevopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Katsamanis</surname>
          </string-name>
          and
          <string-name>
            <surname>A. Potamianos.</surname>
          </string-name>
          ,
          <article-title>"EmpBot: A T5-based Empathetic Chatbot focusing on Sentiments,"</article-title>
          <source>arXiv preprint arXiv:2111</source>
          .
          <fpage>00310</fpage>
          .,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Jamin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Madotto</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <article-title>"Generating empathetic responses by looking ahead the user's sentiment.," in ICASSP 2020-</article-title>
          2020 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Mao</surname>
          </string-name>
          , LiangjunWang,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ruwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gou</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <article-title>"An emotion-based responding model for natural language conversation,"</article-title>
          <source>Springer Science+Business Media</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Q.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>"A Dynamic Emotional Session Generation Model Based on Seq2Seq and a Dictionary-Based Attention Mechanism," Appl</article-title>
          . Sci., p.
          <fpage>10</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          ,
          <article-title>"MultiEmotions-It: a New Dataset for Opinion Polarity and Emotion,"</article-title>
          <source>in Proceedings of the Seventh Italian Conference on Computational Linguistics</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>S. M. Mohammad</surname>
          </string-name>
          ,
          <article-title>"Practical and ethical considerations in the effective use of emotion and sentiment lexicons,"</article-title>
          <source>arXiv preprint arXiv</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Welivita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>"A LargeScale Dataset for Empathetic Response Generation,"</article-title>
          <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , pp.
          <fpage>1251</fpage>
          -
          <lpage>1264</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Rashkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Boureau</surname>
          </string-name>
          ,
          <article-title>"Towards empathetic open-domain conversation models: A new benchmark and dataset,"</article-title>
          arXiv preprint arXiv:
          <year>1811</year>
          .00207,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Welivita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>"A LargeScale Dataset for Empathetic Response Generation,"</article-title>
          <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , pp.
          <fpage>1251</fpage>
          --
          <lpage>1264</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almahairi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Babaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bashlykov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Bhosale and others, "Llama 2: Open foundation and fine-tuned chat models,"</article-title>
          <source>arXiv preprint arXiv:2307.09288</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Gadanho</surname>
          </string-name>
          ,
          <article-title>"Learning BehaviorSelection by Emotions and Cognition in a Multi-Goal Robot Task</article-title>
          .,
          <source>" Journal of Machine Learning Research</source>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>385</fpage>
          -
          <lpage>412</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>