<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CONVERSA: Efective and Eficient Resources and Models for Transformative Conversational AI in Spanish and Co-oficial Languages</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Griol</string-name>
          <email>dgriol@ugr.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ksenia Kharitonova</string-name>
          <email>ksenia@ugr.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Pérez-Férnandez</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asier Gutiérrez-Fandiño</string-name>
          <email>asier@lhf.ai</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zoraida Callejas</string-name>
          <email>zoraida@ugr.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro de Investigación en TIC de la Universidad de Granada (CITIC-UGR)</institution>
          ,
          <addr-line>Granada</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dpto. de Lenguajes y Sistemas Informáticos, Universidad de Granada</institution>
          ,
          <addr-line>Granada</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LHF Labs</institution>
          ,
          <addr-line>Bilbao</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidad Autónoma de Madrid</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Access to information is increasingly conversational. However, there is a lack of conversational AI training material for Spanish and the co-oficial languages, in general and for specific key tasks and domains. Additional barriers include steep computational costs for training conversational agents and challenging inference times, and a lack of guarantees for the safety and transparency of conversational systems. The CONVERSA project (TED2021-132470B-I00) constitutes a step forward to democratize access to conversational AI through computation and data eficient development and testing of innovative, open and safe resources in Spanish and co-oficial languages. The duration of the project is from December 2022 to November 2024.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Conversational AI</kwd>
        <kwd>conversational systems</kwd>
        <kwd>dialogue systems</kwd>
        <kwd>corpus</kwd>
        <kwd>datasets</kwd>
        <kwd>language models</kwd>
        <kwd>open access</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>and e-government [7].</p>
      <p>Recent market reports [8, 9], as well as relevant surveys
The term conversational artificial intelligence, coined re- [7] and both academic and industrial vision statements on
cently in academic research, refers to Natural Language the future of conversationally enabled applications [10]
Processing (NLP) technologies, such as dialog systems, highlight how the use of natural language (and speech
chatbots or intelligent virtual assistants, with which users interaction) is changing the way people connect to the
can engage in a conversation in natural language and ar- information that they need [11].
tificial intelligence techniques are extensively used [ 1, 2]. This has been accentuated during the Covid-19
panThese systems provide a low barrier entry for users, en- demic with the highly successful increased use of
Converabling them to access information, interact in an intuitive sational AI for e-government [12, 13, 14] and customer
way with services, resources, and data on the Internet, purchases [15, 16]. There are predictions by 2024 that
as well as with their surrounding environment [3, 4]. consumer retail spending via chatbots will reach $142</p>
      <p>The development of conversational AI has seen rapid billion worldwide, up from just $2.8 billion in 2019 [9].
progress in recent years, both text- and voice-based, en- Revenue in NLP is estimated to increase from $3.2 billion
abled by large pretrained language models and a number in 2017 to more than $43 billion in 2025 [17].
of new consumer-facing applications and intelligent de- Since the 2000s the emphasis in the scientific spoken
divices, such as personal mobile assistants [5, 6], social alogue systems community has moved from handcrafted
networks, messaging applications, and intelligent speak- systems (symbolic and logic-based AI) to data-driven
sysers. Application domains have increased dramatically tems using machine learning. Machine-learning
techranging from retail, telecommunication, finance, health niques avoid specifying dialogue state machines and
make it possible to address the inconveniences derived
from unexpected patterns that slot filling approaches
cannot anticipate. Key to the development of efective
chatbased conversational AI technology using this paradigm
is the availability of a large volume of training material
[1, 18].</p>
      <p>These new systems rely on high-quality
datasets/corpora for the training of deep-learning
algorithms to develop precise models. The preparation
of a high-quality gold standard corpora for natural
language processing on a large scale is a challenging about users does not change frequently over time and
task due to the need of data cleaning, accurate language are required by one or more tasks. Being able to exploit
identification models, and precise content parsing the history/record for several dialogues would make the
tools. Transformer-based models and self-supervised system more efective and enhance the user experience.
learning mechanisms have shown promising results in Hence, monolithic dialogues negatively afect the quality
key NLP tasks and academia and industry are currently of interactions and users’ satisfaction, since the same
developing large transformer-based linguistic models repetitive questions/answers pairs are followed for every
[1, 19]. user.</p>
      <p>Due to their size, the training and adjustment for A more elaborated case of this kind of situation -that
the implementation of these conversational services cur- we have experienced in past projects- occurs when it
rently requires large computational resources and expert can be inferred (possibly following a machine-learning
knowledge for their optimization [20, 21]. Current mod- process) that particular responses to a question at some
els have a large number of parameters and are trained point allow some subcomponents of a dialogue to be
with huge collections of training examples. GPT-1 had skipped. However, once the first model is constructed
117 million parameters to work with, GPT-2 had 1.5 bil- and deployed, the experience gained over time cannot be
lion. GPT-3 used for its training, among others, 410 bil- exploited, without constructing a new complete model.
lion crawled data tokens, 67 billion book tokens, and 3 Again, data and knowledge sharing between diferent
billion Wikipedia tokens. The full version of OpenAI conversational AI applications could also help boost
GPT-3 has around 175 billion trainable parameters. GPT- learning processes from past experiences and training
3.5 is significantly larger, with a staggering 355 billion data.
parameters. GPT-4 has been trained on a large amount A related challenge is to ensure system proactivity,
of internet content and it is able to handle more com- that is, the ability of the system to start the conversation
plex instructions and produce higher-quality long-form at adequate moments or to show unsolicited behaviours
writing. Bard, powered by Google’s Language Model for and contents that may be helpful for the users. This
reDialogue Applications (LaMDA), was released in 2021 quires having accurate knowledge about the user and the
with 137 billion parameters trained using 1.56 trillion context of the interaction, which involves being able to
words of public dialog data and web text. identify the user, to identify whether the user is actively</p>
      <p>Regarding open-source alternatives, the LLaMA involved in the conversation (and not for example
someproject has encompassed a set of foundational language one just present in the environment), manage turn-taking
models that vary in size from 7 billion to 65 billion param- and distinguishing between multiple topics.
eters and were trained on millions of tokens extracted
exclusively from publicly available datasets. Stanford
Alpaca claims that it can compete with ChatGPT and 2. Description and Objectives of
the training can be completed with less than 600$. Vi- the Project
cuna is finetuned from the LLaMA model on user-shared
conversations collected from ShareGPT. GPT4ALL is a A key aspect of digital transformation is to be able to
encommunity-driven project and was trained on a massive gage, serve, and empower the users it is directed to. This
curated corpus of assistant interactions, including code, implies that digital services must be accessible, provide
stories, depictions, and multi-turn dialogue. open access to reliable information and open interfaces</p>
      <p>In addition, the black box nature of the neural net- for businesses and citizens, make full use of existing
works that these techniques require, makes it dificult online services with more agile discovery and
personalto explain the reasons behind the behavior of a conver- ization, ensure ease of use, and guarantee security and
sational agent at diferent points of the conversation, privacy.
undermining user-system trust and making it dificult to Access to information is increasingly conversational
tackle changes and evolve the system reliably. In addition, and task-oriented chat systems help users achieve their
topic detection and switching is challenging in particular goals eficiently using natural language, ensuring
accesfor end-to-end neural dialogue systems, as it is dificult sibility and personalized omnichannel interaction,
fosto incorporate long-distance context information. tering user-system trust through explainable social
dia</p>
      <p>There are still remaining open issues that concern flow- logues. However, there are three main barriers to digital
ing/navigating through dialogues within a conversational transformation in Spain through the democratic adoption
AI system. One example of it is the monolithic structure of conversational AI technology by a wide spectrum of
that dialogues exhibit and their single-block granular- technological and societal stakeholders:
ity (i.e., without dialogue sub-components), so that it
is not possible to architect dialogues with several en- • A lack of training material for Spanish and the
trance points. This is particularly desirable when data co-oficial languages, in general and for specific
key tasks and domains. Not undertaking this endeavor would only generate a
• Steep computational costs for training conversa- dependence from the solutions ofered by the major tech
tional agents and challenging inference times. companies. At present, state-of-the-art conversational AI
• A lack of guarantees for the safety and trans- technology, resources and development tools are mostly
parency of conversational systems. owned by big tech companies (e.g., Google DialogFlow,</p>
      <p>Google Assistant, Amazon Alexa Skills Kit, IBM
Wat</p>
      <p>The CONVERSA project (TED2021-132470B-I00) ad- son Assistant, or Microsoft Bot Framework). However,
dresses these challenges by addressing the following re- many public and private stakeholders are keen to take
search questions: a more autonomous position concerning the
technol• How can we automatically generate, simulate ogy to interact with citizens and customers for strategic
and transfer training data to produce engaging (competitiveness) reasons, economic reasons, and legal
chat-based conversations in Spanish and the coo- reasons. In particular, accessibility and non-disclosure of
ifcial languages? By using data-driven technol- customer/citizen data to third parties is key.
ogy to create better performing chat-based tech- To be engaging and efective, modern conversational
nology, public institutions and companies can systems are computationally intensive and data hungry,
outsource more citizen/ customer interaction to requiring large-scale, language-specific, domain-specific,
neural network-based agents. and task-specific training data. CONVERSA will address
• How can we develop and adapt chat-based conver- the development of efective methods and tools for
dosational systems in a computationally and data- main adaptation and data augmentation for
conversaeficient manner? By further developing and im- tional AI, in order to be fast and deployable in lightweight
proving neural network-based models for con- computing environments, and overcome the current
limversational systems for less resource-intensive ited transferability of task knowledge between tasks.
scenarios, we will directly contribute to the way Another important aspect is to enhance trust and
encomputers can communicate with humans. sure the reliability of the processes. Conversational
in• How can we do all of this securely and transpar- terfaces to digital services pose the risk of harming users
ently, without violating privacy, and with pro- by conversing with them inappropriately, revealing
senvisions for explainability and data provenance? sitive private information about them or incurring in
Accountability also implies security transparency bias (gender and racial bias implicitly learned from
auand explainability. Moreover, explainability is a tomatically crawled language resources). A particular
fundamental aspect of knowledge systems, which challenge for conversational systems is that they often
is also enforced in European regulations. directly confront users and so the impact can be
immediate. CONVERSA specifically targets the development of</p>
      <p>The Multilingual Technology Alliance and the META- explainability techniques to prevent harmful utterances,
NET Network of Excellence Europe’s Languages in the both safety and privacy-wise, and design trustworthy
Digital Age, highlight that the lack of natural language conversational systems more resilient against malicious
processing resources for European languages are the manipulations.
most significant impediment that must be overcome for CONVERSA has the general objective of
democratizfurther conversational AI adoption. Although Spanish ing access to conversational AI through computation
is currently the second language in the world by num- and data eficient development and testing of
innovaber of native speakers, it is not a well-resourced lan- tive, open and safe resources in Spanish and co-oficial
guage in terms of the core technologies and datasets languages. The specific objectives are:
needed to build state-of-the-art language-based solutions.</p>
      <p>This problem is even more accentuated in the case of • To create and annotate open-access multi domain
co-oficial languages. Spain has a clear opportunity to dialog corpora in Spanish and the co-oficial
lanlead the development of openly accessible digital ser- guages to train the specific models required to
vices in Spanish and the co-oficial languages that can develop neural-bases AI conversational systems.
be exploited not only in Spain but also in the multiple • To develop compute- and data- eficient neural
countries with Spanish-speaking population. architectures and open-access models for
con</p>
      <p>Through CONVERSA, the generation of high-quality versational AI, including methods and tools for
and accessible labeled data and pretrained models will domain adaptation and data augmentation.
drive the development, optimization and deployment of • To develop open source tools for corpus creation,
conversational AI for Spanish and co-oficial languages, data debugging, bias analysis, and privacy
analyfueling the generation and adoption of transformative sis.
conversational AI solutions in an environmentally re- • To provide a replicable evaluation of the project
sponsible way.
3. Scientific and Technical Impact
outcomes with particular attention to replicabil- implementations will be available as open-source
softity, bias avoidance and privacy preservation. ware and open corpora, and indirectly because the
devel• To showcase the use of the resources and mod- oped methods are in principle technology independent.
els generated through the development and de- CONVERSA results have wide implications for
techployment of an initial pilot demonstrator of con- nology areas (human-computer interaction, artificial
inversationally enabled innovative e-government telligence, natural language processing, conversational
services. assistants, user modelling, service orchestration) and can
revolutionize the development of conversational systems
for Spanish and co-oficial languages. Impact will be
pursued at multiple levels:
CONVERSA is meant to facilitate the following changes
in the way that public and private industrial end users
develop and deploy conversationally enabled applications:
• Change 1: Faster prototyping and retraining of</p>
      <p>conversationally enabled applications
• Change 2: Shift development of conversationally
enabled applications from a supervised learning
paradigm to a transfer learning-based paradigm;
• Change 3: Reduced development times and
ecologically pernicious efects (waste or computation
resources) of conversationally enabled
applications for novel domains, tasks, and scenarios;
• Change 4: Improve customer trust by means of
explanations of actions taken by a chat-based
conversational assistant</p>
      <p>In particular, the outcomes that will facilitate these
changes are:
• Outcome 1: Eficient implementations and
adaptation of conversational technology for new
domains
• Outcome 2: Enriched domain-specific language
models for the Spanish and co-oficial languages
retail and service domains
• Outcome 3: Better coverage of stakeholders’
prod</p>
      <p>uct bases in conversational agents
• Outcome 4: Open data and open-source
conver</p>
      <p>sational technology
• Outcome 5: Improved user satisfaction and user</p>
      <p>trust in conversational interaction
• Outcome 6: Development of R&amp;D technology in</p>
      <p>new contexts</p>
      <p>The algorithmic methods and insights developed will
be published as open-source software, and the data and
corpora will be also openly available. Validated versions
of the packages will be integrated in the open-source
technology of Rasa Technologies, that is available for any
organization to use and implement in their own service
and retail environments.</p>
      <p>The project outcomes therefore impact society directly
and indirectly: directly because the resulting models and
• Information Retrieval (IR): Conversational search
is an important research direction in information
retrieval and question answering applications. By
addressing the domain-specific challenges of
conversational intelligence for e-goverment
applications, sales and service-oriented chatbots, we
hope to give a boost to the Spanish and
cooficial languages conversational IR community with
novel methods, models, datasets, and evaluation,
especially for less-resourced environments.
• Natural Language Processing (NLP): by utilizing
and adapting neural-based architectures and
generating tools to this domain, we expect to
develop new transformational models, along several
dimensions (output quality, memory eficiency,
and inference time). Designing novel methods
for knowledge-based, multilingual, cross-domain
adaptation for conversational AI is expected to
transfer learning from transformer-based
linguistic modeling, due to the specific challenges of this
application.
• Trusted AI (i.e., human-centric AI) and ML
security/privacy: Our goal is to develop techniques
to avoid harmful utterances, non-disclosure of
private data and through the prevention of bias
in corpus and generator models (such as gender,
race or ideology).</p>
      <sec id="sec-1-1">
        <title>The international impact of CONVERSA is based, not</title>
        <p>only on the profound implications of its results, but also
on the solid international production of the members of
the research team, their well-established international
connections with ongoing H2020 projects and the
dissemination and communication activities envisaged.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Acknowledgments</title>
      <sec id="sec-2-1">
        <title>This publication is part of the “CONVERSA: Efective</title>
        <p>and eficient resources and models for transformative
conversational AI in Spanish and co-oficial languages”
project with reference TED2021-132470B-I00, funded by
MCIN/AEI/10.13039/501100011033 and by the European
Union “NextGenerationEU"/PRTR”.
[12] A. Androutsopoulou, N. Karacapilidis, E. Loukis,</p>
        <p>Y. Charalabidis, Transforming the
communica[1] M. McTear, Conversational AI. Dialogue systems, tion between citizens and government through
Conversational Agents, and Chatbots, Morgan and AI-guided chatbots, Government
InformaClaypool Publishers, 2020. doi:10.1007/978-3- tion Quarterly 36 (2019) 358–367. doi:10.1016/
031-02176-3. j.giq.2018.10.001.
[2] M. McTear, Z. Callejas, D. Griol, The Conversational [13] J. Cabot, Chatbots y asistentes de voz, una
oportuInterface: Talking to Smart Devices, Springer, 2016. nidad en la gestión de crisis sanitarias, The
Condoi:10.1007/978-3-319-32967-3. versation (2020).
[3] P. Cañas, D. Griol, Z. Callejas, Towards versa- [14] D. Griol, D. Pérez Fernández, Z. Callejas,
Hispabottile conversations with data-driven dialog manage- Covid19: the oficial Spanish conversational system
ment and its integration in commercial platforms, about Covid-19, in: Proc. of IberSPEECH 2021,
Journal of Computational Science 55 (2021) 101443. Valladolid, Spain, 2021, pp. 139–142. doi:10.21437/
doi:10.1016/j.jocs.2021.101443. IberSPEECH.2021-30.
[4] T. Fu, S. Gao, X. Zhao, J. rong Wen, R. Yan, [15] L. Gkinko, A. Elbanna, Designing trust: The
Learning towards conversational AI: A sur- formation of employees’ trust in conversational
vey, AI Open 3 (2022) 14–28. doi:10.1016/ AI in the digital workplace, Journal of
Busij.aiopen.2022.02.001. ness Research 158 (2023) 113707. doi:10.1016/
[5] S. Lee, R. Jha, Zero-shot adaptive transfer for j.jbusres.2023.113707.</p>
        <p>conversational language understanding, in: Proc. [16] I. U. Jan, S. Ji, C. Kim, What (de) motivates
cusof AAAI’19 Conference on Artificial Intelligence, tomers to use AI-powered conversational agents
Honolulu, Hawaii, USA, 2019, pp. 6642–6649. for shopping? The extended behavioral
reasondoi:10.1609/aaai.v33i01.33016642. ing perspective, Journal of Retailing and
Con[6] A. Rastogi, X. Zang, S. Sunkara, R. Gupta, P. Khai- sumer Services 75 (2023) 103440. doi:10.1016/
tan, Towards scalable multi-domain conversational j.jretconser.2023.103440.
agents: The schema-guided dialogue dataset, in: [17] Statista, Revenues from the natural language
Proc. of AAAI’20, New York, NY, USA, 2020, pp. processing (NLP) market worldwide from 2017 to
8689–8696. doi:10.1609/aaai.v34i05.6394. 2025, https://es.statista.com/estadisticas/1130045/
[7] C. Gao, W. Lei, X. He, M. de Rijke, T.-S. Chua, Ad-
mercado-global-de-procesamiento-de-lenguajevances and Challenges in Conversational Recom- natural/, 2022. Acceded: June 2023.
mender Systems: A Survey, AI Open 2 (2021) 100– [18] P. Su, N. Mrksic, I. Casanueva, I. Vulic, Deep
learn126. doi:10.1016/j.aiopen.2021.06.002. ing for conversational AI, in: Proc. of the 2018
[8] Gartner, Making sense of the chatbot Conference of the North American Chapter of the
and conversational AI platform market, Association for Computational Linguistics, New
Orhttps://www.gartner.com/en/documents/ leans, Louisiana, USA, 2018, pp. 27–32.
3993709/making-sense-of-the-chatbot-and- [19] S. Young, Hey Cyba. The Inner Workings of a
conversational-aiplatfo, 2020. Acceded: June 2023. Virtual Personal Assistant, Cambridge University
[9] BusinessInsider, Chatbot market in 2021: Stats, Press, 2021.</p>
        <p>trends, and companies in the growing AI chatbot in- [20] A. Gutiérrez-Fandiño, D. Pérez-Fernández,
dustry, https://www.businessinsider.com/chatbot- J. Armengol-Estapé, D. Griol, Z. Callejas,
esmarket-stats-trends, 2021. Acceded: June 2023. Corpius: A Massive Spanish Crawling Corpus,
[10] Y. K. Dwivedi, N. Kshetri, L. Hughes, E. L. Slade, in: Proc. IberSPEECH 2022, 2022, pp. 126–130.</p>
        <p>A. Jeyaraj, et alt., Opinion Paper: So what if doi:10.21437/IberSPEECH.2022-26.
ChatGPT wrote it? Multidisciplinary perspectives [21] J. Ni, T. Young, V. Pandelea, X. F., E. Cambria,
Reon opportunities, challenges and implications of cent advances in deep learning based dialogue
sysgenerative conversational AI for research, prac- tems: a systematic survey, Artificial Intelligence
tice and policy, International Journal of Informa- Review (2023) 3055–3155.
doi:10.1007/s10462tion Management 71 (2023) 102642. doi:10.1016/ 022-10248-8.</p>
        <p>j.ijinfomgt.2023.102642.
[11] J. Wei, Y. Tay, R. Bommasani, C. Rafel, B. Zoph,</p>
        <p>S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou,
D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals,
P. Liang, J. Dean, W. Fedus, Emergent abilities of
large language models, Transactions on Machine
Learning Research (2022).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>