<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Japanese Dialect-aware Chatbot: Adapting NLP Models for Japanese Dialect Variation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kinga Lasek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Ptaszynski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fumito Masui</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Text Information Processing Laboratory, Faculty of Engineering, Kitami Institute of Technology</institution>
          ,
          <addr-line>165 Koen-cho, Kitami-shi, Hokkaido, 090-8507</addr-line>
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <fpage>30</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>Chatbots are said to soon become the most frequently used consumer channel service in the world. Users' expectations are growing, especially regarding the usage of the natural language. In this study, we present our attempts of training NLP models for Japanese dialect variation with the aim of creating Japanese dialect-aware chatbot in the future. Firstly, we describe the current situation on chat-bot market. Secondly, we focus on describing attempts that have already been made to create chatbots sensitive to dialects and low-resourced languages. Furthermore, the process of collecting data needed for our research is presented. Finally, evaluation results after model training are shown and described. The results reveal key insights into the models' strengths and limitations in understanding Japanese dialects. We discuss noted occurrences and finally outline directions for our future work.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Natural language processing (NLP)</kwd>
        <kwd>Japanese dialects</kwd>
        <kwd>dialect-specific models</kwd>
        <kwd>chatbot</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>the correctness of the answers given and the understanding of the input text were identified as the
biggest problems. It led company carrying out the survey to conclusion that technical aspects directly
related to response content, such as natural language processing and chatbot response accuracy need
to be improved.</p>
      <p>The outline of the paper is as follows. Firstly, we discuss the current state of chatbot technology and
actions taken so far to build chatbots sensitive to dialects. Then, we shortly describe collecting data
process and how we utilized it. What is more, we explore one of the fundamental steps essential for
creating Japanese dialect-aware chatbot, meaning training models assumed to handle Japanese dialects.
Results of experiments are presented and discussion points are highlighted.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Chatbots today</title>
        <p>
          Around 2027, chatbots are supposed to become the dominant customer service channel for approximately
25% of organizations 3. Moreover, the global chatbot market size was calculated at USD 5,132.8 million
in 2022 and is predicted to achieve a compound annual growth rate (CAGR) of 23.3% from 2023 to 20304.
Such a scale indicates the improvements taking place in the world of Artificial Intelligence (AI) as well
as Natural Language Processing (NLP)‘s realm. Technologies that enhance the knowledge of chatbots
encompass, among others, deep learning, machine learning and NLP [
          <xref ref-type="bibr" rid="ref10">9</xref>
          ]. The mutual understanding
of the system and humans makes it possible to reach customer satisfaction. The better and
errorfree the responses of the chatbot, the greater the customer’s trust. Additionally, it is supposed that
users’ expectations will grow as the chatbots upgrade their communication abilities or even so-called
social skills [
          <xref ref-type="bibr" rid="ref11">10</xref>
          ]. For a long time, research on chatbots and their development focused mainly on
appropriate grammar structures and quick responses. Still, the design of chatbots faces many problems,
including biases, (lack of) emotional involvement, ethical issues or keeping privacy [
          <xref ref-type="bibr" rid="ref10">9</xref>
          ]. Posed dilemmas
involve the need for new improvements, so research into the capabilities of chatbots is also being
expanded. In consequence, chatbots‘ system design and language design are in the spotlight. Therefore,
psycholinguistic, sociolinguistic, dialectometry approaches and so forth have been taken. Previous
research has proved that variation within a language often comes from factors such as individual
author/speaker style, dialect or genre [
          <xref ref-type="bibr" rid="ref11">10</xref>
          ], so scholars took steps to bring machine language closer
to human-like style and make them more personalized and attractive to potential users. Elsholz et
al. [
          <xref ref-type="bibr" rid="ref12">11</xref>
          ] conducted an experiment using two chatbots to sell theater tickets for a Shakespeare play:
one communicating in modern English and one in a Shakespearean-style dialect. The results were
thought-provoking for researchers as modern chatbot reached a higher usability score, whereas scores
regarding interest in play were similar for both chatbots. This highlights the necessity of developing
chatbot‘s responses that will be as close to humans as possible. Marcellus Amadeus et al. [
          <xref ref-type="bibr" rid="ref6">5</xref>
          ] also pointed
out that: integrating language variation into conversational AI will build near-real language inventories
and boost user engagement. Users might use more or less formal language, its vernacular variations,
dialect or slang. Realization of language variety leads to expanding research. Scholars as well as big
companies and organizations started paying more attention to the conversational style of chatbots
[
          <xref ref-type="bibr" rid="ref13">12</xref>
          ] and exploring their abilities of participating in discussions conducted in diverse language aspects.
Among slowly developing enhancements are chatbots/voicebots able to understand and provide an
answer in low-resourced languages and a wide range of dialects. Thanks to this improvement, users will
not feel obligated to use the standard form of their language, but will input information expressed in
phrases most similar to those they use every day. Therefore, among papers about English and Chinese
chatbots, one can find also research on conversational agents and derivatives focused on, for example,
Vietnamese [
          <xref ref-type="bibr" rid="ref14">13</xref>
          ], Irish [
          <xref ref-type="bibr" rid="ref15">14</xref>
          ] or Indian languages [
          <xref ref-type="bibr" rid="ref16">15</xref>
          ]. Big steps regarding chatbot technology were also
taken in the Arabic world [
          <xref ref-type="bibr" rid="ref17 ref4">16</xref>
          ].
3https://www.gartner.com/en/newsroom/press-releases/2022-07-27-gartner-predicts-chatbots-will-become
-a-primary-customer-service-channel-within-five-years [access: 2024/04/15].
4https://www.grandviewresearch.com/industry-analysis/chatbot-market [access: 2024/04/15].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Existing dialect-aware and low-resource language-aware chatbots: overview</title>
        <p>
          Researchers have been working on developing chatbots that are dialect-aware and capable of handling
under-resourced languages to improve their performance and usability across diverse linguistic
communities. We cannot yet speak of a significant number of them, however, on the basis of those available
we can observe certain trends and draw conclusions. Chatbots created to understand language varieties
in linguistically rich countries or communities are coming to the fore. Among 21 countries in the Arab
world, each has its own dialect. In the case of Arabic, nine dialectal categories are distinguished and
each of them has more varieties depending on the particular city or town [
          <xref ref-type="bibr" rid="ref18">17</xref>
          ]. Probably the most known
Arabic dialect chatbots are Nabiha [
          <xref ref-type="bibr" rid="ref19">18</xref>
          ] and Botta [
          <xref ref-type="bibr" rid="ref20">19</xref>
          ]. Nabiha, developed by Dana Al-Ghadhban and
Nora Al-Twairesh, is able to understand and respond in the Saudi Arabic dialect. In the early stages,
authors collected, preprocessed and classified data into several text files in order to build a dialogue
corpus. Authors admitted that despite good first evaluation, Nabiha should be based on a bigger dataset.
Botta, another example of Arabic chatbots, was created earlier, in 2016. It speaks in Egyptian Arabic
(Cairene) dialect. Botta‘s persona – female chatbot – was supposed to simulate friendly conversation
and connect with as many Arab users as possible. Both Botta and Nabiha were built using artificial
intelligence markup language (AIML) and launched on the Pandorabots platform. However, Botta also
possesses some unique sets, being able to: detect dialectal bad language that can be ofending, separate
sets of months based on each dialect it identifies or words that indicate the dialect of the users.
        </p>
        <p>
          Boulesnane et al. [
          <xref ref-type="bibr" rid="ref21">20</xref>
          ] created a communication platform for patients in need of consultation.
They named it DZchatbot and developed it with a focus on the Algerian Arabic dialect. Diferently
from Botta and Nabiha, this chatbot was hinged on the sequence-to-sequence model (seq2seq) with
RNN encoder and decoder. Larger, more detailed healthcare chatbot, MedicalBot [
          <xref ref-type="bibr" rid="ref22">21</xref>
          ], was proposed
through application of three deep learning techniques, but any particular dialect was itemized. Scholars
mentioned, though, the importance of the multitude of dialects. Among Arabic dialect-aware chatbots
one can also mention a medical chatbot for Tunisian dialect [
          <xref ref-type="bibr" rid="ref23">22</xref>
          ] or chatbot system focused on dialogue
acts, which was built in order to cope with Levantine Arabic dialect [
          <xref ref-type="bibr" rid="ref24">23</xref>
          ].
        </p>
        <p>
          In India, despite the common perception that English is mainly spoken there, other so-called 21
modern Indian languages are used and English is spoken by only ten percent of India population [
          <xref ref-type="bibr" rid="ref16">15</xref>
          ].
Among Indian languages one can also observe a quite wide scope of chatbots successfully communicating
in its variations. For instance, Golpo, which can converse in Bengali, was developed with a focus on
storytelling and engaging users with narrative-based interactions [
          <xref ref-type="bibr" rid="ref16">15</xref>
          ]. On the contrary, Doly [
          <xref ref-type="bibr" rid="ref25">24</xref>
          ]
was supposed to be a support for Bengali-speaking users. The development process involved training
Doly on a diverse dataset of Bengali conversations to enhance its linguistic capabilities and dialect
recognition. It can provide the appropriate response thanks to the list using Naive Bayesian. However,
probably the most known Hindi chatbot is AskDISHA, which nowadays is available as AskDISHA2.0
(Digital Interaction To Seek Help Anytime)5 as it has been upgraded. Currently this chatbot, created
thanks to AI and machine learning, can answer its interlocutor both through text message and voice
message in Hindi, English and so-called Hinglish (macaronic hybrid use of Indian English and the
Hindustani language). Its main management is coping with railway reservations. Kaleem et al. [
          <xref ref-type="bibr" rid="ref26">25</xref>
          ]
faced the challenge of designing an Urdu conversational agent, which covers novel features such as
the Word Order Wizard (WOW) algorithm and scripting language in its architecture. They named it
Umair and developed it with the aim of applying it as customer service representative for Pakistan’s
National Database and Registration Authority (NADRA). To improve the relevance and coherence of
its responses, UMAIR may utilize context-aware techniques to understand the ongoing conversation
and maintain context over multiple turns. Some years later, Shabbir et al. [
          <xref ref-type="bibr" rid="ref27">26</xref>
          ] scrutinized whether and
how it is possible to automate the process of user‘s intent generation by using AI and deep learning
techniques so that human endeavor could be reduced. They made their experiment on grounds of
Roman Urdu and RASA Framework. They took into consideration two major factors of this particular
framework: RASA NLU, as it can perform intent classification and entity extraction from the training
dataset, and Dialog Management Model (DMM) that prepares the specific response according to intent.
A knowledge graph with RASA Framework has also been embedded to preserve the dialog history for
semantic based natural language mechanism for chatbot communication [
          <xref ref-type="bibr" rid="ref27">26</xref>
          ]. Brixey and Traum [
          <xref ref-type="bibr" rid="ref28">27</xref>
          ]’s
chatbot, Masheli, represents a significant advancement in exploiting technology to support language
revitalization eforts and cultural preservation. They used ChoCo, a Choctaw language corpus to create
the chatbot’s replies and later formed questions to place them in QA corpus. In order to build Covid-19
chatbot sensitive to African dialects Aymen Ben, Mabrouk et al. [
          <xref ref-type="bibr" rid="ref29">28</xref>
          ] gathered needed data and divided it
into two main categories: Frequently Asked Questions (FAQ) and chitchat. Chatbot possesses an ability
of providing answers in English, French, Arabic, Tunisian as well as spoken in Nigeria Igbo, Yoru‘ba
and Hausa. Sarma and Pathak [
          <xref ref-type="bibr" rid="ref30">29</xref>
          ] presented “Shiksha Mitra”, an artificial intelligence chatbot able to
reply to user queries in Assamese. Just as for chatbots using the best-known languages, the Artificial
Intelligence Markup Language (AIML) method is used for under-resourced language chatbots [
          <xref ref-type="bibr" rid="ref31">30</xref>
          ],
[
          <xref ref-type="bibr" rid="ref32">31</xref>
          ]. The AIML conversational agent operates on the principle of pattern matching, where responses
are generated through the mapping of keywords within each request to their corresponding patterns.
Utilizing the AIML Interpreter facilitates the pattern matching process between queries and responses.
Sandhini, Binu et al. [
          <xref ref-type="bibr" rid="ref32">31</xref>
          ] introduced Malayalam (Dravidian language spoken in the Indian state of
Kerala and assigned as classical language of India) chatbot. The comparison between AIML based
version of this chatbot and machine learning type of this chatbot was made. The performance of
machine learning based chatbot was better. Compared to AIML, machine learning requires substantial
amounts of high-quality training data to perform efectively.
        </p>
        <p>
          As the interest in language varieties is constantly growing [
          <xref ref-type="bibr" rid="ref33">32</xref>
          ], one can assume that deployment of
chatbots related to dialects and low-resource languages will also take on meaning. One of the most
serious challenges is scarcity of data. That is why among scholars, a corpus-driven approach can be
observed. Researchers collect large datasets of conversations in diferent dialects, annotate them with
linguistic features, and use them to train dialect-aware chatbots. Usually, due to a small amount of data,
creators of a chatbot must collect data themselves. This is what Al-Ghadhban and Al-Twairesh [
          <xref ref-type="bibr" rid="ref19">18</xref>
          ]
did to build Nabiha. It has been trained on a diverse corpus of Arabic conversations, encompassing
various dialectal variations, to learn from real-world linguistic data and adapt to diferent linguistic
contexts. Brixey and Traum [
          <xref ref-type="bibr" rid="ref28">27</xref>
          ] also chose this path to bridge the gap between Choctaw language
(native American language) and English. They used ChoCo, a Choctaw language corpus to create the
chatbot’s replies and later formed questions to place them in QA corpus. The methodology of Orosoo et
al.[
          <xref ref-type="bibr" rid="ref34">33</xref>
          ] was also to collect diverse linguistic dataset from diferent languages in order to propose system
that improves NLP in multilingual chatbots. Authors gave prominence to significant obstacles, such as
dificulties with low-resource languages, biases in training data and dynamic between language and
culture. The attempts of researchers mentioned above prove the complexity of studies on chatbots.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data Collection and Data Preprocessing</title>
      <sec id="sec-3-1">
        <title>3.1. Used corpuses and datasets</title>
        <p>In this study, a comprehensive data collection approach was employed, utilizing several key datasets,
including:
• Japanese Dialect Conversations Data6: includes transcripts of short dialogues in diferent</p>
        <p>
          Japanese dialects in diverse formality of language.
• Corpus of Japanese Dialects: COJADS7: a parallel corpus consisting of standard Japanese,
(hyo¯jungo) text and dialect texts as well as audio. It has been sorted according to 47 Japanese
prefectures. Among those corpuses and datasets that we decided to incorporate in our research,
COJADS has the the widest variety of dialectal data.
6Ho¯gen ro¯ru purei kaiwa deetabeesu [Japanese Dialect Conversations Data] http://hougen-db.sakuraweb.com/[access:
2024/06/27].
7Corpus of Japanese Dialects (COJADS). NINJAL. We used version published in March 2023.
https://www2.ninjal.ac.jp/cojads/index.html [access: 2024/06/27].
• The Corpus of Kansai Vernacular Japanese [
          <xref ref-type="bibr" rid="ref35">34</xref>
          ]: was divided according to the place
(OsakaKobe area, Kyoto area, towns of Takacho and Nishiwaki) consisted of four files. Among them,
one was created on the basis of interviews with international students living in the Kansai area.
Aiming for natural language used by Japanese people, we did not include students speech based
data in our research.
• Crowdsourced Parallel Speech Corpus of Japanese Dialects [
          <xref ref-type="bibr" rid="ref36">35</xref>
          ]: consists of parallel text
and speech data of 21 Japanese dialects. Dialogues for each dialect were short and included 250
sentences/turns.
• JMD: Japanese multi-dialect corpus8: that was made of audio and text versions of Osaka and
Kumamoto dialects. It was constructed so that each sentence was used as a separate entity with
a separate recording. Files for both dialects consist of 1300 lines of text and present the same
content.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data preprocessing</title>
        <p>By incorporating those, it was necessary to pay heed to the Japanese writing system. Japanese is written
as a combination of three types of symbols as well as roman letters and Arabic numerals. The oldest
symbols, kanji, came to Japan from China and were adapted to the Japanese language. These logographic
characters are used for conceptual words and names. Two syllabaries, hiragana and katakana, were
developed later. Hiragana‘s main function is recording words and grammatical endings that are not
written in kanji, while katakana serves as a method of writing words of foreign origin [36]. In COJADS,
mixed kanji and hiragana script marks text in standard form of Japanese. On the contrary, dialect texts
are written in katakana.
8Shinnosuke Takamichi, "JMD: Japanese multi-dialect corpus," https://sites.google.com/site/shinnosuketakamichi/publication/
research-topics/jmd_corpus [access: 2024/08/08].</p>
        <p>We decided to convert dialectal texts from the other four corpuses into katakana. Before that, we
adjusted data for our needs: Firstly, we decided not to include tagged parts of each sentence. For
instance, we cleaned the text from non-verbal sounds, like cough or laugh, that were originally indicated
by curly brackets as we focus on the text, not the audio version. The same was done to the hyo¯jungo
version. Consequently, we did the same with other corpus that treated non-verbal sounds in the same
way in the text, namely Japanese Dialect Conversations Data. For the same reason, we did not take into
consideration symbol /// that was supposed to highlight unknown meanings of words. Moreover, in
Japanese Dialect Conversations Data, one could find the arrows that indicated up or down intonation.
Those were also deleted, together with signs // which meant that the immediately following utterance
started at the same time as the next speaker’s utterance [37]. Some of the talks were held in standard
Japanese, however, it needs to be emphasized that each dialogue difers and we do not possess parallel
data. When it comes to the Corpus of Kansai Vernacular Japanese, we used a morphologically untagged
version of the texts. In JMD: Japanese multi-dialect corpus for speech synthesis neither tags nor special
signs were used. We only had to clear the text of some kanji readings (probably those considered as
hard), which were taken in brackets and written in the hiragana syllabary, and transform the text into
katakana.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Katakana Conversion</title>
        <p>For katakana conversion, we used MeCab [38], a highly optimized Japanese morphological analyzer.
It efectively processes mixed kanji and kana sentences, converting them into uniform katakana or
hiragana versions. It is believed to achieve a high eficiency at breaking down Japanese text into
individual morphemes, which we find essential for accurate katakana transcription. We started from
converting each character into katakana (e.g., お父さん became オトウサン, read: oto¯san). Then
we focused on transcription that can provide a standardized phonetic nuance representation (so here
お父さん became オトーサン), which is believed to be useful when dealing with various Japanese
dialects. We hope that this method will benefit in capturing the subtle diferences between accents and
pronunciation among dialects.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Data distribution</title>
        <p>In the next stages of our study, removing empty lines was needed. Possessing all data in their katakana
version, we were able to incorporate them into relevant classes. We have adhered to the division used
in COJADS to a large extent as grouping dialects by prefectures can provide a convenient framework
for study. However, due to the apparent dominance of hyo¯jungo and the Kansai dialect, we decided to
combine Nara, Kyoto, Osaka and Hyo¯go based texts into one data, as Kansai-ben. These dialects share
many features, which can make them more compatible for grouping. Due to scarcity of data, dialect
imbalances in our dataset were, for the time being, unavoidable. On the other hand, dataset constraints
made it impossible to capture dialects spoken in territories more extensive or narrower than prefectures
and between diferent social groups. For instance, Tsugaru dialect (hereafter, Tsugaru-ben), spoken on
the west side of Aomori prefecture 9, was not mentioned as specified group, but still covers some of the
data from class called Aomori. This choice, with all of its advantages and disadvantages, enabled us
to start testing. We also made a decision to include hyo¯jungo in our research. It was developed from
the educated middle-upper-class Tokyo dialect for the purpose of spreading a sense of a nation unity
within the country in the Meiji era (1868-1912) [39]. This denotes that hyo¯jungo is also a variation of a
language. It is believed that recognizing the standard language as one of many varieties, among that
each with separate rules and functions, enables scholars to perceive it within the broader context of
language variations, meaning dialects and social variants [40]. We also felt an obligation to discriminate
between hyo¯jungo and the rest of the data operated under the name of Tokyo. In COJADS, Tokyo was
treated purely territorially, but authors provided users information about particular places (namely:
district, villages, towns etc.). It could be easily observed that language used within a metropolis was
standardized, while language spoken in rural areas near Tokyo (like, for example, Hinohara village)
possessed noticeable diferences from it. That is why we draw a line between those two.</p>
        <p>Next stage consisted in making dialect distribution. We checked the number of lines (sentences/turns)
per each dialect area. The greatest number was observed in the Kansai dialect (more than 300,000 lines),
9See more here: https://tsugaruben.net [access: 2024/10/11].
followed by hyo¯jungo that noted only 175,849 lines due to the division we have applied. There is a
significant gap between those two classes and the rest of the groups. Remaining groups did not even
reach 100,000 lines or at least 50,000 lines. Third place went to Gunma, with 18,165 lines. The lowest
number of lines was recorded in Yamanashi, Saga, Nagasaki and Tottori, respectively. Tottori had less
than 200 lines. Next step involved an investigation of the average length of each line, calculating the
average length of sentence for each dialect and finally checking their standard deviation as it is known
as a universal and functional measure as it shows the average distance of each score from the mean.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental results</title>
      <p>First of all, in order to create a Japanese dialect-aware chatbot, we need a dialect identification model that
will be able to detect dialectal text written in katakana quickly and successfully. Texts are categorized
by dialect/area, and the neural network is trained on katakana representations for each dialect category.
In the experiments, for training we used 525,569 out of 656,826 sentences, while the rest of them was
left for testing. All four models were trained using learning rate 5e-05 and a batch size of 32. We set 3
as a number of epochs. In the first place, we trained four models with preprocessed data:
• tohoku-nlp/bert-base-japanese10: A BERT model pre-trained on a Japanese corpus and trained
on Japanese Wikipedia.
• tohoku-nlp/bert-large-japanese11: A bigger BERT model that contains around 30M sentences.
• tohoku-nlp/bert-large-japanese-char12: A character-level version of the larger BERT model,
trained with the whole word masking enabled for the masked language modeling (MLM) objective.
• tohoku-nlp/bert-large-japanese-v213: A word-level variant of the large BERT model, enriched
by the Japanese portion of CC-100 dataset that incorporates approximately 392M sentences alone.</p>
      <p>Our evaluation measures, namely Precision, Recall and F1 Score, are widely used for Text Classification
[41]. Their formulas demonstrate the significance of retrieval of positive examples in text classification.
Table 1 presents evaluation results for models mentioned above.</p>
      <p>All four models were based on Bidirectional Encoder Representations from Transformers (BERT)
[42] - language framework that deals with pretraining techniques such as Mask Language Modeling
(MLM) and Next Sentence Prediction (NSP) [43]. The tohoku-nlp/bert-base-japanese model entirely
outperforms others, reaching the highest overall F1 score. The tohoku-nlp/bert-large-japanese model
achieved, surprisingly, slightly lower scores in comparison to the base version, despite possessing more
parameters. Moreover, the tohoku-nlp/bert-large-japanese-char model performance is poor across
almost all dialects. It may suggest that the character-level approach is not efective for this task. The
tohoku-nlp/bert-large-japanese-v2 model showed some improvement over the character-level model.
However, it still noticeably shows small efectiveness compared to the base and large models.</p>
      <p>After getting unfavorable results, we decided to compare the performance of these with two other
models:
• google-bert_bert-base-multilingual-uncased14: trained on multilingual data, does not make
distinction between capital and lowercase letters. It was described in detail by Jacob Delvin et al.
[42]
• google-bert_bert-base-multilingual-cased15: almost identical model, but prepared to detect
diferences between uppercase and lowercase letters.</p>
      <p>Letter case might can heavily impact meaning in some languages. Texts in languages like Japanese,
Chinese or Korean usually do not have any spaces between words. This fact was considered during the
development process.</p>
      <p>The results are summarized in Table 2. Both models performed very well dealing with hyo¯jungo and
Kansai dialect. Strong performance could be also observed for Okinawian dialect, especially high F1
score for cased model. A score of over 0.50 was achieved by Aomori and Oita for both models. Both for
tohoku-nlp models as well as google-bert models, poor performance across all metrics can be noted
among Tottori, Tochigi, Shiga and Hiroshima regions. Most of the dialects can be found in the group of
dialects with visible moderate performance, rating F1 score between 0.20 and 0.40.</p>
      <p>Significant disparity between Precision and Recall might be observed within some dialects, like
Fukuoka (for models 1, 2, 5 and 6). Same thing applies to Gifu. Discrepancy is highly noticeable among
dialects with overall bad performance that mark very low precision and even lower recall, like Aichi,
Niigata, Kanagawa or Mie.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>Based on these results, we can observe a significantly strong correlation between results and amount of
data. Kansai dialect, with the largest amount of sentences, reaches the best results whereas low-sourced
dialects note scores close to 0. It indicates the need for more extensive data collection and consequently
searching for more training data in general. The google-bert_bert-base-multilingual-uncased model
generally performs better than the cased model, especially in dialects with lower performance. This
leads to an assumption that case insensitivity might be helpful in efectively handling diverse language
forms.
tohoku-nlp/bert-large-japanese-char-v2. Model 4: tohoku-nlp/bert-large-japanese-v2.
n 2 4 0 0 9 6 5 0 3 6 9 3 4 6 2 1 1 2 7 5 9 3 4 2 6 4 9 6 4 1 9 7 8 8 3 0 5 3 1 9 6 7 7
iso .70 .10 .50 .70 .20 .10 .20 .50 .50 .40 .50 .00 .20 .30 .50 .50 .10 .40 .80 .10 .30 .20 .30 .30 .30 .60 .50 .60 .50 .10 .10 .30 .50 .00 .10 .00 .20 .02 .05 .03 .02 .02 .03
/trceaA ¯jgooyun iichA itakA irooAm iabhC ieEhm ikFuu aokkFuu isakuhm ifuG aunGm iirsaohm iaookkd iIraakb iIsaakhw aaagKw isaagohm aaaagnw isaanK icohK iM izaayM aagN saagN iigN O aakyOm iaknwO agS itaaSm iSh iaShm izoSuh coTh isokuhm ttooT aoyTm aaakym caagum saaanm aoumm okT
e ik on iak taa ita a a a a a e a i h h t y
g n k ig a ir a a i i o o
le h Fu H H K K T
W Y Y K</p>
      <p>O
s 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 C
s</p>
      <p>1 1 1 1 1 1 1 1 1 1 2 2 2 2 2 2 2 2 2 2 3 3 3 3 3 3 3 3 3 3 4 4 S
a
l
b en 28 5 4 5 5 82 4 0 4 3 6 9 3 9 3 0 9 8 8 8
3
m t 6 1 5 7 7 1 6 4 3 2 3 5 9 9
u n 0 3 0 1 8 7 0 2 3 9 5 6 4 5 3 9 9 6 9 5 0 5 0 8 8 6 0 0 4 7 2 3 2 7 3 4 7 1 7 6 2 1
e 4 4 8 2 1 1 9 8 9 8 4 2 3 1 1 0 9 4 5 1 8 0 6 3 0 2 6 2 8 0 1 0 9 6 1 4 2 0 6 4 4 7
N s 1 9 1 6 7 1 3 5 1 2 1 2 1 7 8 1 1 3 2 1 2 1 4 5 1 9 9 2 5 1 1 1 9 8 1 1 2 8 9 7 9 7
multilingual- cased.</p>
      <p>Evaluation results. Model 5: google-bert_bert-base-multilingual-uncased, Model 6:
google-bert_bert-baseNumber of</p>
      <p>Class</p>
      <p>Dialect/Area</p>
      <p>Model 5</p>
      <p>Model 6</p>
      <p>An interesting case in an Okinawa region that reaches a good performance, despite the small amount
of data. The reasons for this may be found in many diferences between Japanese and Ryukyuan
languages. It is essential to notice that Ryukyuan languages have been downgraded to the status of
Japanese dialects [44] despite being mutually unintelligible in many aspects with standard Japanese.
Aleksandra Jarosz [45] as major features of Ryukyuan languages points out, among others, the raising
of short mid-close vowels and a phonological appearance of the glottal stop. The question is, whether
other models will achieve similar ratings for Okinawa or is it only used models specificity.</p>
      <p>Striking variability between Precision and Recall results was noticeable in some classes. Recall is
the proportion of Real Positive cases that are accurately Predicted Positive while Precision represents
the proportion of predicted positive cases that are correctly identified as true positives [ 46]. Imbalance
between them might suggest that while some models are fairly efective at recognizing instances they
can detect, they still fail to capture a significant amount of relevant dialect data. This situation might a
a consequence of a lack of suficient feature representation. Moreover, results achieved by some classes
like Aichi, Niigata or Mie indicate that models are often incorrect with their predictions. Their low
recall reflects inefectiveness in identifying the majority of actual instances.</p>
      <p>For certain dialects, specialization might be necessary since specialized models can be fine-tuned
with dialect-specific data, resulting in higher Precision, Recall, and F1 scores for dialects with poor
performance. This poses a dilemma of whether to create several specialized models for various dialects
or a single, highly generalized model.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>In our study we monitored the performance of some models working on Japanese dialects. We presented
the results and outlined conclusions. We also introduced some already existing dialect-aware chatbots
and chatbots that were made with the thought of low-rescourced languages. The data collection process
was also explained in detail.</p>
      <p>Models successfully dealt with standard Japanese and Kansai dialect, but showed an inability in
detecting less-resourced dialects. This contrast suggests that standard NLP models might have problems
in generalizing across dialectal features without adequate data representation. This stresses also the need
for more targeted training or extensive data collection. Moreover, it gave us a hint to try implementing
other learning rates and scrutinize what impact they will have on the model performance.</p>
      <p>Our study highlights the challenges in dialect detection due to data scarcity, emphasizing the
complexity of dialect detection within Japanese language. Facing mentioned challenges is crucial for advancing
NLP systems that can accurately reflect the rich linguistic landscape of Japan.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Future work</title>
      <p>In the future, besides learning rates implementation, we plan to check the performance of some other
models in order to have a comparison and choose the best performing model as a way to focus on its
application, namely Japanese dialect-aware chatbot. This paper revolved around dialect recognition,
but it cannot be excluded that in the future we will focus on style shifts within dialects. Style adaptation
might greatly influence the chatbot performance and users‘ impression. With the increasing need of
personalization [47], incorporating both dialect recognition and the ability to adjust for conversational
style can provide a more customized and responsive user experience.</p>
      <p>In the case of Japanese, formality’s level of the language plays an important role in customer service
and everyday relationships. Speaker‘s awareness of his position in particular situation in Japan is
reflected by honorific language, keigo [48]. However, among dialects keigo expressions can also difer.
On the other hand, there are also regions, like southern part of the Tohoku Region, that do not have the
linguistic forms corresponding to the honorific and humble forms of the national lingua franca, and
where honorific expressions are considered to be uncommon 16. Hence, profound research on honorific
forms among dialects will be needed. Taking into consideration the scarcity of dialectal data, for the
time being successful distinguishing between keigo expressions within dialects might be impossible.
Nonetheless, collecting dialectal data focused on honorific forms might be another goal to be set.
B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, S. Piperidis, T. Tokunaga (Eds.), Proceedings
of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018),
European Language Resources Association (ELRA), Miyazaki, Japan, 2018, pp. 434–437. URL:
https://aclanthology.org/L18-1067.
[36] W. Hadamitzky, M. Spahn, Japanese Kanji &amp; Kana:(JLPT All Levels) A Complete Guide to the</p>
      <p>Japanese Writing System (2,136 Kanji and 92 Kana), Tuttle Publishing, 2013.
[37] I. Fumiko, A Study of Regional and Generation Diferences in Discourse Pattern, Technical Report,
The National Institute for Japanese Language and Linguistics, 2014. Application/pdf, NINJAL
Collaborative Research Project Reports ; 13-04.
[38] Y. Sim, A morphological analyzer for japanese nouns, verbs and adjectives, ArXiv abs/1410.0291
(2014). URL: https://api.semanticscholar.org/CorpusID:15195083.
[39] S. Saito, M. Turner, Language diversity in “monolingual” japan: Language awareness among
high school teachers of english, Journal of Language, Identity &amp; Education (2024) 1–15. URL:
http://dx.doi.org/10.1080/15348458.2024.2385837. doi:10.1080/15348458.2024.2385837.
[40] R. Quirk, Language varieties and standard language, English Today 6 (1990) 3–10. doi:10.1017/</p>
      <p>S0266078400004454.
[41] M. Sokolova, G. Lapalme, A systematic analysis of performance measures for classification tasks,</p>
      <p>Information Processing &amp; Management 45 (2009) 427–437. doi:10.1016/j.ipm.2009.03.002.
[42] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: Pre-training of deep bidirectional transformers
for language understanding, in: J. Burstein, C. Doran, T. Solorio (Eds.), Proceedings of the 2019
Conference of the North American Chapter of the Association for Computational Linguistics:
Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational
Linguistics, Minneapolis, Minnesota, 2019, pp. 4171–4186. URL: https://aclanthology.org/N19-1423.
doi:10.18653/v1/N19-1423.
[43] C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, et al., A
comprehensive survey on pretrained foundation models: A history from bert to chatgpt, arXiv preprint
arXiv:2302.09419 (2023).
[44] P. Heinrich, S. Miyara, M. Shimoji, Handbook of the Ryukyuan Languages: History, Structure, and</p>
      <p>Use, 2015. doi:10.1515/9781614511151.
[45] A. Jarosz, Japonic languages: an overview, Silva Iaponicarum (2017). doi:10.14746/sijp.2017.</p>
      <p>41/42.6.
[46] D. M. Powers, Evaluation: from precision, recall and f-measure to roc, informedness, markedness
and correlation, arXiv preprint arXiv:2010.16061 (2020).
[47] T. Ait Baha, M. El Hajji, Y. Es-Saady, H. Fadili, The power of personalization: A systematic review
of personality-adaptive chatbots, SN Computer Science 4 (2023). URL: http://dx.doi.org/10.1007/
s42979-023-02092-6. doi:10.1007/s42979-023-02092-6.
[48] E. T. Rahayu, Japanese honorific language in various domains, in: Proceedings of the Fourth
Prasasti International Seminar on Linguistics (Prasasti 2018), Atlantis Press, 2018/08, pp. 25–34.
URL: https://doi.org/10.2991/prasasti-18.2018.5. doi:10.2991/prasasti-18.2018.5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bhirud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tataale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Randive</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nahar</surname>
          </string-name>
          ,
          <article-title>A literature review on chatbots in healthcare domain</article-title>
          ,
          <source>International journal of scientific &amp; technology research 8</source>
          (
          <year>2019</year>
          )
          <fpage>225</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Patil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence in financial services: Customer chatbot advisor adoption</article-title>
          ,
          <source>Int. J. Innov. Technol. Explor. Eng</source>
          <volume>9</volume>
          (
          <year>2019</year>
          )
          <fpage>4296</fpage>
          -
          <lpage>4303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Følstad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nordheim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bjørkli</surname>
          </string-name>
          ,
          <article-title>What makes users trust a chatbot for customer service? an exploratory interview study</article-title>
          ,
          <source>in: Internet Science</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>194</fpage>
          -
          <lpage>208</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>030</fpage>
          -01437-7_
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>16Keigo no shishin (Guide to keigo</article-title>
          ),
          <source>Bunkashingikai (Agency for Cultural Afairs)</source>
          ,
          <year>2019</year>
          : https://www.bunka.go.jp/seisaku/ bunkashingikai/kokugo/hokoku/pdf/keigo_tosin.pdf, [access:
          <year>2024</year>
          /10/8].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Chaves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Gerosa</surname>
          </string-name>
          ,
          <article-title>How should my chatbot interact? a survey on social characteristics in human-chatbot interaction design</article-title>
          ,
          <source>International Journal of Human-Computer Interaction</source>
          <volume>37</volume>
          (
          <year>2019</year>
          )
          <fpage>729</fpage>
          -
          <lpage>758</lpage>
          . URL: https://api.semanticscholar.org/CorpusID:102350801.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Amadeus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Homeli</surname>
          </string-name>
          da Silva,
          <string-name>
            <given-names>J. V.</given-names>
            <surname>Pessoa</surname>
          </string-name>
          <string-name>
            <surname>Rocha</surname>
          </string-name>
          ,
          <article-title>Bridging the language gap: Integrating language variations into conversational AI agents for enhanced user engagement</article-title>
          , in: N. HosseiniKivanani, S. Höhn,
          <string-name>
            <given-names>D.</given-names>
            <surname>Anastasiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Migge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soltan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dippold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kamlovskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Philippy</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 1st Worskhop on Towards Ethical and Inclusive Conversational AI: Language Attitudes</source>
          ,
          <article-title>Linguistic Diversity, and Language Rights (TEICAI</article-title>
          <year>2024</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , St Julians, Malta,
          <year>2024</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>20</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          . teicai-
          <volume>1</volume>
          .3.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jenkins</surname>
          </string-name>
          ,
          <article-title>Speaking your language: The psychological impact of dialect integration in artificial intelligence systems</article-title>
          ,
          <source>Current Opinion in Psychology</source>
          <volume>58</volume>
          (
          <year>2024</year>
          )
          <article-title>101840</article-title>
          . doi:
          <volume>10</volume>
          .1016/j. copsyc.
          <year>2024</year>
          .
          <volume>101840</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jenkins</surname>
          </string-name>
          ,
          <article-title>Speaking your language: The psychological impact of dialect integration in artificial intelligence systems</article-title>
          , Current Opinion in Psychology (
          <year>2024</year>
          )
          <fpage>101840</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kawaguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Inoue</surname>
          </string-name>
          ,
          <article-title>Japanese dialectology in historical perspectives</article-title>
          ,
          <source>Revue belge de philologie et d'histoire 80</source>
          (
          <year>2002</year>
          )
          <fpage>801</fpage>
          -
          <lpage>829</lpage>
          . doi:
          <volume>10</volume>
          .3406/rbph.
          <year>2002</year>
          .
          <volume>4642</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mitra</surname>
          </string-name>
          <article-title>, Multi-purpose nlp chatbot : Design, methodology &amp; conclusion</article-title>
          ,
          <source>ArXiv abs/2310</source>
          .08977 (
          <year>2023</year>
          ). URL: https://api.semanticscholar.org/CorpusID:264127993.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Chaves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Egbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hocking</surname>
          </string-name>
          , E. Doerry,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Gerosa</surname>
          </string-name>
          ,
          <article-title>Chatbots language design: The influence of language variation on user experience with tourist assistant chatbots</article-title>
          ,
          <source>ACM Trans. Comput.-Hum. Interact</source>
          .
          <volume>29</volume>
          (
          <year>2022</year>
          ). URL: https://doi.org/10.1145/3487193. doi:
          <volume>10</volume>
          .1145/3487193.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Elsholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamberlain</surname>
          </string-name>
          , U. Kruschwitz,
          <article-title>Exploring language style in chatbots to increase perceived product value and user engagement</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Human Information Interaction and Retrieval</source>
          , CHIIR '19,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>301</fpage>
          -
          <lpage>305</lpage>
          . URL: https://doi.org/10.1145/3295750.3298956. doi:
          <volume>10</volume>
          .1145/ 3295750.3298956.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoegen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Aneja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>McDuf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Czerwinski</surname>
          </string-name>
          ,
          <article-title>An end-to-end conversational style matching agent</article-title>
          ,
          <source>in: Proceedings of the 19th ACM International Conference on Intelligent Virtual Agents, IVA '19</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>111</fpage>
          -
          <lpage>118</lpage>
          . URL: https://doi.org/10.1145/3308532.3329473. doi:
          <volume>10</volume>
          .1145/3308532.3329473.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O. T. K.</given-names>
            <surname>Tran</surname>
          </string-name>
          , T. C.
          <article-title>Luong, Understanding what the users say in chatbots: A case study for the vietnamese language</article-title>
          ,
          <source>Eng. Appl. Artif. Intell</source>
          .
          <volume>87</volume>
          (
          <year>2020</year>
          ). URL: https://api.semanticscholar.org/ CorpusID:208947162.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N. N.</given-names>
            <surname>Chiaráin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Chasaide</surname>
          </string-name>
          ,
          <article-title>Chatbot technology with synthetic voices in the acquisition of an endangered language: Motivation, development and evaluation of a platform for Irish</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Grobelnik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mazo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (Eds.),
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Portorož, Slovenia,
          <year>2016</year>
          , pp.
          <fpage>3429</fpage>
          -
          <lpage>3435</lpage>
          . URL: https://aclanthology.org/L16-1547.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>George</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Muralikrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ninan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Varrier</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. L K</surname>
          </string-name>
          ,
          <article-title>Survey on the design and development of indian language chatbots</article-title>
          ,
          <source>2021 International Conference on Communication, Control and Information Sciences (ICCISc)</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCISc52257.
          <year>2021</year>
          .
          <volume>9484891</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gammoudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yassine</surname>
          </string-name>
          ,
          <article-title>Trends and challenges of arabic chatbots: Literature review</article-title>
          ,
          <source>Jordanian Journal of Computers and Information Technology</source>
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <article-title>1</article-title>
          . doi:
          <volume>10</volume>
          .5455/jjcit. 71-
          <fpage>1685381801</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aliwy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Taher</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          <article-title>AboAltaheen, Arabic dialects identification for all Arabic countries</article-title>
          , in: I. Zitouni,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdul-Mageed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bouamor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>El-Haj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tomeh</surname>
          </string-name>
          , W. Zaghouani (Eds.),
          <source>Proceedings of the Fifth Arabic Natural Language Processing Workshop</source>
          , Association for Computational Linguistics, Barcelona,
          <source>Spain (Online)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>302</fpage>
          -
          <lpage>307</lpage>
          . URL: https://aclanthology. org/
          <year>2020</year>
          .wanlp-
          <volume>1</volume>
          .
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Al-Ghadhban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Al-Twairesh</surname>
          </string-name>
          ,
          <article-title>Nabiha: An arabic dialect chatbot</article-title>
          ,
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>11</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .14569/IJACSA.
          <year>2020</year>
          .
          <volume>0110357</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Abu Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Habash</surname>
          </string-name>
          ,
          <article-title>Botta: An Arabic dialect chatbot</article-title>
          , in: H.
          <string-name>
            <surname>Watanabe</surname>
          </string-name>
          (Ed.),
          <source>Proceedings of COLING</source>
          <year>2016</year>
          ,
          <article-title>the 26th International Conference on Computational Linguistics: System Demonstrations, The COLING 2016 Organizing Committee</article-title>
          , Osaka, Japan,
          <year>2016</year>
          , pp.
          <fpage>208</fpage>
          -
          <lpage>212</lpage>
          . URL: https://aclanthology.org/C16-2044.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Boulesnane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Saidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kamel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Bouhamed</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Mennour</surname>
          </string-name>
          ,
          <article-title>Dzchatbot: A medical assistant chatbot in the algerian arabic dialect using seq2seq model</article-title>
          ,
          <source>in: 2022 4th international conference on pattern analysis and intelligent systems (PAIS)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . doi:
          <volume>10</volume>
          .1109/PAIS56586.
          <year>2022</year>
          .
          <volume>9946867</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdelhay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohammed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hefny</surname>
          </string-name>
          ,
          <article-title>Deep learning for arabic healthcare: Medicalbot, Social Network Analysis and Mining 13 (</article-title>
          <year>2023</year>
          )
          <article-title>71</article-title>
          . doi:
          <volume>10</volume>
          .1007/s13278-023-01077-w.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rekik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elamine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Belguith</surname>
          </string-name>
          ,
          <article-title>A medical chatbot for tunisian dialect using a rule-based and machine learning approach</article-title>
          ,
          <source>in: 20th ACS/IEEE International Conference on Computer Systems and Applications</source>
          ,
          <string-name>
            <surname>AICCSA</surname>
          </string-name>
          <year>2023</year>
          , Giza, Egypt, December 4-
          <issue>7</issue>
          ,
          <year>2023</year>
          , IEEE,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . URL: https: //doi.org/10.1109/AICCSA59173.
          <year>2023</year>
          .
          <volume>10479249</volume>
          . doi:
          <volume>10</volume>
          .1109/AICCSA59173.
          <year>2023</year>
          .
          <volume>10479249</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joukhadar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saghergy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kweider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ghneim</surname>
          </string-name>
          ,
          <article-title>Arabic dialogue act recognition for textual chatbot systems</article-title>
          ,
          <source>in: Proceedings of The First International Workshop on NLP Solutions for Under Resourced Languages (NSURL</source>
          <year>2019</year>
          )
          <article-title>co-located with</article-title>
          <string-name>
            <surname>ICNLSP 2019-Short Papers</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kowsher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Tithi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Alam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Huda</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Moheuddin</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Rosul</surname>
          </string-name>
          , Doly:
          <article-title>Bengali chatbot for bengali education</article-title>
          ,
          <source>in: 2019 1st International Conference on Advances in Science, Engineering and Robotics Technology (ICASERT)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . URL: https://api.semanticscholar. org/CorpusID:209460125.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kaleem</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. O'Shea</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Crockett</surname>
          </string-name>
          ,
          <article-title>Development of umair the urdu conversational agent for customer service</article-title>
          ,
          <source>Lecture Notes in Engineering and Computer Science</source>
          <volume>1</volume>
          (
          <year>2014</year>
          )
          <fpage>86</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Shabbir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. U.</given-names>
            <surname>Arshad</surname>
          </string-name>
          , W. Shahzad, Nubot:
          <article-title>Embedded knowledge graph with rasa framework for generating semantic intents responses in roman urdu</article-title>
          ,
          <source>ArXiv abs/2102</source>
          .10410 (
          <year>2021</year>
          ). URL: https://api.semanticscholar.org/CorpusID:231986364.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brixey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Traum</surname>
          </string-name>
          ,
          <article-title>Masheli: A choctaw-english bilingual chatbot</article-title>
          ,
          <source>in: International Workshop on Spoken Dialogue Systems Technology</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . URL: https://api.semanticscholar.org/ CorpusID:226325503.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [28]
          <string-name>
            <surname>A. B. E. Mabrouk</surname>
            ,
            <given-names>M. B. H.</given-names>
          </string-name>
          <string-name>
            <surname>Hmida</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Fourati</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Haddad</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Messaoudi</surname>
          </string-name>
          ,
          <article-title>A multilingual african embedding for faq chatbots</article-title>
          ,
          <source>ArXiv abs/2103</source>
          .09185 (
          <year>2021</year>
          ). URL: https://api.semanticscholar.org/ CorpusID:232240431.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pathak</surname>
          </string-name>
          ,
          <article-title>Shiksha mitra: An assamese language ai chatbot using deep learning</article-title>
          ,
          <source>International Journal of Scientific Research in Computer Science Engineering and Information Technology</source>
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <fpage>48</fpage>
          -
          <lpage>57</lpage>
          . doi:
          <volume>10</volume>
          .32628/CSEIT2390572.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Hailu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Welay</surname>
          </string-name>
          ,
          <article-title>Deep learning based amharic chatbot for faqs in universities</article-title>
          ,
          <source>arXiv preprint arXiv:2402.01720</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sandhini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Binu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rajeev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Reshma</surname>
          </string-name>
          ,
          <article-title>A proposal of chatbot for malayalam</article-title>
          ,
          <source>International Journal of Computer Sciences and Engineering</source>
          <volume>06</volume>
          (
          <year>2018</year>
          )
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          . doi:
          <volume>10</volume>
          .26438/ijcse/v6si6.
          <fpage>2125</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dabre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kanojia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hafari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dippold</surname>
          </string-name>
          ,
          <article-title>Natural language processing for dialects of a language: A survey</article-title>
          ,
          <source>ArXiv abs/2401</source>
          .05632 (
          <year>2024</year>
          ). URL: https://api.semanticscholar. org/CorpusID:266933497.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>M.</given-names>
            <surname>Orosoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. R.</given-names>
            <surname>Alphonse</surname>
          </string-name>
          , G. Fatma,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rengarajan</surname>
          </string-name>
          , B. Kiran Bala,
          <article-title>Enhancing natural language processing in multilingual chatbots for cross-cultural communication</article-title>
          ,
          <source>in: 2024 5th International Conference on Intelligent Communication Technologies and Virtual Mobile Networks (ICICV)</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>127</fpage>
          -
          <lpage>133</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICICV62344.
          <year>2024</year>
          .
          <volume>00027</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hefernan</surname>
          </string-name>
          , &lt; research note&gt;
          <article-title>an introduction to the corpus of kansai spoken japanese</article-title>
          ,
          <source>Journal of Policy Studies</source>
          (
          <year>2012</year>
          )
          <fpage>157</fpage>
          -
          <lpage>163</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>S.</given-names>
            <surname>Takamichi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saruwatari</surname>
          </string-name>
          ,
          <article-title>CPJD corpus: Crowdsourced parallel speech corpus of Japanese dialects</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Hasida</surname>
          </string-name>
          , H. Isahara,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>