<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.48550/ARXIV.1910.01108</article-id>
      <title-group>
        <article-title>QUA4I: Question Answering for the Industry 4.0 Domain. An Application of Intelligent Virtual Assistants</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Iñigo López-Riobóo-Botana</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dana Gallent-Iglesias</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Gonzalez-Vázquez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto Tecnológico de Galicia - ITG - Centro Tecnológico Nacional</institution>
          ,
          <addr-line>Cantón Grande 9, Planta 2, 15003, A Coruña</addr-line>
          ,
          <country>España</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The recent advancements with LLMs (Large Language Models) have led to many natural language applications solving diferent NLP (Natural Language Processing) tasks. In such systems, both NLG (Natural Language Generation) and NLU (Natural Language Understanding) play a crucial role. The notion of instruction-based LLMs caused an important impact in the development of chatbot-styled NLP applications. In the past months, we have seen an incredible amount of LLMs making use of this idea to solve tasks including, but not limited to, QA (Question Answering), information extraction, intent recognition, dialogue or language generation. Whereas these systems have very good generalisation capabilities, in which the AGI (Artificial General Intelligence) idea has spread to the NLP research field, we still lack of good domain adaptation techniques for tailored ADI (Artificial Domain Intelligence) systems. Apart from prompt engineering, we cannot fully control the outputs of this new chatbot-styled LLMs to obtain our desired output given a specific domain. In the context of the CEL.IA network, we present QUA4I (QUestion Answering for the Industry 4.0), a chatbot-oriented application or IVA (Intelligent Virtual Assistant) for the industry 4.0 domain, mixing the NLU and NLG techniques using the Rasa chatbot framework. We designed a custom demo for question answering and information extraction about the industry 4.0 topic, including a dialogue system which can also generate automatic responses in natural language. We included both ASR (Automatic Speech Recognition) and TTS (Text To Speech) modules, so we can also interact with the bot using spoken language in Spanish.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;question answering</kwd>
        <kwd>information extraction</kwd>
        <kwd>IVA</kwd>
        <kwd>NLU</kwd>
        <kwd>NLG</kwd>
        <kwd>ASR</kwd>
        <kwd>TTS</kwd>
        <kwd>industry 4</kwd>
        <kwd>0</kwd>
        <kwd>Rasa</kwd>
        <kwd>chatbot</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        only transformer models, such as GPT-3 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (OpenAI),
BLOOM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (open source), PaLM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (Google) or LLaMA
In recent years, we have seen an incredible amount of [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (Meta). The main task of these models is to predict
new chatbot-oriented applications [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Chatbots and the next token given a sequence, but we have seen a drift
IVAs (Intelligent Virtual Assistants) can automate tasks towards dialogue capabilities recently, which is a clear
without human intervention, saving a lot of time and ef- improvement for instruction-based applications. Recent
fort with automation. Chatbots can be applied to several and well-known examples are DialoGPT [7] (Microsoft),
domains, including healthcare, retail, banking, among InstructGPT [8] and ChatGPT [9] (OpenAI) or LaMDA
others. For example, these systems can be integrated [10] and Bard [11] (Google). LLMs have grown to the
in websites or VR (Virtual Reality) environments. We point that they can mimic or even outperform human
can provide chatbots with more interfaces in addition answers in some NLP tasks [12]. These breakthroughs
to text, like human voice, so we can interact with them are possible thanks to the continuous and competitive
in diferent ways. Moreover, the great advancements in scaling of deep learning models [13] and the proposal of
the NLP (Natural Language Processing) field have led new architectures [14].
to a very good understanding of the language and in- In Section 1.1, we present our motivation to carry out
credible improvements in generation capabilities, thanks this project. In Section 1.2 we enumerate our main
conto LLMs (Large Language Models). LLMs are decoder- tributions. In Section 2, we depict the architecture and
methodology followed for this project, describing its
limitations in Section 3. Finally, we conclude with some ideas
and future work in Section 4.
      </p>
      <sec id="sec-1-1">
        <title>1.1. Motivation</title>
        <p>For domain-specific and ad hoc conversational agents
(i.e., to fulfil special needs of a use case), we need to adjust
and control the chatbot output thoroughly. The
chatbotoriented LLMs fit in AGI (Artificial General Intelligence)
contexts and they can be somewhat “fine-tuned” with
prompt engineering [15, 16, 17, 18], but this is not enough
for guided ADI (Artificial Domain Intelligence) systems.
Adaptation and customisation of these chatbots is still a
work-in-progress [19, 20].</p>
        <p>In the context of the CEL.IA network1, we propose
a domain-specific IVA (Intelligent Virtual Assistant) for
the Industry 4.0 domain. We mixed the concepts of NLU
(Natural Language Understanding) and NLG (Natural
Language Generation) using the Rasa chatbot framework
[21]. In this paper, we describe our first demo version
for the question answering and information extraction
NLP tasks, including a dialogue system which can also
generate answers in natural language. We considered
both ASR (Automatic Speech Recognition) and TTS (Text
To Speech) modules to enhance the communication
interfaces with the chatbot.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Contributions</title>
        <p>Our main contributions are as follows:
• We implemented a domain-specific IVA
system with Rasa framework for NLU and
dialogue. We solve the QA (Question Answering)
and information extraction NLP task in the
context of Industry 4.0.
• We avoided relying on third-party APIs and
deployed our own on-premise models. These
models are currently available through REST
services to improve the capabilities of the chatbot.
• We provide both voice and text interfaces, so
that we can communicate with spoken language
in Spanish with the chatbot, as well as listen to
the answers.
• We handle out-of-scope questions with
automatic dialogue generation using our own
GPTbased Spanish model for chitchat, fine-tuning a
DialoGPT-2 [7] model2.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>We designed and implemented the following modules:
• NLU and dialogue modules: These are the core
modules for intent recognition and dialogue flow,
respectively. The Rasa framework is in charge of
this management.
• QA module: This module is in charge of
providing the answer according to an user information
need, following an information extraction task
with a QA approach. We manage our own
document store in our servers, containing information
about the Industry 4.0 domain in several files.
This knowledge base, which basically contains
descriptions and definitions, can be extended if
new information needs arise. This module was
implemented as a REST service.
• NLG module: This module is in charge of
generating automatic responses when the answer is
neither known by Rasa framework nor by the QA
system. In such situation, we implemented a
single turn dialogue or chitchat mechanism with the
chatbot, trying to preserve a logical conversation
lfow. This module was also deployed as a REST
endpoint.
• ASR module: The speech recognition module
is implemented in our custom NLP services,
enabling us to transcribe from audio to text and
then send it back to the chatbot. After that, it is
processed in the NLU module. We based our
implementation on the Spanish stt_es_citrinet_512
model, publicly available in the NVIDIA NeMo
toolkit3.
• TTS module: When fetching the chatbot
response, we transform the text output to human
voice again so that we can easily interact with
the IVA system. This service is also implemented
in our custom NLP services. We based our
implementation on the Spanish glow-speak:es_tux
model, publicly available for production-ready
environments using the OpenTTS framework4.</p>
      <p>The general diagram of the IVA system is depicted
in Figure 1. In Section 2.1, we study in more detail the
NLU and dialogue implementations for the chatbot. In
Section 2.2, we present our approach for the question
answering functionality of the chatbot. In Section 2.3, we
describe our language generation method for the chitchat
situations. We provide an example conversation in Figure
2.</p>
      <sec id="sec-2-1">
        <title>2.1. NLU and dialogue module</title>
        <p>These two modules are the principal components of Rasa
open source [21].</p>
        <p>1. The NLU subsystem is in charge of the intent
recognition NLP task. Rasa projects follow a
datadriven approach, providing several data files with
text samples for each intent and configuration
ifles to adjust the pipelines for model training.
2. The dialogue module relies on a combination
of a rule-based system and a “user stories”
mechanism to infer the next step in the
conversation. These actions can be (1) direct chatbot</p>
        <sec id="sec-2-1-1">
          <title>1https://www.redcelia.es/</title>
          <p>2https://huggingface.co/ITG/DialoGPT-medium-spanishchitchat</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>3https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/</title>
          <p>models/stt_es_citrinet_512
4https://github.com/synesthesiam/opentts#voices</p>
          <p>responses (when we already fulfilled the
information need from Rasa) or (2) delegations in
the independent SDK action server to access
third-party services (i.e., our aforementioned
custom NLP services).</p>
          <p>In order to train our Rasa model for intent recognition,
we used the following Rasa components [22]:
1. Tokenizer: Rasa component in charge of
splitting each sentence into tokens or words. We used
the simple WhiteSpaceTokenizer to get the tokens
splitting with white spaces.
2. Feature extractor: Rasa component in charge
of feature engineering to transform the
corresponding tokens into numerical vector
representations. We made use of feature extractors
from pre-trained transformer-encoder backbones.</p>
          <p>Specifically, we used the deep word embeddings
from the last encoder layer of a mDistilBERT
transformer model [23]. We used the
multilingual version with support for Spanish, with a
total size of 134M parameters (compared to 177M
parameters for mBERT-base [24]). On average
mDistilBERT is twice as fast as mBERT-base [25]. 2.2. QA module
3. Intent classifier : Rasa component in charge of
the intent recognition NLP task. We used the
DIET classifier from the Rasa framework authors
[26]. DIET is a multi-task modular transformer
architecture that handles both intent
classification and entity recognition together. It provides
the ability to plug and play various pre-trained
embeddings like BERT (and variants), GloVe,
ConveRT, among others [27].
4. Fallback classifier : This Rasa component is in
charge of triggering the default intent when the
intent prediction from the DIET classifier does not
have a confidence above a pre-specified
threshold. If so, the default intent will diverge from the
regular conversation, following a custom action.</p>
          <p>This is the situation in which the Rasa SDK action
server uses our custom NLP services.</p>
          <p>Nowadays, the most amount of NLP pipeline
approaches are end-to-end and avoid the feature
extraction step [28]. However, we are following a more
classical approach. We do not have enough data for an
end-to-end fine-tuning of a transformer model for the
specific task of Industry 4.0 intent classification.
Moreover, relying on multilingual deep word embeddings to
define the input features for the DIET classifier has some
advantages, considering that this model outperforms
finetuned BERT and is about six times faster to train [26, 27].</p>
          <p>For the QA task, we followed a transformer-based
approach using encoder models (BERT-like) for information
extraction using an input context [29]. We propose an
extractive QA approach, using question answering models
already fine-tuned in Spanish corpora. Basically, these
models add to a pre-trained encoder backbone two difer- some single-turn professional-styled flows 8. We present
ent classification heads: (1) one of the heads is in charge the training hyper-parameters in Table 2. This model is
of predicting the starting point (index) of the answer and publicly available in our Huggingface repository9.
(2) the other is in charge of predicting the ending point.</p>
          <p>These models receive both the question and the context Table 2
as the input. They can generalise well to many difer- Fine-tuning hyper-parameters for our chitchat 345M
parameent input contexts. Then, the answer is provided if the ters DialoGPT-spanish-medium model.
prediction is above a pre-specified confidence threshold. Hyper-parameter Value</p>
          <p>After exploring all the available models for extractive Validation data partition (%) 20%
QA in Spanish in the Huggingface repository, we ended Training batch size 8
up with the Spanish BERT (BETO)5 fine-tuned in the Learning rate 5e-4
SQuAD-es-v2.0 Spanish dataset6. This fine-tuned model Max training epochs 20
is available in the Huggingface repository7. We used the WarmuWpetirgahintidnegcastyeps (%) 06.0%1
pipeline implementation for question answering from Optimiser ( 1,  2,  ) AdamW (0.9, 0.999, 1 − 08)
Huggingface, so that we can handle long input contexts Monitoring metric (Δ, patience) validation loss (0.1, 3)
(i.e., process long context documents for the Industry 4.0
domain). These industry-related documents, stored in
our servers, define the knowledge base for the topic. We
used the inference parameters depicted in Table 1. Note 3. Limitations
that we set the pipeline’s maximum input length
(question + context) to the maximum input length allowed Some problems could arise when using our chatbot
imby the BERT-based model by design (up to 512 tokens). plementation. These issues are related to the following
Since the context documents are longer than that, we fol- topics:
lowed a window-based approach of the input, applying
overlaps.</p>
          <p>5https://github.com/dccuchile/beto
6https://github.com/ccasimiro88/TranslateAlignRetrieve
7https://huggingface.co/mrm8488/bert-base-spanish-wwmcased-finetuned-spa-squad2-es</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>8https://github.com/microsoft/botframework-cli/blob/main/</title>
          <p>packages/qnamaker/docs/chit-chat-dataset.md
9https://huggingface.co/ITG/DialoGPT-medium-spanishchitchat</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. NLG module</title>
        <p>For the NLG task, we followed a transformer-decoder
approach, using our own version of a GPT-2 model
[30], from the pre-trained 345M parameters
DialoGPTmedium model for dialogue [7].</p>
        <p>We further fine-tuned this model in an auto-regressive
and self-supervised fashion, following the CLM (Causal
Language Modelling) objective. We optimised the next
token prediction task using a chitchat dataset including
• Natural language generation: Considering
that we are using a custom NLG model for
chitchat, when the user asks for out-of-scope
information, we cannot fully control the outputs.
To mitigate this issue, we decided to fine-tune
the model with some overfitting so that the
probability distribution for the next word is guided
in the style of the chitchat behaviour. Moreover,
we adjusted the inference generation parameters
(e.g., temperature) so that we favour the highest
probability words to be predicted.
• Rasa knowledge base for dialogue and intent
recognition: When new information needs in
relation to Industry 4.0 arise, we have to manually
add new data for new intent recognition by the
Rasa NLU module. We also need new tailored
answers. Despite this disadvantage, we gain control
over the output, avoiding generation
hallucination [31], one important problem in the LLM field
nowadays.
• Document store for information extraction:
As stated in the previous point, the same problem
happens in the extractive QA approach. We have
to manually maintain the document store in our
servers to provide additional information when
required.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusion and future work</title>
      <p>In this work, we present QUA4I, a chatbot-oriented
application for the Industry 4.0 domain. We included a
dialogue system which can also generate automatic
responses in natural language. We developed a tailored IVA
tool with controlled text generation capabilities,
avoiding the LLM hallucination problem. We mixed both NLU
and NLG techniques using the Rasa chatbot framework,
designing a custom demo for question answering and
information extraction. We also included both ASR and
TTS modules to improve the chatbot interfaces with the
user.</p>
      <p>We plan to integrate this chatbot service in a AR
(Augmented reality) or MR (mixed reality) environment
for assistance in some industry-related domains. Our
medium-term objective is to fulfil not only information
needs but also take action in response to user commands.
We are considering NLG improvements by fine-tuning
other state-of-the-art dialogue models, including the
Galician version of them. Moreover, we want to explore
NLU improvements following an end-to-end approach
with a custom intent classifier based on
transformerencoders, avoiding the current feature extraction step in
the pipeline. To do so, we are exploring data
augmentation techniques based on cross-language translation and
NLG models.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This project belongs to the CEL.IA network initiative10,
which is supported by the Ministerio de Ciencia e
In10https://itg.es/cervera-celia/
novación through the CDTI (Centro para el Desarrollo
Tecnológico Industrial) (grant CER-20211022).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Caldarini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jaf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McGarry</surname>
          </string-name>
          ,
          <source>A Literature Survey of Recent Advances in Chatbots, Information</source>
          <volume>13</volume>
          (
          <year>2022</year>
          ). URL: https://www.mdpi.com/2078-2489/13/ 1/41. doi:
          <volume>10</volume>
          .3390/info13010041.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Caldarini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jaf</surname>
          </string-name>
          , Recent Advances in Chatbot Algorithms, Techniques, and
          <article-title>Technologies: DESIGNING CHATBOTS</article-title>
          , in: Trends, Applications, and Challenges of Chatbot Technology,
          <source>IGI Global</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>245</fpage>
          -
          <lpage>273</lpage>
          . URL: http://sure.sunderland.ac.uk/id/ eprint/15712/. doi:
          <volume>10</volume>
          .4018/978-1-
          <fpage>6684</fpage>
          -6234- 8.
          <year>ch011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          , et al.,
          <article-title>Language Models are Few-Shot Learners</article-title>
          , in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          . URL: https://proceedings.neurips.cc/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Workshop</surname>
          </string-name>
          , BLOOM:
          <string-name>
            <given-names>A</given-names>
            <surname>176B-Parameter OpenAccess Multilingual Language Model</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2211.05100. doi:
          <volume>10</volume>
          .48550/ ARXIV.2211.05100.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , et al.,
          <source>PaLM: Scaling Language Modeling with Pathways</source>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2204.02311. doi:
          <volume>10</volume>
          . 48550/ARXIV.2204.02311.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          , et al.,
          <source>LLaMA: Open and Eficient Foundation Language Models,</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>