<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Comparison of Natural Language Understanding Services to build a chatbot in Italian?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matteo Zubani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Serina Ivan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfonso Emilio Gerevini</string-name>
          <email>alfonso.gerevinig@unibs.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Brescia</institution>
          ,
          <addr-line>Via Branze 38, Brescia 25123</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mega Italia Media S.P.A.</institution>
          ,
          <addr-line>Via Roncadelle 70A, Castel Mella 25030</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>104</fpage>
      <lpage>117</lpage>
      <abstract>
        <p>All leading IT companies have developed cloud-based platforms that allow building a chatbot in few steps and most times without knowledge about programming languages. These services are based on Natural Language Understanding (NLU) engines which deal with identifying information such as entities and intents from the sentences provided as input. In order to integrate a chatbot on an e-learning platform, we want to study the performance in intent recognition task of major NLU platforms available on the market through a deep and severe comparison, using an Italian dataset which is provided by the owner of the e-learning platform. We focused on the intent recognition task because we believe that it is the core part of an e cient chatbot, which is able to operate in a complex context with thousands of users who have di erent language skills. We carried out di erent experiments and collected performance information about F-score, error rate, response time and robustness of all selected NLU platforms.</p>
      </abstract>
      <kwd-group>
        <kwd>Chatbot • Cloud platform • Natural Language Understanding • E-learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In the last decade more and more companies have replaced their traditional
communication channels with chatbots which can satisfy automatically the users'
requests. A chatbot is a virtual person that can talk to a human user using
textual messages or voice.</p>
      <p>Given the high demand for this technology, leading IT companies have
developed cloud-based platforms which o er NLU engines and fast-developing tools,
with the purpose to allow the user to build virtual assistants only providing a
bunch of examples without knowing anything about NLU algorithms.</p>
      <p>This paper comes from the need of the private company Mega Italia Media
S.P.A., to implement a chatbot in their messaging system. We realised the lack
of a systematic evaluation of cloud base NLU services for the Italian language,
moreover research papers which present real cases of use of these systems do not
explain why one service was preferred over another.</p>
      <p>This research aims to evaluate the performance of leading cloud-based NLU
services in a real context using an Italian dataset. Thus, we used users' requests
provided by Mega Italia Media S.p.A collected through their chat system and we
designed four di erent experiments in order to compare the selected platforms
from di erent points of view such as the ability to recognise intents, response
time and robustness to spelling errors.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        NLU algorithms can be used in complex and critical contexts such as the
extraction and the classi cation of drug-drug interactions in the medical eld
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the classi cation of radiological reports [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or named entity
recognition (NER) task[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. These works do not use commercial NLU platforms that
are only appropriate for not critical and general-purposes like chatbots or Spoken
Dialogue Systems (SDS).
      </p>
      <p>
        Several publications exist which use the cloud-based platforms to create
virtual assistants for di erent domains, e.g. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] uses DialogFlow to implement a
medical dialogue system, [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] presents a virtual teacher for online education based
on IBM Watson and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] shows the use of chatbots to the Internet of things. One
of the architectural elements of these virtual assistants is a cloud-based NLU
service. However, none of these papers discuss how they chose a particular service
instead of another or propose an in-depth analysis concerning the performance
of their system.
      </p>
      <p>
        Since 2017 di erent papers have analysed the performance of the leading
platforms on di erent datasets; e.g. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] builds and analyses two di erent corpora
and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presents a descriptive comparison (o ering a taxonomy) and a
performance evaluation, based on a dataset which includes utterances concerning the
weather. While in 2019 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] proposed a comparison concerning the performance
of 4 NLU services on large datasets over 21 domains and 64 intents, and [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
presents a less extensive and deep comparison than what is shown in this paper
on the same domain.
      </p>
      <p>All the platforms are evolving year after year then also the performance
could change compared to what previous papers presented. Furthermore, to our
knowledge, a comparison of a dataset which is not in English does not exist; this
pushed us to study if the performance in Italian is still the same as in latest
research.</p>
    </sec>
    <sec id="sec-3">
      <title>Natural Language Understanding Cloud Platforms</title>
      <p>The principal goal of NLU algorithms is the extraction of useful and structured
information from a natural language input which by its nature is unstructured.
The structures obtained by the NLU service are two, intents and entities.
{ Intent: the services should understand and classify what the user means in
his sentence, which is not necessarily a request or a question but can be any
user's sentence.
{ Entity: instead of dealing with the overall meaning of the user's sentences,
the entities extraction tries to identify information and parameter values
inside the sentence itself.</p>
      <p>Watson (IBM): It is a complete NLU framework which allows developing
a chatbot using a web-based user interface or to exploit di erent SDKs, in
order to build a virtual assistant fast and with di erent programming languages.
The platform is easy to integrate with a variety of communication services and
provides complete support to Text-to-Speech and Speech-to-Text technology.
Watson allows recognising intents and entities, and some models concerning
general topic are provided by the platform, while for speci c domains the users
have to create custom models giving a quite small set of examples to train the
NLU engine. The context allows, once intents or entities are recognised, to store
data and to reuse them in following dialogue interactions. A Dialogue frame is
developed through a tree structure which allows designing deep and articulate
conversation.</p>
      <p>Dialog ow (Google): Previously known as Api.ai it has recently changed
its name into Dialog ow and it is a platform which develops virtual textual
assistants and quickly adds speech capability. This framework creates chatbots
through a complete web-based user interface, or for complex projects, it uses
a vast number of APIs via Rest (there are also many SDKs for di erent
programming languages). Di erent intents and entities can be created by giving
examples. In Dialog ow the context is an essential component because the data
stored inside it allow developing a sophisticated dialogue between the virtual
assistant and the user.</p>
      <p>Luis (Microsoft): Similar to the previous platforms, Luis is a cloud-based
service available through a web user interface or the API requests. Luis provides
many pre-built domain models including intents, utterances, and entities,
furthermore, it allows de ning custom models. The management of dialogue ow
results to be more intricate than other products. However, it is possible to build
a complete virtual assistant with the same capabilities o ered by other services.
Di erent tools, which facilitate the development of a virtual assistant, are
supported (such as massive test support, Text-to-Speech and sentiment analysis).</p>
      <p>Wit.ai (Facebook): The Wit.ai philosophy is slightly di erent compared
to other platforms. Indeed, the concept of intent does not exist, but every model
in Wit.ai is an entity. There are three di erent types of \Lookup Strategy" that
distinguish among various kinds of entities. \Trait" is used when the entity value
is not inferred from a keyword or speci c phrase in the sentence. \Free Text"
is used when we need to extract a substring of the message, and this substring
does not belong to a prede ned list of possible values. \Keyword" is used when
the entity value belongs to a prede ned list, and we need substring matching to
look it up in the sentence. Wit.ai has both a web-based user interface and APIs
furthermore there are several SDKs.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Case of Study</title>
      <p>The Italian company Mega Italia Media S.P.A. acting in the e-learning sector is
the owner of DynDevice, a cloud-based platform which provides online courses
about \occupational safety". Every day they supply courses to thousands of
users. Therefore, they want to introduce an arti cial assistant on their platform
to respond to the frequently asked questions. Among di erent types of virtual
assistants, they chose to implement a chatbot because they already have a
messaging system on their platform and users are accustomed to using this kind of
interface, but currently, only human operators can respond to users' requests.
Aiming at suggesting the best cloud-based NLU service that ts the company
requirements, we decided to analyse the performance of di erent platforms
exploiting real data which the company has made available in the Italian language.
4.1</p>
      <sec id="sec-4-1">
        <title>Data Collection</title>
        <p>In order to undertake the study proposed in this paper, we use data provided
by the company concerning the conversations between users and human
operators occurred between 2018-01-01 and 2019-12-31. We used only data from the
last two years because the e-learning platform is continuously developing, and
also the requests of the users are evolving. Therefore, introducing an additional
feature in the platform may produce a new class of requests from the users;
moreover, bugs resolution can cause the disuse of one or more class of requests.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Data Classi cation</title>
        <p>Obviously, NLU services have to be trained, thus it is necessary to create a
training dataset which includes labelled data. The requests collected were not
classi ed, so we had to label them manually. For the classi cation phase, we
developed a simple program that randomly extracts 1000 requests and then
shows them one by one to the operator who is assigned to the classi cation. The
operator, through a command-line interface, assigns the label which describes
better the intent of request and then saves the tuple &lt; request; label &gt; in the
dataset. All the sentences extracted from conversations are made anonymous,
and we do not treat them in any other way such as removing special characters
and punctuation marks or correcting spelling mistakes. This anonymising step
is necessary to satisfy the privacy policy of the company.</p>
        <p>It is necessary to highlight that 33.1% of the total amount of requests was
rejected because they were meaningless or because inside of these requests there
were references to emails or phone calls occurred previously between user and
operator. Therefore, only a human operator can solve this type of requests.
Some meaningless requests occurred because some users wrote a text in the chat
interface to try it, with no needs. All 669 classi ed requests are split into 33
di erent intents; the major part of these had just one or few instances associate
to them. So, we selected only the most relevant intents among the entire set
of detected intents. How we can see in table 1, the examples belonging to the
subgroup of principal intents are 551, and they cover the 82.4% of all examples
classi ed. A small part (37) of these are not compliant with the constraints
imposed by platforms. The actual number of examples that can be used to build
the training set and the test set is 514. In table 2 of section 5.1, we show the
distribution of the 551 classi ed sentences over the selected intents which range
from a minimum of 22 up to 190.</p>
        <p>Below, we present the English translation of two requests about the same
intent \Expired course", one simple and another one more complex in order to
show the variety and the complexity of the sentences coming from a real case of
study:
1. \Course expired; is it possible to reactivate it?"
2. \Good morning, I sent a message this morning because I can't join \Workers
- Speci c training course about Low risk in O ces", probably because it has
expired and unfortunately I didn't realise it. What shall I do? Thank you"
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation Experiments</title>
      <p>We selected only services with Italian language support: Watson, Dialog ow,
Luis, Wit.ai. We chose the free version for all NLU services analysed because
there is only a constraint about the number of APIs calls that can be made daily
or monthly, while the NLU engine capabilities are not limited. We tested the
capability of recognising correctly the underlying intent of every single message
sent by users, in the real case described in section 4. To evaluate the results as
thoroughly as possible, we designed four di erent experiments. The experiment
(A) aiming at evaluating the performance on the whole extracted training set,
in other words, we use all the examples available to train the NLU platform. In
experiment (B) we study the performance of the di erent systems at the increase
of the number of examples provided to train the platforms, while the test set
remains xed. While the experiment (C) tests the response time of each platform.
Finally in experiment (D) we test the robustness of the analysed services, so we
built di erent test sets with an increasing number of misspellings contained in
each example and then we calculate the error rate of every single test set.
5.1</p>
      <sec id="sec-5-1">
        <title>Training And Test Sets</title>
        <p>To carry out the rst experiment (A), we use classi ed data divided by main
intents. We extract the training set composed of 75% of entire examples collection
and test set made of the remaining elements (which corresponds to 25% of the
whole dataset). In table 2, we present an overview of how the di erent datasets
are made up for each intent.</p>
        <p>In experiment (B), we use a xed test set, and we build nine training subsets
of the training set previously de ned for the experiment (A). The rst training
subset is created by extracting 10% of elements relating to every single intent
of the initial training set. The second one comprises 20% of the examples of the
initial training set, keeping all instances which are included in the rst training
subset. This process is repeated increasing by 10% of examples for each intent
until we reach the entire initial training set. We show the composition of training
subsets and test set in table 3.</p>
        <p>In order to carry out the experiment (C), we select the dataset built for
the experiment (A). The training set remains unchanged while the test set is
expanded replicating elements in it until the number of elements reaches 1000.</p>
        <p>We select the dataset built for the experiment (A) also for the experiment
(D). We use the training set as in experiment (A) to train the services; from
the test set of experiment (A), we extracted the 25 examples which intents were
recognised correctly by all the platforms and then we create 30 new di erent
test sets: in the rst new test set, we change randomly a character for each
examples belonging to the test set. In order to build the second test set, we
use the test set just create and we change another character randomly for each
example belonging to the test set. In this way, we create new 30 test sets, where
the last test set has for each example 30 characters changed compared to the
initial test set.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Experimental Design</title>
        <p>
          The number of examples for some intents is quite small, and it is not possible
to build a k-Fold Cross-Validation model. Thus, we de ne ten di erent training
sets and corresponding test sets with the structure illustrated in section 5.1
experiment (A), making sure to extract the examples randomly as an alternative
to Cross-Validation. To evaluate if the di erences between the results of the NLU
platforms are statistically signi cant we use Friedman test and then we propose
the pairwise comparison using Post-hoc Conover Friedman test as shown in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>To analyse Experiment (B), we take among previously mentioned datasets
the one that has the F-score closest to the average of the F-score on all datasets,
and we create nine training subsets, as reported in section 5.1.</p>
        <p>For experiment (C) we use a training set of experiment (A) and we built a
test set with 1000 elements (as mention in section 5.1). In order to analyse the
response time of the di erent services we send to the NLU engines each example
in the test set independently, and we collect the time lapses between sending the
message and receiving the reply. We summarise the measured times in a report
where one can nd average, standard deviation, highest and lower response time.
We repeat this experiment three times in di erent hours of the day because the
servers are located in di erent places of the world and the load of each server
may depend on the time when we perform the test.</p>
        <p>In experiment (D) we chose a training set used in experiment (A) and we
built 30 di erent test sets as reported in section 5.1. We examine all the test sets
and for each one, we collect how many times the service recognises the correct
intent and then we calculate the error rate on all test sets analysed.</p>
        <p>We developed a Python application which uses the SDKs a orded by the
owners of each platform with the aim to train, test and evaluate the
performance of each service. Wit.ai supports Python programming language, however,
the instructions set is limited and to overcome this de ciency, we wrote a speci c
code that allowed us to invoke a set of HTTP APIs. The program receives two
CSV les as input, one for the training set and one for the test set, and then
it produces a report. The application is divided into three modules completely
independent that means each one can be used individually.</p>
        <p>Training module: for each intent de ned, all examples related to it are sent to
the NLU platform, and then the module waits until the service ends the training
phase.</p>
        <p>Testing module: for each element in the test set, the module sends a message
containing the text of the element. The application waits for the NLU service
reply. All the analysed platforms produce information about the intent
recognised and the con dence associated with it. The module compares the intent
predicted and the real intent, and then it saves the result of the comparison in a
le. We want to underline that the application sends and analyses every instance
of test set independently. In this module, there is an option that allows testing
the response time using a timer which is activated before sending the message
and it is ended when the service reply arrives.</p>
        <p>Reporting module: this is responsible for the reading of the le saved which
contains results of the test module and then calculates the statistical indexes,
both for each intent and the entire model. The indexes calculated are Recall,
Accuracy, F-score and Error rate. This module also allows saving a report about
response time where one can nd average, standard deviation, maximum and
minimum response time.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Results And Discussion</title>
      <p>As mentioned in section 5, we create 4 di erent experiments. In experiment(A)
we studied the performance on the training set with all elements available, in
experiment (B) instead, we evaluated the results with the increase of training set
instances, in the experiment (C) we measured the response time of all selected
platforms and in the experiment (D) we evaluated the robustness of services.
Experiment (A) What we expected in the rst experiment is that the
performance might not be uniform among di erent NLU platforms, but we supposed
that the outcomes would be relatively constant on di erent datasets randomly
built as described in section 5.1 (from here called D1, D2...D10).</p>
      <p>Table 4 presents the performance for each service in terms of Error rate and
F-score 3 and the last two columns provide the average and standard deviation
on the ten di erent datasets.
As we had supposed the general performance among platforms is di erent.
The average of F-score is roughly equal between Dialog ow and Watson over
0.86; however, the latter has a more signi cant standard deviation. Luis' results
are slightly worse than the other two with 0.82, while Wit.ai is the worst with an
average just above 0.71. So to con rm this, we can look at table 4 where the best
F-score for each dataset is bold, and in no case, Luis or Wit.ai have managed
to overcome Dialog ow or Watson. These considerations are evident in gure 1
(a), where the median of Dialog ow is the highest while that of Wit.ai is the
3 We used python function f1 score belonging to the sklearn module. Being dataset
unbalanced, we set the parameter \average" equal \weighted".
lowest. We can see almost all the lengths of boxes are quite short, this indicates
low dispersion, and it means that the performance of services is stable moving
through the ten di erent datasets.</p>
      <p>The number of examples associated with each intent is not balanced.
Consequently, the performances on the single intent can be various, so in gure 2 we
decided to break down the F-score outcome in every single intent. The legends
show the intent ID and number of examples used to train the services. Looking
at the diagrams, we notice that intent I9, which uses the highest number of
elements, has one of the highest median and lower dispersion for all services. I2
is trained with the lowest number of examples, and its performances are quite
variable, but in all platforms, the median is far below 0.8, and the boxes are very
stretched. Dialog ow has better results in general, and the boxes are shorter than
others, but there is an anomaly, in fact, I2 is signi cantly worse than Watson
and Luis with a very high index of dispersion. The outcomes of Watson and
Luis are pretty good; indeed, the median is always over 0.7. To con rm previous
analysis on table 4 Wit.ai does not perform well, it has two intents under 0.6
and only two over 0.8, also for some intents the dispersion is very high.</p>
      <p>Figure 3 presents the result of Post-hoc Conover Friedman test and it shows
if the di erence between results produced by all the NLU platforms analysed is
signi cant with di erent p values. We can observe that with p &lt; 0:001, Watson
and Dialog ow perform better than Luis and Wit.ai and we can also assert that
the di erence between Dialog ow and Watson is not statistically signi cant with
the same p value.</p>
      <p>Experiment (B) In this experiment, we selected the rst dataset DS1, and
we split it into nine training subsets (as described in section 5.1), while the
test set remains the same. We expected that increasing the size of the training
set corresponds to an increase of performance until they reach the same
Fscore obtained on the entire training set. The graph (a) in gure 4 con rms
our assumption that F-score of all four platforms starts low and then increases.
What we can see on the same graph is that the curves of Watson, Dialog ow and
also Wit.ai grow quite fast until 40% and then uctuate or grow slowly, while
Luis rises steadily up until it reaches its maximum. Watson and Dialog ow are
signi cantly better than Luis and Wit.ai with small training sub-datasets which
is showed by the graph (b) in gure 4 where the error rate, especially on the
left, for the rst two services is signi cantly lower than the other two.
Experiment (C) The response time of chatbots is generally not a critical
value. In some speci c applications such as a Spoken Dialogue Systems (SDS),
the time of intent recognition task should be short so that increase the speaking
experience when a user tries to use the SDS.</p>
      <p>In this experiment, we want to understand if the response time is very
different between the various platforms. It is not possible to place the servers in
the same location, so we execute the test three times in three di erent moments
of the day, then we calculate the average. In gure 5 each column is the
average of the response times on the entire test set, while the yellow line represents
the average of the three executions during the day. The results are expressed
in milliseconds and we use Rome Time Zone (CEST) to express times. Luis is
the fastest platform while Watson and Wit.ai are the slowest. Watson presents
similar response times during the day. Wit.ai's performance seems to su er some
excessive servers load during particular times of the day, in fact, it has excellent
results in the experiment performed at 11 p.m. while in the remaining two tests
it shows the slowest response times. It should be noted that all platforms have
average response times for each moment of the day under 400 ms and this is a
reasonable response time for a messaging system.
Experiment (D) We described in section 5.1 how the test sets for this
experiment are constructed. We want to evaluate the spelling errors robustness of NLU
services; in order to do so, we provide 30 datasets with an increasing number of
errors inside of the examples.</p>
      <p>As shown in gure 6 Watson is the most robust platform, it maintains an
error rate below 0.1 with less than 10 misspellings and error rate below 0.4 with
30 spelling errors. The other three platforms have similar trends. Luis maintains
an almost constant error rate close to 0.5 when the spelling errors are between
20 and 30. Wit.ai and Dialog ow appear to be the least robust platforms in this
experiment, in fact, their error rate exceeds 0.6.
In this paper, we present four di erent experiments which allow comparing from
di erent points of view the ability of di erent cloud based NLU platforms to
recognise the underlying intent of sentences, belonging to an Italian dataset,
coming from a real business case. The idea that pushed us to compare various
services is to create the best chatbot which t the company's needs. The rst
step in order to implement a chatbot is to choose the best platform on the market
through an accurate and severe comparison.</p>
      <p>In our analysis, Dialog ow shows the better overall results, however, Watson
achieves similar performance and in some cases, it overcomes Dialog ow. Luis
also has good performance, but in no case, it provides better results than services
already mentioned. In our experiment, Wit.ai provides the worst results, and we
cannot rule out this is due to the language used. Experiment (B) presents a quite
impressive result, and that is Watson and Dialog ow achieve excellent outcomes,
using just 40% of the whole training set. Experiment (C) shows that Luis is the
fastest platform to recognise the intent associated to an input sentence. All
platforms reply with an average of response times less than 400ms, this can be
considered acceptable for a chat messaging system. In experiment (D) we notice
that Watson is the most robust service when the sentences contain spelling errors.</p>
      <p>To build a chatbot able to answer questions in Italian on an e-learning
platform, the best NLU services are Watson and Dialog ow. In the speci c case of
Mega Italia Media S.p.A, the users of their e-learning platform are students with
di erent language skills and someone has just a basic level or little knowledge of
the Italian language. In this context, Watson is probably the best choice because
it is the most robust service and the performance di erence in term of F-score
is not statistically signi cant compared to Dialog ow as shown in section 6.1.</p>
      <p>As future work, we plan to increase the entire dataset size and add more
intents, aiming to build a more deep and robust analysis. We plan to de ne a
deep linguistic analysis of our dataset and a comparison of performance on other
domains. And nally, to have a complete performance evaluation, we want to
analyse not only the ability to recognise intents but also the ability to identify
entities.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Braun</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Hernandez</given-names>
            <surname>Mendez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Matthes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Langen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Evaluating natural language understanding services for conversational question answering systems</article-title>
          .
          <source>In: Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue</source>
          . pp.
          <volume>174</volume>
          {
          <fpage>185</fpage>
          . Association for Computational Linguistics, Saarbrucken,
          <source>Germany (Aug</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Canonico</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russis</surname>
            ,
            <given-names>L.D.:</given-names>
          </string-name>
          <article-title>A comparison and critique of natural language understanding tools</article-title>
          . pp.
          <volume>110</volume>
          {
          <fpage>115</fpage>
          .
          <string-name>
            <surname>CLOUD</surname>
            <given-names>COMPUTING</given-names>
          </string-name>
          <year>2018</year>
          : The Ninth International Conference on Cloud Computing, GRIDs, and
          <string-name>
            <surname>Virtualization</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maroldi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Squassina</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Automatic classi cation of radiological reports for clinical care</article-title>
          .
          <source>Artif. Intell. Medicine</source>
          <volume>91</volume>
          ,
          <volume>72</volume>
          {
          <fpage>81</fpage>
          (
          <year>2018</year>
          ). https://doi.org/10.1016/j.artmed.
          <year>2018</year>
          .
          <volume>05</volume>
          .006
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polepeddi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Jill watson: A virtual teaching assistant for online education (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haldar</surname>
          </string-name>
          , R.:
          <article-title>Applying chatbots to the internet of things: Opportunities and architectural elements</article-title>
          .
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>7</volume>
          (
          <issue>11</issue>
          ) (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eshghi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swietojanski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rieser</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Benchmarking natural language understanding services for building conversational agents</article-title>
          .
          <source>CoRR abs/1903</source>
          .05566 (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mehmood</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Leveraging multi-task learning for biomedical named entity recognition</article-title>
          . In: Alviano,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Greco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Scarcello</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.)
          <source>AI*IA 2019 - Advances in Arti cial Intelligence - XVIIIth International Conference of the Italian Association for Arti cial Intelligence, November 19-22</source>
          ,
          <year>2019</year>
          ,
          <source>Proceedings. Lecture Notes in Computer Science</source>
          , vol.
          <volume>11946</volume>
          , pp.
          <volume>431</volume>
          {
          <fpage>444</fpage>
          . Springer (
          <year>2019</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -35166-3 31
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mehmood</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>: Multi-task learning applied to biomedical named entity recognition task</article-title>
          . In: Bernardi,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Semeraro</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          , Bari, Italy,
          <source>November 13-15</source>
          ,
          <year>2019</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2481</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2019</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2481</volume>
          /paper47.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mehmood</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Combining multi-task learning with transfer learning for biomedical named entity recognition</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>176</volume>
          ,
          <issue>848</issue>
          {
          <fpage>857</fpage>
          (
          <year>2020</year>
          ),
          <source>knowledge-Based and Intelligent Information &amp; Engineering Systems: Proceedings of the 24th International Conference KES2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Putelli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Applying self-interaction attention for extracting drug-drug interactions</article-title>
          . In: Alviano,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Greco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Scarcello</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.)
          <source>AI*IA 2019 - Advances in Arti cial Intelligence - XVIIIth International Conference of the Italian Association for Arti cial Intelligence, November 19-22</source>
          ,
          <year>2019</year>
          ,
          <source>Proceedings. Lecture Notes in Computer Science</source>
          , vol.
          <volume>11946</volume>
          , pp.
          <volume>445</volume>
          {
          <fpage>460</fpage>
          . Springer (
          <year>2019</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -35166-3 32
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Putelli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olivato</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Deep learning for classi cation of radiology reports with a hierarchical schema</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>176</volume>
          ,
          <issue>349</issue>
          {
          <fpage>359</fpage>
          (
          <year>2020</year>
          ),
          <source>knowledge-Based and Intelligent Information &amp; Engineering Systems: Proceedings of the 24th International Conference KES2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Putelli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The impact of self-interaction attention on the extraction of drug-drug interactions</article-title>
          . In: Bernardi,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Semeraro</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          , Bari, Italy,
          <source>November 13-15</source>
          ,
          <year>2019</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2481</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2019</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2481</volume>
          /paper61.pdf
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rosruen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samanchuen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Chatbot utilization for medical consultant system</article-title>
          .
          <source>2018 3rd Technology Innovation Management and Engineering Science International Conference (TIMES-iCON)</source>
          pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zubani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sigalini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerevini</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          :
          <article-title>Evaluating di erent natural language understanding services in a real business case for the italian language</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>176</volume>
          ,
          <issue>995</issue>
          {
          <fpage>1004</fpage>
          (
          <year>2020</year>
          ),
          <source>knowledge-Based and Intelligent Information &amp; Engineering Systems: Proceedings of the 24th International Conference KES2020</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>