<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatically Predicting User Ratings for Conversational Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A. Cervone</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Gambi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>G. Tortoreto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E.A. Stepanov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>G. Riccardi</string-name>
          <email>giuseppe.riccardig@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Signals and Interactive Systems Lab, University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>VUI, Inc.</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Automatic evaluation models for open-domain conversational agents either correlate poorly with human judgment or require expensive annotations on top of conversation scores. In this work we investigate the feasibility of learning evaluation models without relying on any further annotations besides conversationlevel human ratings. We use a dataset of rated (1-5) open domain spoken conversations between the conversational agent Roving Mind (competing in the Amazon Alexa Prize Challenge 2017) and Amazon Alexa users. First, we assess the complexity of the task by asking two experts to re-annotate a sample of the dataset and observe that the subjectivity of user ratings yields a low upper-bound. Second, through an analysis of the entire dataset we show that automatically extracted features such as user sentiment, Dialogue Acts and conversation length have significant, but low correlation with user ratings. Finally, we report the results of our experiments exploring different combinations of these features to train automatic dialogue evaluation models. Our work suggests that predicting subjective user ratings in open domain conversations is a challenging task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. I modelli stato dell’arte per la
valutazione automatica di agenti
conversazionali open-domain hanno una scarsa
correlazione con il giudizio umano
oppure richiedono costose annotazioni oltre
al punteggio dato alla conversazione. In
questo lavoro investighiamo la possibilita`
di apprendere modelli di valutazione
attraverso il solo utilizzo di punteggi umani
dati all’intera conversazione. Il corpus
utilizzato e` composto da conversazioni
parlate open-domain tra l’agente
conversazionale Roving Mind (parte della
competizione Amazon Alexa Prize 2017) e
utenti di Amazon Alexa valutate con
punteggi da 1 a 5. In primo luogo, valutiamo
la complessita` del task assegnando a due
esperti il compito di riannotare una parte
del corpus e osserviamo come esso risulti
complesso perfino per annotatori umani
data la sua soggettivita`. In secondo luogo,
tramite un’analisi condotta sull’intero
corpus mostriamo come features estratte
automaticamente (sentimento dell’utente,
Dialogue Acts e lunghezza della
conversazione) hanno bassa, ma significativa
correlazione con il giudizio degli utenti.
Infine, riportiamo i risultati di
esperimenti volti a esplorare diverse
combinazioni di queste features per addestrare
modelli di valutazione automatica del
dialogo. Questo lavoro mostra la difficolta`
del predire i giudizi soggettivi degli utenti
in conversazioni senza un task specifico.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>We are currently witnessing a proliferation of
conversational agents in both industry and academia.
Nevertheless, core questions regarding this
technology remain to be addressed or analysed in
greater depth. This work focuses on one such
question: can we automatically predict user
ratings of a dialogue with a conversational agent?</p>
      <p>
        Metrics for task-based systems are generally
related to the successful completion of the task.
Among these, contextual appropriateness
        <xref ref-type="bibr" rid="ref3">(Danieli
and Gerbino, 1995)</xref>
        evaluates, for example, the
degree of contextual coherence of machine turns
with respect to user queries which are classified
with ternary values for slots (appropriate,
inappropriate, and ambiguous). The approach is
somewhat similar to the attribute-value matrix of the
popular PARADISE dialog evaluation framework
        <xref ref-type="bibr" rid="ref15">(Walker et al., 1997)</xref>
        , where there are matrices
representing the information exchange requirements
between the machine and users towards solving
the dialog task, as a measure of task success rate.
      </p>
      <p>Unlike task-based systems, non-task-based
conversational agents (also known as chitchat
models) do not have a specific task to accomplish (e.g.
booking a restaurant). The goal of these can
arguably be defined as the conversation itself, i.e.
the entertainment of the human it is conversing
with. Thus, human judgment is still the most
reliable evaluation tool we have for such
conversational agents. Collecting user ratings for a system,
however, is expensive and time-consuming.</p>
      <p>
        In order to deal with these issues, researchers
have been investigating automatic metrics for
nontask based dialogue evaluation. The most
popular of these metrics (e.g. BLEU
        <xref ref-type="bibr" rid="ref10">(Papineni et al.,
2002)</xref>
        , METEOR
        <xref ref-type="bibr" rid="ref1">(Banerjee and Lavie, 2005)</xref>
        ) rely
on surface text similarity (word overlaps) between
machine and reference responses to the same
utterances. Notwithstanding their popularity, such
metrics are hardly compatible with the nature of
human dialogue, since there could be multiple
appropriate responses to the same utterance with no
word overlap. Moreover, these metrics correlate
weakly with human judgments
        <xref ref-type="bibr" rid="ref7">(Liu et al., 2016)</xref>
        .
      </p>
      <p>
        Recently, a few studies proposed metrics
having a better correlation with human judgment.
ADEM
        <xref ref-type="bibr" rid="ref8">(Lowe et al., 2017)</xref>
        is a model trained on
appropriateness scores manually annotated at the
response-level. Venkatesh et al. (2017) and Guo
et al. (2017) combine multiple metrics, each
capturing a different aspect of the interaction, and
predict conversation-level ratings. In particular,
Venkatesh et al. (2017) shows the importance of
metrics such as coherence, conversational depth
and topic diversity, while Guo et al. (2017)
proposes topic-based metrics. However, these
studies require extensive manual annotation on top of
conversation-level ratings.
      </p>
      <p>In this work, we investigate non-task based
dialogue evaluation models trained without relying
on any further annotations besides
conversationlevel user ratings. Our goal is twofold:
investigating conversation features which characterize good
interactions with a conversational agent and
exploring the feasibility of training a model able to
predict user ratings in such context.</p>
      <p>
        In order to do so, we utilize a dataset of
nontask based spoken conversations between
Amazon Alexa users and Roving Mind
        <xref ref-type="bibr" rid="ref2">(Cervone et al.,
2017)</xref>
        , our open-domain system for the Amazon
Alexa Prize Challenge 2017
        <xref ref-type="bibr" rid="ref12 ref14">(Ram et al., 2017)</xref>
        .
As an upper bound for the rating prediction task,
we re-annotate a sample of the corpus using
experts and analyse the correlation between expert
and user ratings. Afterwards, we analyse the
entire corpus using well-known automatically
extractable features (user sentiment, Dialogue Acts
(both user and machine), conversation length and
average user turn length), which show a low, but
still significant correlation with user ratings. We
show how different combinations of these
features together with a LSA representation of the
user turns can be used to train a regression model
whose predictions also yield a low, but significant
correlation with user ratings. Our results indicate
the difficulty of predicting how users might rate
interactions with a conversational agent.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Collection</title>
      <p>The dataset analysed in this paper was collected
over a period of 27 days during the Alexa Prize
2017 semifinals and consists of conversations
between our system Roving Mind and Amazon
Alexa users of the United States. The users could
end the conversation whenever they wanted, using
a command. At the end of the interaction users
were asked to rate a conversation on a 1 (not
satisfied at all) to 5 (very satisfied) Likert scale. Out
of all the rated conversations, we selected the ones
longer than 3 turns to yield 4,967 conversations.
Figure 1 shows the distribution (in percentages)
of the ratings in our dataset. The large majority of
conversations are between a system and a
“firsttime” users, as only 5.25% of users had more than
one conversation.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <p>In this section we describe conversation
representation features, experimentation, and evaluation
methodologies used in the paper.
3.1</p>
      <sec id="sec-4-1">
        <title>Conversation Representation Features</title>
        <p>Since in the competition the objective of the
system was to entertain users, we expect the ratings
to reflect how much they have enjoyed the
interaction. User “enjoyment” can be approximated
)
%
(e 30 28
m
u
l
o
V
ion 20
t
a
s
r
e
v
n
oC 10
14
1</p>
        <p>Users
Experts</p>
        <p>
          All ratings
14
11
9
11
7
5
13
34
using different metrics that do not require manual
annotation, such as conversation length (in turns),
mean turn length (in words), assuming that the
more users enjoy the conversation the longer they
talk; sentiment polarity – hypothesizing that
enjoyable conversations should carry a more
positive sentiment. While length metrics are
straightforward to compute, the sentiment score is
computed using a lexicon-based approach
          <xref ref-type="bibr" rid="ref6">(Kennedy
and Inkpen, 2006)</xref>
          .
        </p>
        <p>
          Another representation that could shed a light
on enjoyable conversations is Dialogue Acts (DA)
of user and machine utterances. DAs are
frequently used as a generic representation of intents
and the considered labels often include thanking,
apologies, opinions, statements and alike.
Relative frequencies of these tags potentially can be
useful to distinguish good and bad conversations.
The DA tagger we use is the one described in
Mezza et al. (2018) trained on the Switchboard
Dialogue Acts corpus
          <xref ref-type="bibr" rid="ref13">(Stolcke et al., 2000)</xref>
          , a subset
of Switchboard
          <xref ref-type="bibr" rid="ref4">(Godfrey et al., 1992)</xref>
          annotated
with DAs (42 categories), using Support Vector
Machines. The user and machine DAs are
considered as separate vectors and assessed both
individually and jointly.
        </p>
        <p>Additional to Dialogue Acts, sentiment and
length features, we experiment with word-based
text representation. Latent Semantic Analysis
(LSA) is used to convert a conversation to a
vector. First, we construct a word-document
cooccurrence matrix and normalize it. Then, we
reduce the dimensionality to 100 by applying
Singular Value Decomposition (SVD).
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Correlation Analysis Methodology</title>
        <p>The two widely used correlation metrics are
Pearson correlation coefficient (PCC) and Spearman’s
rank correlation coefficient (SRCC). While the
former evaluates the linear relationship between
variables, the latter evaluates the monotonic one.</p>
        <p>The metrics are used to assess correlations of
different conversation features, such as sentiment
score or conversation length, with the provided
human ratings for those conversations; as well as to
assess the correlation of the predicted scores of the
regression models to those ratings. For the
assessment of the correlation of both features and
regression models raw rating predictions are used.
3.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Prediction Methodology</title>
        <p>
          Using the conversation features described above,
we train regression models to predict human
ratings. We experiment with both Linear Regression
and Support Vector Regression (SVR) with radial
basis function (RBF) kernel using scikit-learn
          <xref ref-type="bibr" rid="ref11">(Pedregosa et al., 2011)</xref>
          . Since the latter consistently
outperforms the former, we report only the results
for the SVR. The performance of the regression
models is evaluated using the standard metrics
of Root Mean Squared Error (RMSE) and Mean
Absolute Error (MAE). Additionally, we compute
Pearson and Spearman’s Rank Correlation
Coefficients for the predictions with respect to the
reference human ratings.
        </p>
        <p>We experiment with the 10-fold
crossvalidation setting. The performance of the
regression models is compared to two baselines:
(1) mean baseline, where all instances in the
testing fold are assigned as a score the mean of
the training set ratings, and (2) chance baseline,
where an instance is randomly assigned a rating
from 1 to 5 with respect to their distribution in
the training set. The models are compared for
statistical significance to these baselines using
paired two-tail T-test with p &lt; 0:05. In Section
6 we report average RMSE and MAE as well as
average correlation coefficients.</p>
        <p>Exp 1 vs. Exp 2
Exp 1 vs. Users
Exp 2 vs. Users
Since human ratings are inherently subjective, and
different users can rate the same conversation
differently, it is difficult to expect the models to yield
perfect correlations or very low RMSE and MAE.
In order to test this hypothesis two human experts
(members of our Alexa Prize team) were asked to
rate a random subset of the corpus (100
conversations). The rating distributions for both experts
and users on the sample is reported in Figure 1.
We observe that expert ratings tend to be closer to
the middle of the Likert scale (i.e. from 2 to 4),
while users had more conversations with ratings at
both extremes of the scale (i.e. 1 and 5).</p>
        <p>The RMSE, MAE and Pearson and Spearman’s
rank correlation coefficients of expert and user
ratings are reported in Table 1. We observe that
the experts tend to agree with each other more
than they agree individually with users, since
compared to each other the experts have the highest
Pearson and Spearman correlation scores (0.705
and 0.694, respectively) and the lowest RMSE and
MAE (0.875 and 0.660, respectively). The fact
that expert ratings do not correlate with user
ratings as well as they correlate among themselves,
confirms the difficulty of the task of predicting
subjective user ratings even for humans.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Correlation Analysis Results</title>
      <p>The results of the correlation analysis are reported
in Table 2. From the table, we can observe
that conversation length has a positive correlation
with human judgment, while the average user turn
length has a negative correlation. The positive
correlation with conversation length confirms the
expectation that users tend to have longer
conversations with the system when they enjoy it. The
negative correlation with average user turn length, on
the other hand, is unexpected. As expected,
sentiment score has a significant positive correlation
with human judgments.</p>
      <p>Feature
Conversation Length
Av. User Turn Length
User Sentiment
yes-answer
appreciation
thanking
action-directive
statement-non-opinion
...</p>
      <p>Due to the space considerations, we report only
a portion of the DAs that have significant
correlations with human ratings. The analysis confirms
our expectations that user DAs, such as thanking
and appreciation, have significant positive
correlations. We also observe that the action-directive
DA has a negative correlation. Since this DA label
covers the turns where a user issues control
commands to the system, we hypothesize this
correlation could be due to the fact that in such cases
users were using a task-based approach with our
system which was instead designed for chitchat
and might therefore feel disappointed (e.g.
requesting the Roving Mind system to perform
actions it was not designed to perform, such as
playing music).</p>
      <p>Regarding machine DAs, we observe that even
though some DAs exhibit significant correlations,
overall they are lower than user DAs. In particular,
yes-no-question has a significant positive
correlation with human judgments, indicating that some
users appreciate machine initiative in the
conversation. The analysis confirms the utility of length
and sentiment features, as well as the importance
of some DAs (generic intents) for estimating user
ratings.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Prediction Results</title>
      <p>The results of the experiments using 10-fold
crossvalidation and Support Vector Regression are
reported in Table 3. We report performances of each
feature representation is isolation and their
combinations. We consider two baselines – chance and
mean. For the chance baseline an instance is
randomly assigned a rating with respect to the
training set distribution. For the mean baseline, on the
other hand, all the instances are assigned the mean
of the training set as a rating. The mean
baseline yields better RMSE and MAE scores;
consequently, we compare the regression models to it.</p>
      <p>Sentiment and length features (conversation and
average user turn) both yield RMSE higher than
the mean baseline and MAE significantly lower
than it. Nonetheless, their predictions have
significant positive correlations with reference
human ratings. The picture is similar for the
models trained on user and machine DAs alone and
their combination. The RMSE scores are higher
or insignificantly lower and MAE scores are
significantly lower than the mean baseline.</p>
      <p>For the LSA representation of conversations we
consider ngram sizes between 1 and 4. The
representation that considers 4-grams and the SVD
dimension of 100 yields better performances; thus,
we report the performances of this models only,
and use it for feature combination experiments.
The LSA model yields significantly lower error
both in terms of RMSE and MAE. Additionally,
the correlation of the predictions is higher than for
the other features (and combinations).</p>
      <p>The regression model trained on all features but
LSA, yields performances significantly better than
the mean baseline. However, they are inferior to
that of LSA alone. Combination of all the
features retains the best RMSE of the LSA model, but
achieves a little worse MAE score. While it yields
the best Pearson and Spearman’s rank correlation
coefficients among all the models, the difference
from LSA only model is not statistically relevant
using Fisher r-to-z transformation.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>In this work we experimented with a set of
automatically extractable black-box features which
correlate with the human perception of the quality
of interactions with a conversational agent.
Furthermore, we showed how these features can be
combined to train automatic non-task-based
dialogue evaluation models which correlate with
human judgments without further expensive
annotations.</p>
      <p>The results of our experiments and analysis
contribute to the body of observations that indicate
that there still remains a lot of research to be done
in order to understand characteristics of enjoyable
conversations with open-domain non-task oriented
agents. In particular, our analysis of expert vs.
user ratings suggests that the task of estimating
subjective user ratings is a difficult one, since the
same conversation might be rated quite differently.</p>
      <p>For the future work, we plan to extend our
corpus to include interactions with multiple
conversational agents and task-based systems, as well as to
explore other features that might be relevant for
assessing human judgment of interaction with a
conversational agent (e.g. emotion recognition).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Satanjeev</given-names>
            <surname>Banerjee</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alon</given-names>
            <surname>Lavie</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Meteor: An automatic metric for mt evaluation with improved correlation with human judgments</article-title>
          .
          <source>In Proceedings of the acl workshop on</source>
          intrinsic and
          <article-title>extrinsic evaluation measures for machine translation and/or summarization</article-title>
          , volume
          <volume>29</volume>
          , pages
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Cervone</surname>
          </string-name>
          , Giuliano Tortoreto, Stefano Mezza, Enrico Gambi, and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Riccardi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Roving mind: a balancing act between opendomain and engaging dialogue systems</article-title>
          .
          <source>In Alexa Prize Proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Morena</given-names>
            <surname>Danieli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Gerbino</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Metrics for evaluating dialogue strategies in a spoken language system</article-title>
          .
          <source>In Proceedings of the 1995 AAAI spring symposium on Empirical Methods in Discourse Interpretation and Generation</source>
          , volume
          <volume>16</volume>
          , pages
          <fpage>34</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>John J Godfrey</given-names>
            , Edward C Holliman, and
            <surname>Jane McDaniel</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Switchboard: Telephone speech corpus for research and development</article-title>
          .
          <source>In Acoustics, Speech, and Signal Processing</source>
          ,
          <year>1992</year>
          . ICASSP-
          <volume>92</volume>
          ., 1992 IEEE International Conference on, volume
          <volume>1</volume>
          , pages
          <fpage>517</fpage>
          -
          <lpage>520</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Fenfei</given-names>
            <surname>Guo</surname>
          </string-name>
          , Angeliki Metallinou, Chandra Khatri, Anirudh Raju, Anu Venkatesh, and
          <string-name>
            <given-names>Ashwin</given-names>
            <surname>Ram</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Topic-based evaluation for conversational bots</article-title>
          .
          <source>In NIPS 2017 Conversational AI workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Alistair</given-names>
            <surname>Kennedy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Diana</given-names>
            <surname>Inkpen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Sentiment classification of movie reviews using contextual valence shifters</article-title>
          .
          <source>Computational intelligence</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ):
          <fpage>110</fpage>
          -
          <lpage>125</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Chia-Wei</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and
          <string-name>
            <given-names>Joelle</given-names>
            <surname>Pineau</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>2122</fpage>
          -
          <lpage>2132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Lowe</surname>
          </string-name>
          , Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier,
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Joelle</given-names>
            <surname>Pineau</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards an automatic turing test: Learning to evaluate dialogue responses</article-title>
          .
          <source>In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>1116</fpage>
          -
          <lpage>1126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Mezza</surname>
          </string-name>
          , Alessandra Cervone, Giuliano Tortoreto, Evgeny A.
          <string-name>
            <surname>Stepanov</surname>
            , and
            <given-names>Giuseppe</given-names>
          </string-name>
          <string-name>
            <surname>Riccardi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Iso-standard domain-independent dialogue act tagging for conversational agents</article-title>
          .
          <source>In Proceedings of COLING</source>
          <year>2018</year>
          ,
          <source>the 27th International Conference on Computational Linguistics: Technical Papers</source>
          , pages
          <fpage>3539</fpage>
          -
          <lpage>3551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Kishore</given-names>
            <surname>Papineni</surname>
          </string-name>
          , Salim Roukos, Todd Ward, and
          <string-name>
            <given-names>WeiJing</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In Proceedings of the 40th annual meeting on association for computational linguistics</source>
          , pages
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Ashwin</given-names>
            <surname>Ram</surname>
          </string-name>
          , Rohit Prasad, Chandra Khatri, Anu Venkatesh, Raefer Gabriel, Qing Liu, Jeff Nunn, Behnam Hedayatnia, Ming Cheng, Ashish Nagar, Eric King, Kate Bland, Amanda Wartick, Yi Pan, Han Song, Sk Jayadevan, Gene Hwang, and
          <string-name>
            <given-names>Art</given-names>
            <surname>Pettigrue</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Conversational ai: The science behind the alexa prize</article-title>
          .
          <source>In Alexa Prize Proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Stolcke</surname>
          </string-name>
          , Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, Daniel Jurafsky, Paul Taylor, Rachel Martin,
          <string-name>
            <surname>Carol Van</surname>
            Ess-Dykema, and
            <given-names>Marie</given-names>
          </string-name>
          <string-name>
            <surname>Meteer</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Dialogue act modeling for automatic tagging and recognition of conversational speech</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>26</volume>
          (
          <issue>3</issue>
          ):
          <fpage>339</fpage>
          -
          <lpage>373</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Anu</given-names>
            <surname>Venkatesh</surname>
          </string-name>
          , Chandra Khatri, Ashwin Ram, Fenfei Guo, Raefer Gabriel, Ashish Nagar, Rohit Prasad, Ming Cheng, Benham Hedayatnia, Angeliki Metallinou, Rahul Goel,
          <string-name>
            <given-names>Shaohua</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Anirudh</given-names>
            <surname>Raju</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>On evaluating and comparing conversational agents</article-title>
          .
          <source>In NIPS 2017 Conversational AI workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Marilyn A Walker</surname>
          </string-name>
          ,
          <article-title>Diane J Litman, Candace A Kamm,</article-title>
          and
          <string-name>
            <given-names>Alicia</given-names>
            <surname>Abella</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Paradise: A framework for evaluating spoken dialogue agents</article-title>
          .
          <source>In Proceedings of the eighth conference on European chapter of the Association for Computational Linguistics</source>
          , pages
          <fpage>271</fpage>
          -
          <lpage>280</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>