<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Markovian Kernel-based Approach for itaLIan Speech acT labEliNg</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Basili</string-name>
          <email>basilig@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Roma</institution>
          ,
          <addr-line>Tor Vergata Via del Politecnico 1, Rome, 00133</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>English. This paper describes the
UNITOR system that participated to the
itaLIan Speech acT labEliNg task within
the context of EvalIta 2018. A
Structured Kernel-based Support Vector
Machine has been here applied to make the
classification of the dialogue turns
sensitive to the syntactic and semantic
information of each utterance, without relying on
any task-specific manual feature
engineering. Moreover, a specific Markovian
formulation of the SVM is adopted, so that
the labeling of each utterance depends on
speech acts assigned to the previous turns.
The UNITOR system ranked first in the
competition, suggesting that the
combination of the adopted structured kernel and
the Markovian modeling is beneficial.</p>
      <p>Italian. Questo lavoro descrive il sistema
UNITOR che ha partecipato all’itaLIan
Speech acT labEliNg task organizzato
nell’ambito di EvalIta 2018. Il
sistema e` basato su una Structured
Kernelbased Support Vector Machine (SVM) che
rende la classificazione dei turni di
dialogo dipendente dalle informazioni
sintattiche e semantiche della frase, evitando la
progettazione di alcuna feature specifica
per il task. Una specifica formulazione
Markoviana dell’algoritmo di
apprendimento SVM permette inoltre di etichettare
ciascun turno in funzione delle
classificazioni dei turni precedenti. Il sistema
UNITOR si e´ classificato al primo posto
nella competizione, e questo conferma
come la combinazione della funzione
kernel e del modello Markoviano adottati sia
molto utile allo sviluppo di sistemi di
dialoghi robusti.</p>
      <p>
        A dialogue agent is designed to interact and
communicate with other agents, in a coherent manner,
not just through one-shot messages, but according
to sequences of meaningful and related messages
on an underlying topic or in support to an overall
goal
        <xref ref-type="bibr" rid="ref13">(Traum, 1999)</xref>
        . These communications are
seen not just as transmitting information but as
actions that change the state of the world, e.g., the
mental states of the agents involved in the
conversation, as well as the state, or context, of the
dialogue. In other words, speech act theory
allows to design an agent in order to place its
communication within the same general framework as
the agent’s actions. In such a context, the robust
recognition of the speech acts characterizing an
interaction is crucial for the design and deployment
of artificial dialogue agents.
      </p>
      <p>
        This specific task has been considered in the
first itaLIan Speech acT labEliNg task at EvalIta
(iLISTEN,
        <xref ref-type="bibr" rid="ref9">(Novielli and Basile, 2018)</xref>
        ): given a
dialogue between an agent and a user, the task
consists in automatically annotating the user’s
dialogue turns with speech act labels, i.e. with the
communicative intention of the speaker, such as
statement, request for information or agreement.
Table 1 reports the full set of speech act labels
considered in the challenge, with definition and
examples.
      </p>
      <p>
        In this paper, the UNITOR system
participating in the iLISTEN task within the EvalIta 2018
evaluation campaign is described. The system
realize the classification task through a Structured
Kernel-based Support Vector Machine
        <xref ref-type="bibr" rid="ref14">(Vapnik,
1998)</xref>
        classifier. A structured kernel, namely a
Smoothed Partial Tree Kernel (SPTK,
        <xref ref-type="bibr" rid="ref3">(Croce et
al., 2011)</xref>
        ) is applied in order to make the
classification of each utterance dependent from the
syntactic and semantic information of each individual
utterance.
      </p>
      <p>
        Since turns are not observed in isolation, but
immersed in a dialogue, we adopted a Markovian
forSpeech Act
OPENING
CLOSING
INFO-REQUEST
SOLICIT-REQ-CLARIF
STATEMENT
GENERIC-ANSWER
AGREE
REJECT
KIND-ATT-SMALLTALK
mulation of SVM, namely SVMhmm
        <xref ref-type="bibr" rid="ref1">(Altun et al.,
2003)</xref>
        so that the classification of the ith utterance
also depends from the dialogue act assigned at the
i 1th utterance.
      </p>
      <p>The UNITOR system ranked first in the
competition, suggesting that the combination of the
adopted structured kernel and the Markovian
learning algorithm is beneficial.</p>
      <p>In the rest of the paper, Section 2 describes the
adopted machine learning method and the
underlying semantic kernel functions. In Section 3, the
performance measures of the system are reported
while Section 4 derives the conclusions.
2</p>
      <p>A Markovian Kernel-based Approach
The UNITOR system implements a
Markovian formulation of the Support Vector Machine
(SVM) learning algorithm. The SVM adopts a
structured kernel function in order to estimate the
syntactic and semantic similarity between
utterances in a dialogue, without the need of any
taskspecific manual feature engineering (only the
dependency parse of each sentence is required). In
the rest of this section, first the learning algorithm
is presented, then the adopted kernel method is
discussed.
2.1</p>
      <p>A Markovian Support Vector Machine
The aim of a Markovian formulation of SVM is to
make the classification of a input example xi 2 Rn
(belonging to a sequence of examples) dependent
on the label assigned to the previous elements in a
history of length m, i.e., xi m; : : : ; xi 1.</p>
      <p>In our classification task, a dialogue is a
sequence of utterances x = (x1; : : : ; xs) each of
them representing an example xi, i.e., the specific
i-th utterance. Given the corresponding sequence
of expected labels y = (y1; : : : ; ys), a sequence of
m step-specific labels (from a dictionary of d
different dialogue acts) can be retrieved, in the form
yi m; : : : ; yi 1.</p>
      <p>In order to make the classification of xi
dependent also from the previous decisions, we
augment the feature vector of xi introducing a
projection function m(xi) 2 Rmd that associates to
each example a md dimensional feature vector
where each dimension set to 1 corresponds to the
presence of one of the d possible labels observed
in a history of length m, i.e. m steps before the
target element xi.</p>
      <p>In order to apply a SVM, a projection function
m( ) can be defined to consider both the
observations xi and the transitions m(xi) by
concatenating the two representations as follows:</p>
      <p>m(xi) = xi jj m(xi)
with m(xi) 2 Rn+md. Notice that the symbol
jj here denotes the vector concatenation, so that
m(xi) does not interfere with the original feature
space, where xi lies.</p>
      <p>
        Kernel-based methods can be applied in order to
model meaningful representation spaces,
encoding both the feature representing individual
examples together with the information about the
transitions. According to kernel-based learning
        <xref ref-type="bibr" rid="ref12">(ShaweTaylor and Cristianini, 2004)</xref>
        , we can define a
kernel function Km(xi; zj ) between a generic item of
a sequence xi and another generic item zj from the
same or a different sequence, parametric in the
history length m: it surrogates the product between
m( ) such that:
      </p>
      <p>Km(xi; zj ) =</p>
      <p>m(xi) m(zj ) =
= Kobs(xi; zj ) + Ktr
m(xi); m(zj )
In other words, we define a kernel that is the
linear combination of two further kernels: Kobs
operating over the individual examples xi and a Ktr
operating over the feature vectors encoding the
involved transitions. It is worth noticing that Kobs
does not depend on the position nor the context of
individual examples in line with Markov
assumption characterizing a large class of these generative
models, e.g. HMM. For simplicity, we define Ktr
as a linear kernel between input instances, i.e. a
dot-product in the space generated by m( ):
Km(xi; zj ) = Kobs(xi; xj ) +
m(xi) m(zj )</p>
      <p>At training time, we use the kernel-based
SVM in a One-Vs-All schema over the feature
space derived by Km( ; ). The learning
process provides a family of classification functions
f (xi; m) Rn+md Rd, which associate each
xi to a distribution of scores with respect to the
different d labels, depending on the context size
m. At classification time, all possible sequences
y 2 Y+ should be considered in order to
determine the best labeling y^, where m is the size of
the history used to enrich xi, that is:
y^ = aryg2mY+axf</p>
      <p>X f (xi; m)g
i=1:::m</p>
      <p>In order to reduce the computational cost, a
Viterbi-like decoding algorithm is adopted1. The
next section defines the kernel function Kobs
applied to specific utterances.
2.2 Structured Kernel Methods for Speech</p>
      <p>
        Act Labeling
Several NLP tasks require the explorations of
complex semantic and syntactic phenomena.
For instance, in Paraphrase Detection, verifying
whether two sentences are valid paraphrases
involves the analysis of some rewriting rules in
which the syntax plays a fundamental role. In
Question Answering, the syntactic information is
crucial, as largely demonstrated in
        <xref ref-type="bibr" rid="ref3">(Croce et al.,
2011)</xref>
        .
      </p>
      <p>1When applying f (xi; m) the classification scores are
normalized through a softmax function and probability scores
are derived.
AUX</p>
      <p>ADVMOD</p>
      <p>DET
DET:POSS
Devo controllare maggiormente il mio peso
AUX VERB ADV DET DET NOUN</p>
      <p>
        A natural approach to exploit such linguistic
information consists in applying kernel methods
        <xref ref-type="bibr" rid="ref10 ref12">(Robert Mu¨ller et al., 2001; Shawe-Taylor and
Cristianini, 2004)</xref>
        on structured representations of
data objects, e.g., documents. A sentence s can
be represented as a parse tree2 that expresses the
grammatical relations implied by s. Tree kernels
(TKs)
        <xref ref-type="bibr" rid="ref2">(Collins and Duffy, 2001)</xref>
        can be employed
to directly operate on such parse trees, evaluating
the tree fragments shared by the input trees. This
operation corresponds to a dot product in the
implicit feature space of all possible tree fragments.
      </p>
      <p>Whenever the dot product is available in the
implicit feature space, kernel-based learning
algorithms, such as SVMs, can operate in order to
automatically generate robust prediction models.
TKs thus allow to estimate the similarity among
texts, directly from sentence syntactic structures,
that can be represented by parse trees. The
underlying idea is that the similarity between
two trees T1 and T2 can be derived from the
number of shared tree fragments. Let the set
T = ft1; t2; : : : ; tjT jg be the space of all the
possible substructures and i(n2) be an indicator
function that is equal to 1 if the target ti is rooted
at the node n2 and 0 otherwise. A tree-kernel
function over T1 and T2 is defined as follows:
T K(T1; T2) = Pn12NT1 Pn22NT2 (n1; n2)
where NT1 and NT2 are the sets of
nodes of T1 and T2 respectively, and
(n1; n2) = PjT j</p>
      <p>k=1 k(n1) k(n2) which
computes the number of common fragments between
trees rooted at nodes n1 and n2. The feature space
generated by the structural kernels obviously
depends on the input structures. Notice that
different tree representations embody different
linguistic theories and may produce more or less
effective syntactic/semantic feature spaces for a
2Parse trees can be extracted using automatic parsers. In
our experiments, we used SpaCy https://spacy.io/.
given task.</p>
      <p>
        Dependency grammars produce a significantly
different representation which is exemplified in
Figure 1. Since tree kernels are not tailored to
model the labeled edges that are typical of
dependency graphs, these latter are rewritten into
explicit hierarchical representations. Different
rewriting strategies are possible, as discussed in
        <xref ref-type="bibr" rid="ref3">(Croce et al., 2011)</xref>
        : a representation that is shown
to be effective in several tasks is the Grammatical
Relation Centered Tree (GRCT) illustrated in
Figure 2: the PoS-Tags are children of grammatical
function nodes and direct ancestors of their
associated lexical items.
      </p>
      <p>
        Different tree kernels can be defined
according to the types of tree fragments considered in
the evaluation of the matching structures. In the
Subtree Kernel
        <xref ref-type="bibr" rid="ref15">(Vishwanathan and Smola, 2002)</xref>
        ,
valid fragments are only the grammatically well
formed and complete subtrees: every node in a
subtree corresponds to a context free rule whose
left hand side is the node label and the right hand
side is completely described by the node
descendants. Subset trees are exploited by the
Subset Tree Kernel
        <xref ref-type="bibr" rid="ref2">(Collins and Duffy, 2001)</xref>
        , which
is usually referred to as Syntactic Tree Kernel
(STK); they are more general structures since their
leaves can be non-terminal symbols. The subset
trees satisfy the constraint that grammatical rules
cannot be broken and every tree exhaustively
represents a CFG rule. Partial Tree Kernel (PTK)
        <xref ref-type="bibr" rid="ref8">(Moschitti, 2006)</xref>
        relaxes this constraint
considering partial trees, i.e., fragments generated by the
application of partial production rules (e.g.
sequences of non terminal with gaps). The strict
constraint imposed by the STK may be
problematic especially when the training dataset is small
and only few syntactic tree configurations can be
observed. The Partial Tree Kernel (PTK)
overcomes this limitation, and usually leads to higher
accuracy, as shown in
        <xref ref-type="bibr" rid="ref8">(Moschitti, 2006)</xref>
        .
      </p>
      <p>OBJ
DET DET:POSS NOUN
il::d mio::d</p>
      <p>
        Capitalizing lexical information in Convolution
Kernels. The tree kernels introduced above
perform a hard match between nodes when
comparing two substructures. In NLP tasks, when nodes
are words, this strict requirement reflects in a too
strict lexical constraint, that poorly reflects
semantic phenomena, such as the synonymy of
different words or the polysemy of a lexical entry. To
overcome this limitation, we adopt Distributional
models of Lexical Semantics
        <xref ref-type="bibr" rid="ref11 ref7">(Sahlgren, 2006;
Mikolov et al., 2013)</xref>
        to generalize the meaning
of individual words by replacing them with
geometrical representations (also called Word
Embeddings) that are automatically derived from the
analysis of large-scale corpora. These
representations are based on the idea that words occurring in
the same contexts tend to have similar meaning:
the adopted distributional models generate
vectors that are similar when the associated words
exhibit a similar usage in large-scale document
collections. Under this perspective, the distance
between vectors reflects semantic relations between
the represented words, such as paradigmatic
relations, e.g., quasi-synonymy. These word spaces
allow to define meaningful soft matching between
lexical nodes, in terms of the distance between
their representative vectors. As a result, it is
possible to obtain more informative kernel functions
which are able to capture syntactic and semantic
phenomena through grammatical and lexical
constraints.
      </p>
      <p>
        The Smoothed Partial Tree Kernel (SPTK)
described in
        <xref ref-type="bibr" rid="ref3">(Croce et al., 2011)</xref>
        exploits this idea
extending the PTK formulation with a similarity
function between nodes:
      </p>
      <p>SP T K (n1; n2) =
SP T K (n1; n2) =
(n1; n2)
(n1; n2) ; if n1 and n2 are leaves
2+
+</p>
      <p>X</p>
      <p>l(I~1)
d(I~1)+d(I~2) Y
I~1;I~2:l(I~1)=l(I~2)
k=1</p>
      <p>SP T K cn1 (i1k); cn2 (i2k)
(1)
In the SPTK formulation, the similarity function
(n1; n2) between two nodes n1 and n2 can be
defined as follows:
if n1 and n2 are both lexical nodes, then
(n1; n2) = LEX (n1; n2) = ~vn1 ~vn2 .
k~vn1 kk~vn2 k
It is the cosine similarity between the word
vectors ~vn1 and ~vn2 associated with the
labels of n1 and n2, respectively. is called
terminal factor and weighs the contribution
of the lexical similarity to the overall kernel
computation.
else if n1 and n2 are nodes sharing the same
label, then (n1; n2) = 1.</p>
      <p>else (n1; n2) = 0.</p>
      <p>
        In the challenge we adopt the SPTK in order to
implement the Kobs function used in the Markovian
SVM. This kernel in fact has been showed very
robust in the classification of (possibly short)
sentences, such as in Question Classification
        <xref ref-type="bibr" rid="ref3">(Croce
et al., 2011)</xref>
        or Semantic Role Labeling
        <xref ref-type="bibr" rid="ref3 ref4">(Croce et
al., 2011; Croce et al., 2012)</xref>
        .
      </p>
      <p>
        Dataset
train
test
complete
#dialogues
In iLISTEN, the reference dataset includes the
transcriptions of 60 dialogues amounting to about
22; 000 words. The detailed statistics regarding
dialogues and turns in the train and test dataset are
reported in Table 2.
In the proposed classification workflow each
utterance from the training/test material is processed
through the SpaCy3 dependency parser whose
outputs are automatically converted into GRCT
structures4 discussed in Section 2. These structures are
used within the Markovian SVM implemented in
KeLP5
        <xref ref-type="bibr" rid="ref5 ref6">(Filice et al., 2015; Filice et al., 2018)</xref>
        .
      </p>
      <p>The learning algorithm is based on a SPTK
combined with a One-Vs-All multi-classification
schema adopted to assign individual utterances to
the targeted classes. All the parameters of the
3It is freely available for several languages (including
Italian) at https://spacy.io/</p>
      <p>4Utterances may include more than one sentence and
potentially generate different trees. These cases are handled as
follows: all trees after the first one are linked through the
creation of an artificial link between their roots and the root of
the tree generated by the first sentence.</p>
      <p>5http://www.kelp-ml.org/?page_id=215
specific kernel (i.e. the contribution of the
lexical nodes in the overall computation) and of the
SVM algorithm have been tuned via 10-cross fold
validation over the training set. In the Markovian
SVM, a history of m = 1 previous steps allowed
to achieve the best results during the
parameterization step. Given the limited size of the dataset,
higher values of m led to sparse representation of
the transitions m that are not helpful. As a
multiclassification task, results are reported in terms
of precision (P), recall (R) and F1-score with
respect to the gold standard, as reported in Table 3.
These are averaged across each utterance
(microstatistics) and per class (macro-statistics).</p>
      <p>Among the two submitted systems, UNITOR
(reported on top of the table) achieved best
results, both considering micro and macro statistics,
where a F1 of 0:733 and 0:653 are achieved,
respectively. These results are higher with respect
to the other participant (namely System2) and far
higher than the baseline (that confirms the
difficulty of the task).</p>
      <p>Given the unbalanced number of examples
for each class, UNITOR achieves higher results
w.r.t. the micro statistics, while lower results are
achieved w.r.t. classes with a reduced number of
examples. The confusion matrix reported in
Table 4 shows that some recurrent misclassifications
(e.g. the STATEMENT class with respect to the
REJECT class) need to be carefully addressed.
Clearly this is a very challenging task, also for
the annotators, where the differences between the
speech act is not strictly defined: as for example,
given the stimulus of the system “Bisognerebbe
mangiare solo se si ha fame, ed aspettare che la
digestione sia completata prima di assumere altri
cibi6” the answer “a volte il lavoro non mi
permette di mangiare con ritmi regolari!7” should
be labeled as REJECT while the system provides
STATEMENT. Overall this results is
straightforward, also considering that the system did not
required any task specific feature modeling, but
the adopted structured kernel based method allows
capitalizing the syntactic and semantic
information useful for the task. The only requirement
of the system is the availability of a dependency
parser.</p>
      <p>6Translated in English: “You should only eat if you are
hungry, and wait until digestion is complete before eating
again.”</p>
      <p>7Translated in English: “sometimes work doesn’t allow
me to eat at a regular pace!”</p>
    </sec>
    <sec id="sec-2">
      <title>Conclusions</title>
      <p>In this paper the description of the UNITOR
system participating to the iLISTEN task at EvalIta
2018 has been provided. The system ranked first
in all the evaluations. Thus, the proposed
classification strategy shows the beneficial impact of the
combination of a structured kernel-based method
with a Markovian classifier, capitalizing the
contribution of the dialogue modeling in deciding the
speech act of individual sentences. One
important finding of this evaluation is that a quite
robust speech classifier can be obtained with
almost no requirement in term of task-specific
feature and system engineering: results are
appealing mostly considering the reduced size of the
dataset. Further work is needed to improve the
overall F1 scores, possibly extending the adopted
kernel function by addressing other dimensions
of the linguistic information or also making the
kernel more sensitive to task-specific knowledge.
Also the combination of the adopted strategy with
recurrent neural approaches is an interesting
research direction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Y.</given-names>
            <surname>Altun</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Tsochantaridis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Hidden Markov support vector machines</article-title>
          .
          <source>In Proceedings of the International Conference on Machine Learning.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Collins</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nigel</given-names>
            <surname>Duffy</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Convolution kernels for natural language</article-title>
          .
          <source>In Proceedings of Neural Information Processing Systems</source>
          (NIPS'
          <year>2001</year>
          ), pages
          <fpage>625</fpage>
          -
          <lpage>632</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alessandro Moschitti, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Structured lexical similarity via convolution kernels on dependency trees</article-title>
          .
          <source>In Proceedings of EMNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alessandro Moschitti, Roberto Basili, and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Palmer</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Verb classification using distributional similarity in syntactic and semantic structures</article-title>
          .
          <source>In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Simone</given-names>
            <surname>Filice</surname>
          </string-name>
          , Giuseppe Castellucci, Danilo Croce, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Kelp: a kernelbased learning platform for natural language processing</article-title>
          .
          <source>In Proceedings of ACL-IJCNLP 2015 System Demonstrations</source>
          , pages
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Simone</given-names>
            <surname>Filice</surname>
          </string-name>
          , Giuseppe Castellucci, Giovanni Da San Martino, Alessandro Moschitti, Danilo Croce, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Kelp: a kernel-based learning platform</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>18</volume>
          (
          <issue>191</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>CoRR, abs/1301</source>
          .3781.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Efficient convolution kernels for dependency and constituent syntactic trees</article-title>
          .
          <source>In ECML</source>
          , Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Nicole</given-names>
            <surname>Novielli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the evalita 2018 italian speech act labeling (ilisten) task</article-title>
          . In Tommaso Caselli, Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of the 6th evaluation campaign of Natural Language Processing and Speech tools for Italian (EVALITA'18)</source>
          , Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Robert</surname>
          </string-name>
          <article-title>Mu¨ller, Sebastian Mika, Gunnar Ra¨tsch, Koji Tsuda</article-title>
          , and Bernhard Scho¨lkopf.
          <year>2001</year>
          .
          <article-title>An introduction to kernel-based learning algorithms</article-title>
          .
          <source>IEEE Transactions on Neural Networks</source>
          ,
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <fpage>181</fpage>
          -
          <lpage>201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The Word-Space Model</article-title>
          .
          <source>Ph.D. thesis</source>
          , Stockholm University.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>John</surname>
            Shawe-Taylor and
            <given-names>Nello</given-names>
          </string-name>
          <string-name>
            <surname>Cristianini</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Kernel Methods for Pattern Analysis</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>David R.</given-names>
            <surname>Traum</surname>
          </string-name>
          ,
          <year>1999</year>
          .
          <article-title>Speech Acts for Dialogue Agents</article-title>
          , pages
          <fpage>169</fpage>
          -
          <lpage>201</lpage>
          . Springer Netherlands, Dordrecht.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Vladimir N.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Statistical Learning Theory</article-title>
          . Wiley-Interscience.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>S.V.N.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexander J.</given-names>
            <surname>Smola</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Fast kernels on strings and trees</article-title>
          .
          <source>In Proceedings of Neural Information Processing Systems</source>
          , pages
          <fpage>569</fpage>
          -
          <lpage>576</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>