<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PESInet: Automatic Recognition of Italian Statements, Questions, and Exclamations With Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sonia Cenceschi</string-name>
          <email>sonia.cenceschi@supsi.ch</email>
          <email>sonia.cenceschi@supsi.ch Roberto Tedesco Politecnico di Milano P.za L. da Vinci 32, Milano, Italy roberto.tedesco@polimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Licia Sbattella</string-name>
          <email>licia.sbattella@polimi.it</email>
          <email>licia.sbattella@polimi.it Davide Losio Politecnico di Milano P.za L. da Vinci 32, Milano davide.losio@mail.polimi.it Mauro Luchetti Politecnico di Milano P.za L. da Vinci 32, Milano, Italy mauro.luchetti@mail.polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano</institution>
          ,
          <addr-line>P.za L. da Vinci 32, Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SUPSI University</institution>
          ,
          <addr-line>Via Pobiette 11, Manno</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>PESInet is an Automatic Prosody Recognition system aiming at classifying Information Units as Statement, Question or Exclamation. PESInet adopts a modular architecture, with a master NN evaluating the results of two independent BLSTM NNs that work on audio and its transcription. PESInet has been trained with our own three-class, balanced corpus composed of about 1.5 million text phrases and 60 000 utterances of recited and spontaneous speech. PESInet reached an accuracy of 80% on three classes, and 91% on two classes (Question vs Non-question). Finally PESInet, compared against human listeners on a two-class test based on a different corpus, reached a better Accuracy (89% for PESInet, against 80% for human listeners).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The goal of PESInet was to investigate whether
clues derived from text could improve the
recognition of simple prosodic forms in Information Units
1http://www.polisocial.polimi.it
2https://www.esagramma.net
(IUs). In particular, we focused on Statement,
Question, and Exclamation which are
proposition’s structures and are independent of the
pragmatic function of the corresponding IU: each one
can assume a large set of illocutionary acts, as
explained into the Language into Act theory (L-AcT)
described in Cresti (2014). An IU is composed
of a textual realisation (i.e., a written phrase) and
an acoustic realisation (an audio recording of a
speaker uttering such a phrase), and conveys a
specific informative intention
        <xref ref-type="bibr" rid="ref1 ref8">(Austin, 1975; Cresti,
2000)</xref>
        . We designed a modular model based on
Neural Networks (NNs), able to highlight how
much audio and text affected recognition accuracy.
Moreover, to validate our results, we compared
our NN model against human listeners, on a set of
IUs that did not overlap with the corpus we used
to train the model.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        The majority of studies on prosody regards the
automatic recognition or detection of single prosodic
clues
        <xref ref-type="bibr" rid="ref11 ref15 ref16 ref17 ref20 ref22">(Ren et al., 2004; Jeon and Liu, 2009;
Tamburini and Wagner, 2007; Taylor, 1993)</xref>
        . Others,
deal with the detection of phrase boundaries or
prosodic phrases
        <xref ref-type="bibr" rid="ref13 ref17 ref23">(Liu et al., 2006; Wightman and
Ostendorf, 1991; Rosenberg, 2009)</xref>
        . Just a few
works, however, focus on modality detection. In
the following we briefly introduce some of them.
Question detection is investigated in Tang et al.
(2016) using Recurrent Neural Networks (RNN),
in the Mandarin language. Authors propose
sev
      </p>
      <p>Copyright 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
eral RNN and Bidirectional RNN (BRNN)
models, trained on a simulated call-centre recordings
consisting of just 2850 Question and 3142
Nonquestion IUs. The best result is an F1 score of
85.5%.</p>
      <p>The work described in Yuan and Jurafsky
(2005) focuses on Question and Statement
detection, from text and audio, for Chinese; authors
investigate the influence of text in prosody
comprehension, on a telephone corpus (with
transcriptions). Their classifier achieves an error rate of
14.9% with respect to a 50% chance-level rate.
Quang et al. (2007) use decision trees to
automatically detect Questions in a small elicited French
and Vietnamese corpora, leveraging both
acoustic and lexical features (unigrams, bigrams, and
presence of so-called “interrogative terms”). The
best result is an F1 of 80% for the Vietnamese
language.</p>
      <p>
        Finally, the work described in
        <xref ref-type="bibr" rid="ref12">Li et al.
(2016)</xref>
        combines Convolutional NNs (CNNs) and
Bidirectional Long Short-Term Memory NNs
(BLSTM) to extract textual and acoustic
features for recognising stances (Affirmative,
Neutral, Negative opinions) in the Mandarin language.
It exploits a small, manually-tagged corpus of four
debate videos (1254 IUs). Combining both
audio and text this system reaches an Accuracy of
90.3%.
      </p>
      <p>None of the works mentioned above is perfectly
comparable with ours and, on the other hand, all
of them are based on ah-hoc corpora (as we did).
This makes impossible to compare the results we
obtained against other approaches. We, however,
validated our results comparing our model against
human listeners.
4</p>
    </sec>
    <sec id="sec-3">
      <title>The corpus</title>
      <p>
        Our own corpus is composed of eBooks, EPUB3
audio-books (an EPUB3 audio-book contains both
text and audio recording, time-aligned at the level
of sentence), and the LIT/DIA-LIT corpus
        <xref ref-type="bibr" rid="ref2 ref5">(Biffi,
1976; Buroni, 2009)</xref>
        , which contains audio
recordings of Italian TV shows, with transcriptions.
      </p>
      <p>From eBooks, the textual part of EPUB3
audiobooks, and transcriptions of LIT/DIA-LIT we
extracted about 1.5 million sentences, balanced on
the three target classes: Statement, Question, and
Exclamation.</p>
      <p>From LIT/DIA-LIT audio recordings and the
audio part of EPUB3 audio-books, we collected
about 60 000 utterances (again, balanced on
the three target classes). Both sentences and
utterances were tagged with the correct class,
leveraging the punctuation marks we found in
text/transcriptions. Of course, we removed such
punctuation marks from the textual part of the
corpus. Moreover, we discarded all the sentences
containing a sub-phrase or other complex
syntactic structures. In doing so we aimed at retaining
plain simple examples of statements, questions,
and exclamations.</p>
      <p>We are aware that leveraging punctuation marks
for tagging sentences can lead to confounds, as
exclamation marks is also used for Vocatives and
Orders, while the full stop is also used for Orders.
Anyway, it was simply not possible to manually
review the text collection and manually solve the
problem. Thus, we assume our corpus is affected
by a small amount of noise (in other words, we
assume Exclamations and Statements are way more
frequent than Vocatives and Orders).</p>
      <p>Notice that the question marks might be
used for different question typologies
(rhetorical, information-seeking, confirmation-seeking,
biassed), and that question could be further
partitioned into open questions, polar questions, etc.
Thus, the question mark is used to tag sentences
with wildly divergent phonetic forms. This is not,
however, a blocking issue: it only makes harder
for the classifier to learn the input/output
correlation. In particular, this is one of the reasons that
lead us to the idea of leveraging text to improve
the classification of IUs.</p>
      <p>Summing up, we built three corpora:</p>
      <p>ACorpus: audio corpus composed of about
60 000 .wav labelled samples.</p>
      <p>TCorpus: textual corpus composed of about
1.5 million .txt labelled samples.</p>
      <p>MCorpus: mixed corpus composed of all the
ACorpus files, with their transcriptions (from
the TCorpus); about 60 000 labelled samples.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Features extraction</title>
      <p>From acoustic and textual samples we derived a
set of features that our NNs leveraged for training
and recognition.
5.1</p>
      <sec id="sec-4-1">
        <title>Acoustic features</title>
        <p>With a sample rate of 44.1 kHz, we adopted a
window of 2048 samples with a hop-size of 1024
samples (i.e., every 23 ms a new vector of acoustic
features is produced). Notice that our window is
larger than the one usually adopted by ASRs; in
fact, we are not interested in phone recognition
and, on the other hand, prosody phenomena
appear in larger temporal scale than the one
involving individual phones.</p>
        <p>
          We tried several window sizes, and several
acoustic features; in particular we experimented
with different combinations of Cepstrum
coefficients. At the end, we come up with the following
129 acoustic features, normalised (to minimise
dependency on speakers and recording settings) and
calculated by means of Praat
          <xref ref-type="bibr" rid="ref3">(Boersma and others,
2001)</xref>
          , as they provided the best results:
pitch value, with its delta and delta-delta
energy, with its delta and delta-delta
the first 40 Cepstrum coefficients, with their
deltas and delta-deltas
energy of such 40 Cepstrum coefficients (as
MFCC defines), with its delta and delta-delta
Notice that we did not adopt a true “deep”
architecture, as features were not “discovered” by the
network. The field of audio analysis already
provides a huge set of well-known, informative
features; thus, in our opinion, there is no point in let
the network approximating them. Moreover,
precalculated features permit to simplify the network.
Summing up, each utterance was transformed into
an array that contains a column of 129 real
numbers every 23ms.
5.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Textual features</title>
        <p>
          To feed the model with textual samples we used
the usual word embedding technique, which
represents the vocabulary in a continuous vector space
of 300 dimensions
          <xref ref-type="bibr" rid="ref18">(Sahlgren, 2008)</xref>
          . In
particular we adopted Italian Word Embeddings, a
pretrained model of 700 000 words based on GloVe
          <xref ref-type="bibr" rid="ref14">(Pennington et al., 2014)</xref>
          .
        </p>
        <p>Summing up, each sentence was transformed
into an array that contains a column of 300 real
numbers for each token. Notice that punctuation
marks were discarded and no lemmalisation was
applied.</p>
        <p>Available at:
wordembeddings/
http://hlt.isti.cnr.it/</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Architecture</title>
      <p>PESInet is composed of three different NNs:
Acoustic and textual features defined in Section 5
generated low-level pieces of information,
looking at very local phenomena. For considering
higher-level phenomena, both the Audio-based
and the Text-based NNs relied on the same
architecture, leveraging an initial multi-layer
convolutional block.</p>
      <p>A convolutional layer is composed of several
kernels with a predefined width, which “scan” the
input array. Each kernel, after the training phase,
specialises in finding certain patterns in the
input sequence. The network learns “high level”
features (i.e., common prosody contours, for the
Audio-based NN, or particular word sequences for
the Text-based NN) from our low-level feature set.</p>
      <p>
        Features related to prosody unfold along
different time extents
        <xref ref-type="bibr" rid="ref9">(Cutugno et al., 2005)</xref>
        : we
found dependencies both in short and long time
periods. So the idea was to use different kernel
widths, in order to allow the network to consider
different pattern lengths. The hint to adopt this
technique come from various papers
        <xref ref-type="bibr" rid="ref10 ref11 ref17 ref19 ref4">(Sbattella et
al., 2014; Gussenhoven, 2008; Bu¨ring and others,
2009)</xref>
        , which thoroughly analysed the idea of
simultaneously analysing the input at different
temporal granularities with the use of differently-sized
kernels.
      </p>
      <p>In particular, our convolutional block is
composed of three layers, which “scan” at three
different temporal granularity levels. In general, if
s is the stride adopted for kernels at any
temporal granularity level and di is the kernel height
at the i-th temporal granularity level, the kernel
height at the (i + 1)-th temporal granularity level
is di+1 = di + s; see Figure 1, for a simplified
example with two levels. Stride is chosen so that,
after each shift of the filter, the kernel will include
a small subset of the previously analysed input.</p>
      <p>Finally, padding is applied to the input
sequence, so that the shorter kernels (and, by
construction, all the other, longer kernels) fit the
sequence length.</p>
      <p>Being the kernels of different heights, they will
cause the outputs to have different dimension as
well, relatively to the layers they come from.
These dimensions are adjusted in the following
layer of the network. Figure 2 shows a simplified
schema with two differently-sized kernel groups.
6.2</p>
      <sec id="sec-5-1">
        <title>Audio-based and Text-based NNs</title>
        <p>Both the Audio-based and the Text-based NNs
relied on a multi-layer network. The general
architecture is composed of three BLSTM layers on top
of the convolutional block. We connected the first
convolutional layer to the first BLSTM layer; then,
the second convolutional layer is connected,
together with the output from the first BLSTM, to
the second BLSTM layer; finally, the third
convolutional layer is connected, together with the
output of the second BLSTM layer, to the third
BLSTM layer. Figure 3 shows the way in which
the convolutional block is used.</p>
        <p>The Softmax layer shown in the Figure 3 is used
during the training phase and then removed, as the
Text-based and Audio-based NNs are combined
together with the Master NN.
6.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>Master NN and PESInet</title>
        <p>The Master NN is composed of a fully-connected
layer, and a Softmax layer. PESInet, the
resulting network, is shown in Figure 4. Notice that
PESInet is supposed to works on utterances, while
the text is generated by means of an ASR; in fact,
this is the setting we expect to be adopted during
actual usage of PESInet. Our corpus, conversely,
was based on human-generated text; we are aware
that in doing so we did not consider the errors due
to the ASR and, as a consequence, overestimated
the figures obtained during the training/validation
procedure. The rationale was highlighting the
contribution of text-related features to the recognition
of prosodic forms, and thus we decided to avoid
the “noise” introduced by ASR-related errors.</p>
        <p>As a final remark on the ASR, notice that it is
supposed to not add any punctuation mark to the
transcription it generates.
7</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Training</title>
      <p>The architecture was implemented, trained, and
tested using the TensorFlow library along with
Python 3.6. The code itself was run on a machine
equipped with 32GB of RAM, a Xeon Intel
processor and a Nvidia Titan X (Pascal) GPU. During
training, we adopted the early stopping (using
Accuracy as reference index); moreover, to improve
the learning effectiveness, we used the variational
drop-out on recurrent layers. We started
training, independently, the Audio-based and the
Textbased NNs, on 80% of their respective corpora:
ACorpus and TCorpus. Then, once removed the
final Softmax layer from them, these NNs where
attached to the Master NN, and a further training
–involving 80% of the MCorpus– was performed
on PESInet. In particular, we investigated three
approaches:
1. Allowing PESInet to train only the Master</p>
      <p>NN weights (all the others remain fixed).
2. Allowing PESInet to change all its internal
weights (also those already trained).
3. Training PESInet from scratch, skipping
training of Audio-based and Text-based NNs.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Evaluation</title>
      <p>Validation was performed using 20% of the
corpus. We experimented with several feature
combinations, hyperparameter values, and network
structures, before reaching the final models.</p>
      <p>The Audio-based and Text-based NNs gave the
following Accuracies: 0.68 and 0.79. It’s
interesting that the Text-based NN gave a better
Accuracy than the Audio-based NN. This was
surprising, as, after all, prosody is an acoustic
phenomenon. Nevertheless, data seem to show that
the words composing the utterance are indeed a
good predictor of prosody. Moreover, considering
that ACorpus was much smaller than TCorpus, the
surprisingly low results of Audio-based NN can be
explained.</p>
      <p>Table 1 and Table 2 show the confusion
matrices for the two NNs. It’s interesting to notice that
Audio-based NN predicted Statements much
better than the other two classes, while Text-based
NN was also very good in recognising Questions.</p>
      <p>About PESInet, Table 3 shows that the approach
2 obtained, as expected, the best results. As the
confusion matrix of Table 4 shows, audio and text
cooperated to improve recognition of all the three
classes.</p>
      <p>As a further experiment, we trained and tested
PESInet on two classes: Question vs
Nonquestion, adapting the same PESInet
architecture to handle 2 classes. The corpus tags
Predicted</p>
      <p>Excl.
205
1242
215
racy for our NN (or surprisingly bad Accuracy
for human listeners) could be caused by
decontextualisation: in the experiment each IU was
given in isolation, without any dialogue context;
probably, listeners were more affected by that
lacking of context than our NN. Anyway, this is
just a hypothesis that should be investigated and
deepened with further experiments, as the
comparison could be tainted by a large number of other
confounds, such as the non ecological nature of
the task and the stratification of the repertoire of
Italian speakers.
fExclamation, Statementg were rewritten as
Nonquestion, and we randomly extracted a number
of Non-Question samples equals to the Question
samples. Then, we used 90% of such dataset for
training and 10% for testing. Accuracy reached
91% (Table 5).
8.1</p>
      <sec id="sec-7-1">
        <title>PESInet against human listeners</title>
        <p>
          Finally, to validate the results we obtained, we
conducted a perceptive experiments with 302
Italian speakers
          <xref ref-type="bibr" rid="ref6 ref6 ref7 ref7">(Cenceschi et al., 2018b; Cenceschi
et al., 2018a)</xref>
          . The aim of the experiment was to
understand the role of acoustic clues and textual
clues in the perception of various prosodic forms.
        </p>
        <p>The experiment was divided into several tests;
each test was about a specific prosodic form: users
were asked to listen a set of IUs and select which
of them carried the expected prosodic form. In that
experiment we used an ad-hoc audio/textual
corpus called SI-CALLIOPE, where 14 professional
actors spoke a set of 139 sentences, for a total
of 1946 IUs. Notice that SI-CALLIOPE did not
share anything, in terms of sentences and
speakers, with corpora we used to train PESInet.</p>
        <p>In particular, for the Question/Non-question
test, each user listened to a set of audios randomly
extracted from 714 question IUs and 1232
nonquestion IUs. The average accuracy was 80% (std.
dev.: 7.24%).</p>
        <p>Running the two-class version of PESInet on
the same test, we got an Accuracy of 89%.</p>
        <p>We argue that this surprisingly good
AccuPESInet got an Accuracy of 80% on three classes
and and 91% on two classes. Moreover, PESInet
reached very good results when compared to
human listeners on a totally different corpus.
Although this human/NN comparison should be
taken with a grain of salt, we believe that it is a
hint that the network works well and the results are
truly promising. As a future work, more
recordings should be added to ACorpus and MCorpus to
improve the performance of the Audio-based NN
and, as a consequence, of the whole PESInet.</p>
        <p>Currently, we are working for cleaning the code
and streamlining the training procedure, as we
plan to release the code.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          John Langshaw Austin.
          <year>1975</year>
          .
          <article-title>How to do things with words</article-title>
          , volume
          <volume>88</volume>
          . Oxford university press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Biffi</surname>
          </string-name>
          .
          <year>1976</year>
          .
          <article-title>Il lit-lessico italiano televisivo: l'italiano televisivo in rete</article-title>
          . L'italiano televisivo:
          <fpage>1976</fpage>
          -
          <lpage>2006</lpage>
          .
          <article-title>Atti del convegno-</article-title>
          <string-name>
            <surname>Milano</surname>
          </string-name>
          ,
          <fpage>15</fpage>
          -
          <lpage>16</lpage>
          giugno
          <year>2009</year>
          , pages
          <fpage>35</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Paul</given-names>
            <surname>Boersma</surname>
          </string-name>
          et al.
          <year>2001</year>
          .
          <article-title>Praat, a system for doing phonetics by computer</article-title>
          . Glot international,
          <volume>5</volume>
          (
          <issue>9</issue>
          /10):
          <fpage>341</fpage>
          -
          <lpage>345</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          Daniel Bu¨ring et al.
          <year>2009</year>
          .
          <article-title>Towards a typology of focus realization</article-title>
          .
          <source>Information structure</source>
          , pages
          <fpage>177</fpage>
          -
          <lpage>205</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Edoardo</given-names>
            <surname>Buroni</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>La voce del telegiornale. aspetti prosodici del parlato telegiornalistico italiano in chiave diacronica</article-title>
          .
          <source>l'italiano televisivo 1976- 2006. Atti del Convegno</source>
          , Milano, pages
          <fpage>15</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Sonia</given-names>
            <surname>Cenceschi</surname>
          </string-name>
          , Licia Sbattella, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Tedesco</surname>
          </string-name>
          . 2018a.
          <article-title>Towards automatic recognition of prosody</article-title>
          .
          <source>In Proc. 9th International Conference on Speech Prosody</source>
          <year>2018</year>
          , pages
          <fpage>319</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Sonia</given-names>
            <surname>Cenceschi</surname>
          </string-name>
          , Licia Sbattella, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Tedesco</surname>
          </string-name>
          . 2018b.
          <article-title>Verso il riconoscimento automatico della prosodia</article-title>
          .
          <source>STUDI AISV</source>
          , pages
          <fpage>433</fpage>
          -
          <lpage>440</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Emanuela</given-names>
            <surname>Cresti</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Corpus di italiano parlato: Introduzione</article-title>
          , volume
          <volume>1</volume>
          . Accademia della Crusca.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Cutugno</surname>
          </string-name>
          , G. Coro, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Petrillo</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Multigranular scale speech recognizers: Technological and cognitive view</article-title>
          . In Springer, editor,
          <source>Congress of the Italian Association for Artificial Intelligence</source>
          ,
          <fpage>227</fpage>
          -
          <lpage>330</lpage>
          , Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Gussenhoven</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Types of focus in english</article-title>
          .
          <source>In Topic and focus</source>
          , pages
          <fpage>83</fpage>
          -
          <lpage>100</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>Je Hun Jeon and Yang Liu</source>
          .
          <year>2009</year>
          .
          <article-title>Automatic prosodic events detection using syllable-based acoustic and syntactic features</article-title>
          .
          <source>In 2009 IEEE International Conference on Acoustics, Speech and Signal Processing</source>
          , pages
          <fpage>4565</fpage>
          -
          <lpage>4568</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Linchuan</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Zhiyong Wu</given-names>
            , Mingxing Xu,
            <surname>Helen M Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Lianhong</given-names>
            <surname>Cai</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Combining cnn and blstm to extract textual and acoustic features for recognizing stances in mandarin ideological debate competition</article-title>
          .
          <source>In INTERSPEECH</source>
          , pages
          <fpage>1392</fpage>
          -
          <lpage>1396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Yang</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, and
          <string-name>
            <given-names>Mary</given-names>
            <surname>Harper</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Enriching speech recognition with automatic detection of sentence boundaries and disfluencies</article-title>
          .
          <source>IEEE Transactions on audio, speech, and language processing</source>
          ,
          <volume>14</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1526</fpage>
          -
          <lpage>1540</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Vu˜ Minh Quang</surname>
            , Laurent Besacier, and
            <given-names>Eric</given-names>
          </string-name>
          <string-name>
            <surname>Castelli</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Automatic question detection: prosodiclexical features and crosslingual experiments</article-title>
          . In Eighth Annual Conference of the International Speech Communication Association.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Yuexi</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sung-Suk</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Mark Hasegawa-Johnson, and
          <string-name>
            <given-names>Jennifer</given-names>
            <surname>Cole</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Speaker-independent automatic detection of pitch accent</article-title>
          .
          <source>In Speech Prosody</source>
          <year>2004</year>
          , International Conference.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Rosenberg</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Automatic detection and classification of prosodic events</article-title>
          . Columbia University.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>The distributional hypothesis</article-title>
          .
          <source>Italian Journal of Disability Studies</source>
          ,
          <volume>20</volume>
          :
          <fpage>33</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Licia</given-names>
            <surname>Sbattella</surname>
          </string-name>
          , Roberto Tedesco, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Trivilini</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Forensic examinations: Computational analysis and information extraction</article-title>
          .
          <source>In International Conference on Forensic ScienceCriminalistics Research (FSCR)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Tamburini</surname>
          </string-name>
          and
          <string-name>
            <given-names>Petra</given-names>
            <surname>Wagner</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>On automatic prominence detection for german</article-title>
          . In Eighth Annual Conference of the International Speech Communication Association.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Yaodong</given-names>
            <surname>Tang</surname>
          </string-name>
          , Yuchen Huang, Zhiyong Wu, Helen Meng, Mingxing Xu,
          <string-name>
            <given-names>and Lianhong</given-names>
            <surname>Cai</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Question detection from acoustic features using recurrent neural network with gated recurrent unit</article-title>
          .
          <source>In Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2016</year>
          IEEE International Conference on, pages
          <fpage>6125</fpage>
          -
          <lpage>6129</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Paul A</given-names>
            <surname>Taylor</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Automatic recognition of intonation from f0 contours using the rise/fall/connection model</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>CW</given-names>
            <surname>Wightman and Mari Ostendorf</surname>
          </string-name>
          .
          <year>1991</year>
          .
          <article-title>Automatic recognition of prosodic phrases</article-title>
          .
          <source>In Acoustics, Speech, and Signal Processing</source>
          ,
          <year>1991</year>
          . ICASSP-
          <volume>91</volume>
          ., 1991 International Conference on, pages
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Jiahong</given-names>
            <surname>Yuan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Detection of questions in chinese conversational speech</article-title>
          .
          <source>In Automatic Speech Recognition and Understanding</source>
          ,
          <source>2005 IEEE Workshop on</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>52</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>