<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CorAIt - A Non-native Speech Database for Italian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Roberta Combei</string-name>
          <email>roberta.combei@fileli.unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FiLeLi - University of Pisa, on leave at FAU Erlangen-Nürnberg</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. CorAIt is a non-native speech database for Italian, which is freely accessible online for academic research purposes. It was especially designed to meet the requirements of a larger research project focused on foreign accented Italian speech. The corpus is aimed at providing a uniform collection of speech samples uttered by non-native speakers of Italian. To date, 105 non-native speakers - whose mother tongues are either French, Romanian, Spanish, English, German, or Russian - have been recorded. The corpus includes also a control group made up of 16 Italian speakers. There are almost 8 hours of audio material, both read speech (first and second reading), and spontaneous speech. This paper emphasizes the necessity for this type of database, it describes the steps involved in its construction, and it presents the features of CorAIt.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. CorAIt è un corpus audio di
l’italiano L2 liberamente consultabile
online per scopi di ricerca scientifica. Il
corpus è parte integrante di un progetto di
ricerca che affronta l’accento straniero
nella lingua italiana da una prospettiva
più ampia. E’ stato ideato e costruito con
lo scopo di fornire una raccolta uniforme
di materiale audio prodotto da parlanti di
italiano L2. Ad oggi sono stati registrati
105 parlanti stranieri di madrelingua:
francese, romena, spagnola, inglese,
tedesca, e russa. In aggiunta, il corpus è
dotato di un gruppo di controllo composto
da 16 parlanti italiani. Sono disponibili
circa 8 ore di registrazioni, sia di parlato
letto (prima e seconda lettura) che di
parlato spontaneo. L’articolo evidenzia la
necessità di costruire questo tipo di
database, e descrive la progettazione e le
caratteristiche di CorAIt.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        It has become clear that accurately designed
speech corpora are of essential importance for the
development of efficient speech technologies.
Investigating how native and foreign-accented
speech differ is a necessary step in non-native
speech recognition
        <xref ref-type="bibr" rid="ref14">(Tomokiyo, 2001)</xref>
        .
      </p>
      <p>Currently, the number of non-native speech
databases seems almost insignificant if compared
to corpora of native speech.</p>
      <p>
        Moreover, until recently, the majority of the
research has focused on English. Therefore, some
of the largest non-native speech databases are
available for this language: TED
        <xref ref-type="bibr" rid="ref15">(Lamel et al.,
1994)</xref>
        , Duke-Arslan
        <xref ref-type="bibr" rid="ref1">(Arslan &amp; Hansen, 1997)</xref>
        ,
ISLE
        <xref ref-type="bibr" rid="ref19">(Menzel et al., 2000)</xref>
        , IBM-Fisher
        <xref ref-type="bibr" rid="ref10">(Fisher et
al., 2003)</xref>
        , ATR-Gruhn
        <xref ref-type="bibr" rid="ref11 ref6 ref7">(Gruhn et al., 2004)</xref>
        , CSLU
        <xref ref-type="bibr" rid="ref16">(Lander, 2007)</xref>
        , NATO M-ATC
        <xref ref-type="bibr" rid="ref25 ref26">(Pigeon et al.,
2007)</xref>
        , and Speech Accent Archive
        <xref ref-type="bibr" rid="ref33">(Weinberger,
2015)</xref>
        . Large speech corpora of foreign-accented
English are owned by Speechocean, and they are
specifically built for commercial purposes,
especially for training and testing speech
recognizers, but some of them are also available
for academic research on the KilingLine Data
Center platform
        <xref ref-type="bibr" rid="ref30">(Speechocean, 2017)</xref>
        .
      </p>
      <p>
        Only in the last few years has there been an
interest in other languages. Without claiming to be
exhaustive, some of the largest non-native speech
databases for languages other than English will be
mentioned: BAS Strange I+II
        <xref ref-type="bibr" rid="ref18">(University of
Munich, 1998)</xref>
        for German, WP Russian
        <xref ref-type="bibr" rid="ref17">(La
Rocca &amp; Tomei, 2003)</xref>
        for Russian, Tokyo-Kikuko
        <xref ref-type="bibr" rid="ref22">(Nishina, 2004)</xref>
        for Japanese, TC-STAR
        <xref ref-type="bibr" rid="ref13 ref31 ref35">(van den
Heuvel et al., 2006)</xref>
        and WP Spanish
        <xref ref-type="bibr" rid="ref21">(Morgan,
2006)</xref>
        for Spanish, SINOD
        <xref ref-type="bibr" rid="ref13 ref31 ref35">(Žgank et al., 2006)</xref>
        for
Slovenian, and iCALL
        <xref ref-type="bibr" rid="ref5">(Chen et al., 2015)</xref>
        for
Chinese.
      </p>
      <p>
        However, as a result of that fact that many
nonnative speech databases are built for commercial
purposes within private research centres, it is
actually quite difficult to map all the resources of
this type ever built
        <xref ref-type="bibr" rid="ref12">(Cf. Gruhn et al., 2011, for an
overview of the non-native speech databases
available at the date their study was published)</xref>
        .
      </p>
      <p>This paper presents CorAIt, a non-native
speech database for Italian. The database is part of
a Ph.D. project which intends to study foreign
accented Italian speech both from a computational
perspective (automatic identification and
classification of non-native accent) and a
perceptual perspective (interpretation of
quantitative and qualitative judgments delivered
by expert and naïve native Italian speakers with
respect to non-native pronunciations). The design
and the development of this corpus were
determined by several factors, which are outlined
below.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Motivations</title>
      <p>Currently, the automatic speech recognition
systems for Italian which are integrated into
generally available virtual assistant software (e.g.
Google Now, Google Assistant, Siri, Cortana,
etc.) perform quite well on native speech.
However, despite recent advances in this field,
non-native accents still represent a challenge. This
may be due to fact that there is significantly less
training data available for automatic speech
recognition systems on non-native pronunciations.
Considering that Italy is a multicultural country,
with over 5 million foreign citizens, representing
8.3% of the entire population residing on its
territory1, it would be desirable to provide services
to users who speak Italian with non-native
accents.</p>
      <p>Apart from acting like training sets for
automatic speech recognition systems or for
textto-speech systems, non-native speech databases
might be beneficial in the fields of computer
assisted language learning (CALL) and
mobileassisted language learning (MALL), as well as for
linguistic profiling tasks. Glottologists and
scholars working on Italian as a foreign language
might also benefit from the presence of these
resources.</p>
      <p>
        At the date this research began, there was only
one audio corpus for foreign-accented Italian
speech, freely available for online consultation,
1 The data were provided by the National demographic
balance (year 2016) produced by the Italian National Institute
for Statistics (ISTAT). The full report is available at:
http://www.istat.it/it/files/2017/06/bilanciodemografico2016_13giugno2017.pdf
namely DILS - Dialoghi in Italiano Lingua
Straniera
        <xref ref-type="bibr" rid="ref29">(Savy et al., 2012)</xref>
        , consisting of
semispontaneous audio material obtained by means of
the task-oriented dialogue elicitation technique.
DILS contains 9 large audio samples (for a total
duration of 100 minutes) uttered by 18 speakers:
12 Dutch females, 3 Spanish females and 3
Spanish males.
      </p>
      <p>
        It is worthwhile to mention that there are
several other learner corpora for Italian: VALICO
- Varietà Apprendimento Lingua Italiana Corpus
Online
        <xref ref-type="bibr" rid="ref3">(Barbera &amp; Marello, 2004)</xref>
        , which is a
collection of non-native written Italian; LIPS
Lessico dell’italiano parlato da stranieri
        <xref ref-type="bibr" rid="ref13 ref31 ref32 ref35">(Vedovelli et al., 2006)</xref>
        ; and Corpus Parlato di
Italiano L2
        <xref ref-type="bibr" rid="ref13 ref31 ref35">(Spina et al., 2006)</xref>
        . The last two
corpora consist of transcriptions of audio samples
produced by non-native speakers.
      </p>
      <p>
        In addition to the above-mentioned corpora,
there exists a database of written and spoken
nonnative Italian, entitled ADIL2 - Archivio Digitale
di Italiano L2
        <xref ref-type="bibr" rid="ref23">(Palermo, 2009)</xref>
        , which is
purchasable in the form of a DVD. However,
despite the sophistications of its search tool, the
accurate transcription, as well as the admirable
amount of data collected, ADIL2 presents a series
of issues that cannot be ignored, such as:
imbalance with respect to the speakers’ mother
tongues (i.e. some languages are underrepresented
while others are overrepresented) and elicitation
technique used for some samples (i.e. interviews
repeated various times over variable time-frames
to the same subjects). These aspects render ADIL2
unsuitable for the type of research to be taken on.
Therefore, it became necessary to collect a
database of non-native Italian speech.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Data Collection</title>
      <p>The corpus was designed, collected and developed
from January 2016 through July 2017, and it was
aimed at providing a uniform collection of audio
material produced by adult non-native speakers of
Italian residing in Bologna2.</p>
      <p>Initially, the intent was to collect data for 11
different mother tongues (L1s): Maghrebi Arabic,
Urdu, Mandarin Chinese, Albanian, Russian,
English, German, French, Romanian, Spanish, and
Italian (as a control group). The first 10 groups
correspond to the L1s spoken by some of the
major foreign populations residing in Italy.
However, recruiting speakers for all these groups
2 To simplify the data collection process, the author chose to
recruit people that studied or worked in Bologna (her city of
residence).
proved to be a challenging task. This may be
result of the fact that participation was entirely
voluntary and no material reward was provided to
informants. Since it was not possible to recruit
enough speakers of Maghrebi Arabic, Urdu,
Mandarin Chinese and Albanian, these four
groups were abandoned.
3.1</p>
      <sec id="sec-4-1">
        <title>Speakers’ recruitment</title>
        <p>Specific criteria of quality, quantity and diversity
were observed, as much as possible, for each L1
when the participants were selected (Cf. section
4.1).</p>
        <p>All speakers were recruited locally in Bologna.
Most informants were enrolled as regular or
exchange students in B.A., M.A. and Ph.D.
programmes at the University of Bologna and
they were contacted on their personal e-mail
address. The e-mail message contained a
description of the research project and informed
potential participants about the tasks they would
have performed. Nearly one fourth of them replied
positively to the call.
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Experimental protocol</title>
        <p>All informants were aware that they were
recorded. They gave informed consent in writing
to the use of their speech samples and their
sociolinguistic data for research purposes.</p>
        <p>
          In order to guarantee uniformity, the same
experimental protocol was employed for all
subjects. Before each recording session, speakers
were asked to fill in a detailed form regarding
their sociocultural and sociolinguistic background.
The digital recordings were performed with a
Samson METEOR MIC cardioid
pickup microphone (condenser diaphragms: 25
mm) on the Praat software
          <xref ref-type="bibr" rid="ref4">(Boersma &amp; Weenink,
2017)</xref>
          . The sampling parameters were the
following: mono channel, 16-bit, 44,100 Hz,
linearly encoded WAV.
        </p>
        <p>Each recording session lasted around 60
minutes. The sessions were individual-based, and
they were guided and monitored by the author.
3.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Speech modalities</title>
        <p>
          The speakers were asked to perform two tasks:
reading an article excerpt published on the Italian
newspaper Corriere della Sera3; and describing
spontaneously how they spent their last holidays.
3 The newspaper article is available online at:
http://cinquantamila.corriere.it/storyTellerArticolo.php?storyI
d=0000002228555. The excerpt was included in this frame:
“Don Geretti è un grande affabulatore […] Pietro terminò il
suo cammino terreno e quello, tormentatissimo, verso la
fede.”
That specific reading fragment was chosen
because it presented various levels of complexity
and it contained all Italian phonemes. The reading
task was necessary for triggering difficulties that
could emerge as a result of conflicting
orthographic conventions between the speakers’
mother tongues and Italian
          <xref ref-type="bibr" rid="ref34">(Wottawa &amp;
AddaDecker, 2016)</xref>
          . Moreover, it could allow speakers
comparisons and analyses on the same type of
material.
        </p>
        <p>All participants had two reading attempts and
they were asked to read and speak as naturally as
they could.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Description of CorAIt</title>
      <p>CorAIt is a non-native speech database for Italian,
which has become fully and freely available
online for academic research consultation4. It
contains 2,244 audio samples produced by 105
non-native speakers of Italian. It also includes 300
audio samples obtained from 16 native Italian
speakers. In total there are almost 8 hours of
speech, consisting in roughly 72,000 words.
4.1</p>
      <sec id="sec-5-1">
        <title>Speakers’ statistics</title>
        <p>Originally, it was planned to recruit at least 15
speakers for each L1. This threshold was reached
for French, and it was exceeded for the other six
groups.</p>
        <p>Regarding the age distribution of the
informants, the range is 19-40 years, but most
speakers are older than 20 and younger than 30
years (Cf. Table 1).</p>
        <sec id="sec-5-1-1">
          <title>Mother</title>
          <p>tongue
Russian</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>English</title>
        </sec>
        <sec id="sec-5-1-3">
          <title>German</title>
        </sec>
        <sec id="sec-5-1-4">
          <title>French</title>
        </sec>
        <sec id="sec-5-1-5">
          <title>Romanian</title>
        </sec>
        <sec id="sec-5-1-6">
          <title>Spanish</title>
        </sec>
        <sec id="sec-5-1-7">
          <title>Italian</title>
          <p>Number of
speakers
17
16
17
15
20
20
16</p>
          <p>Age means</p>
          <p>(±S.D.)
27.06 (5.94)
23.75 (4.69)
23.53 (4.06)
23.33 (2.72)
25.40 (2.22)
23.65 (2.95)
29.00 (4.62)</p>
          <p>
            Despite the efforts to call up the same number
of male and female speakers, the corpus is not
perfectly gender-balanced. Apparently, it was
4 CorAIt database is accessible for consultation (prior to
registration) at: www.corait.it
easier to engage female participants. In fact, 80
females (66%) and 41 males (34%) were recorded
for this project. Nonetheless, according to
literature, gender has not been reported as a major
source of pronunciation issues, and for this reason
it was assumed that gender imbalance had minor
effects on this type of studies
            <xref ref-type="bibr" rid="ref12">(Gruhn et al., 2011)</xref>
            .
          </p>
          <p>For the corpus design, the age of Italian
language onset was taken into consideration, and
it was distributed as follows: childhood (23%),
adolescence (26%), and adulthood (51%). The
data are predictable, considering that most
informants learnt Italian naturalistically (64%)
soon after they moved to Italy. In fact, only 36%
of them claimed that they used mainly scholastic
methods for learning Italian, and that they had
already spoken the language at their arrival in
Italy.</p>
          <p>Since most informants were exchange students,
59% of them had spent 6-12 months in Italy at the
time they were recruited for this research. The
remaining part had lived in Italy for 12-24 months
(15%), or for more than 24 months (26%). Not
surprisingly, the great majority of speakers
claimed that they had been exposed only to the
Bolognese variety of Italian.</p>
          <p>Because it was almost impossible to predict the
speakers’ proficiency level5 in Italian before
meeting them, the balancedness is not guaranteed
for all accent groups (e.g. no Romanian speaker
had A2 waystage/elementary level in Italian). For
the sake of brevity, at the general level, this
variable is represented as follows in the database:
waystage/elementary level - A2 (12%),
threshold/intermediate level - B1 (28%),
vantage/upper-intermediate level - B2 (28%), and
advanced/proficiency levels - C1 and C2 (32%).
4.2</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>Speech samples</title>
        <p>An average of 4 minutes and 25 seconds of raw
audio material consisting of read and spontaneous
speech were recorded for each speaker. Some
speakers had to terminate the registration session
earlier than planned, so in those cases it was
possible to record only their first reading attempt.
Regardless of that, the spontaneous speech
collected (23%) is, however, inferior to the
reading speech material (77%). All raw samples
were segmented manually into utterances
corresponding to grammatical sentences for the
reading material, and to phonological sentences
5 All participants self-assessed their Italian level based on the
Common European Framework of Reference for Languages,
available at:
http://www.coe.int/t/dg4/linguistic/Source/Framework_EN.p
df
for the spontaneous speech. In most cases, the
material was not qualitatively altered, so
hesitation phenomena and disfluencies were
generally left as they were.
4.3</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Webapp architecture6</title>
      <p>One contribution of this project is that of making
the corpus available to the research community.
Following the model of similar tools, a website
that would host the database was created. Then,
the embedded webapp could extrapolate and
classify the audio files from the dataset, according
to specific criteria.</p>
      <p>For the creation of the webapp, the web
framework Django as well as several Python
libraries (MySQL-python, django-treebeard,
django-filer, html5lib, sorl, wsgi, polymorphic,
classy-tags, audiofield, appconf, etc.) were
employed. That allowed the use of a powerful
ORM system, equipped with a web interface for
storing multiple data types into our MySQL
database.</p>
      <p>Moreover, the Django web framework ORM
favoured the realization of data collection models:
a model is the only final data source containing
the fields and the essential behaviours of the
dataset and of the reference objects.</p>
      <p>Generally, each model is mapped to a single
database table and each attribute represents a
database field. The queries are performed by
means of ad-hoc APIs for each model.</p>
      <p>The project is hosted on a server with a CentOS
7 operating system. CortAIt is already configured
for various types of SQL and NoSQL databases
(PostgreSQL, MongoDB, Cassandra, etc.). It also
supports the execution of some cloud computing
platforms, such as Amazon Web Services (AWS),
which could improve its performance in case of an
exponential growth of the computational
complexity.
4.4</p>
      <sec id="sec-6-1">
        <title>Front-end presentation</title>
        <p>
          The web database is queryable from the dedicated
section of CorAIt website, prior to registration and
approval. Due to storage issues, and observing the
design of Speech Accent Archive
          <xref ref-type="bibr" rid="ref33">(Weinberger,
2015)</xref>
          , the format of the audio files available on
the online version of CorAIt is .mp3. Samples
coded in other formats (e.g. .wav, .flac, etc.) are
freely available under request.
6 The section 4.3 was written with the contribution of
Antonio Maria Tenace, who provided support on the
graphical implementation of the webapp and was in charge
with the technical aspects of its architecture.
        </p>
        <p>To enable advanced queries, various layers of
metadata were added to each audio file: the
speaker’s mother tongue, gender, age of Italian
language onset, age at the time the sample was
recorded, level of Italian proficiency, Italian
learning method, length of residence in Italy,
proficiency in other foreign languages. Moreover,
information on the type of sample and its quality
was included (Cf. Figure 1).</p>
        <p>
          The corpus has not been transcribed nor
annotated yet. However, following the example of
the Speech Accent Archive
          <xref ref-type="bibr" rid="ref33">(Weinberger, 2015)</xref>
          ,
the sole grammatical sentences from the reading
excerpts were inserted under the samples
corresponding to the reading task.
        </p>
        <p>Besides the embedded audio player – which
allows to listen and download the audio sample –
the window where the single result is displayed
provides biographical and quantitative
information with respect to the speaker who
uttered that speech sample, as well as qualitative
information regarding the audio file (Cf. Figure
2).</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion and future work</title>
      <p>Considering that currently this non-native speech
database presents some imbalance issues as
regards the speakers’ age of Italian language
onset, their proficiency level, as well as the length
of their residence in Italy, the data collection will
be further extended.</p>
      <p>In the future, the web database might also be
enhanced with orthographic and phonetic
transcriptions. Disfluencies (i.e. false starts, filled
and silent pauses, phoneme lengthening,
mispronounced words), mouth clicks, and external
noise could be annotated.
6</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>The author gratefully acknowledges the
informants who participated in this research, and
A. M. Tenace, whose expertise and assistance
were fundamental for the architecture and the
implementation of the webapp.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>L. M. Arslan</surname>
            &amp;
            <given-names>J. H.</given-names>
          </string-name>
          <string-name>
            <surname>Hansen</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Frequency characteristics of foreign accented speech</article-title>
          .
          <source>In Proc. of ICASSP</source>
          , pp.
          <fpage>1123</fpage>
          -
          <lpage>1126</lpage>
          , Munich, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Atagi</surname>
          </string-name>
          &amp;
          <string-name>
            <given-names>T.</given-names>
            <surname>Bent</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Auditory free classification of non-native speech</article-title>
          .
          <source>In J. Phon.</source>
          ,
          <volume>41</volume>
          (
          <issue>6</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Barbera</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>C. Marello</surname>
          </string-name>
          .
          <year>2004</year>
          ,
          <article-title>VALICO (Varietà di Apprendimento della Lingua Italiana Corpus Online): una presentazione</article-title>
          .
          <source>In ITALS 4</source>
          ,
          <string-name>
            <surname>Guerra</surname>
            <given-names>Edizioni</given-names>
          </string-name>
          , Perugia, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>P.</given-names>
            <surname>Boersma</surname>
          </string-name>
          &amp;
          <string-name>
            <given-names>D.</given-names>
            <surname>Weenink</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Praat: doing phonetics by computer</article-title>
          [Computer program].
          <source>Version 6.0</source>
          .
          <fpage>33</fpage>
          . [http://www.praat.org]
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>N. F.</given-names>
            <surname>Chen</surname>
          </string-name>
          et al.
          <year>2015</year>
          . iCALL Corpus:
          <article-title>Mandarin Chinese Spoken by Non-Native Speakers of European Descent</article-title>
          .
          <source>In Proc. of Interspeech</source>
          , pp.
          <fpage>801</fpage>
          -
          <lpage>805</lpage>
          , Dresden, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Cieri</surname>
          </string-name>
          et al.
          <year>2004</year>
          .
          <article-title>The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text</article-title>
          .
          <source>In Proc. of LREC</source>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>71</lpage>
          , Lisbon, Portugal.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Cincarek</surname>
          </string-name>
          et al.
          <year>2004</year>
          .
          <article-title>Speech Recognition for Multiple Non-Accent Groups with Speaker-GroupDependent Acoustic Models</article-title>
          .
          <source>In Proc. of Interspeech</source>
          , pp.
          <fpage>1509</fpage>
          -
          <lpage>1512</lpage>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Cucchiarini &amp; H. Van Hamme</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The JASMIN Speech Corpus: Recordings of Children, Nonnatives and Elderly People</article-title>
          . In P. Spyns &amp; J. Odijk (editors),
          <source>Essential Speech and Language Technology for Dutch</source>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>59</lpage>
          . Springer, Heidelberg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Durand</surname>
          </string-name>
          et al. (editors).
          <source>2014. The Oxford Handbook of Corpus Phonology</source>
          . Oxford University Press, Oxford, UK.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>V.</given-names>
            <surname>Fischer</surname>
          </string-name>
          et al.
          <year>2003</year>
          .
          <article-title>Recent progress in the decoding of non-native speech with multilingual acoustic models</article-title>
          .
          <source>In Proc. of Eurospeech</source>
          , pp.
          <fpage>3105</fpage>
          -
          <lpage>3108</lpage>
          , Geneva, Switzerland.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Gruhn</surname>
          </string-name>
          et al.
          <year>2004</year>
          .
          <article-title>A multi-accent non-native English database</article-title>
          .
          <source>In Proc. of Acoustical Society of Japan</source>
          , Kyoto, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Gruhn</surname>
          </string-name>
          et al.
          <year>2011</year>
          .
          <article-title>Statistical Pronunciation Modelling for Non-native Speech Processing</article-title>
          . Springer, Heidelberg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Heuvel</surname>
          </string-name>
          et al.
          <year>2006</year>
          .
          <article-title>TC-STAR: New language resources for ASR and SLT purposes</article-title>
          .
          <source>In LREC</source>
          , pp.
          <fpage>2570</fpage>
          -
          <lpage>2573</lpage>
          , Genoa, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Tomokiyo</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Recognizing non-native speech. Characterizing and adapting to non-native usage in LVCSR</article-title>
          .
          <source>PhD thesis</source>
          . Carnegie Mellon University, Pittsburgh, USA.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>L. F.</given-names>
            <surname>Lamel</surname>
          </string-name>
          et al.
          <year>1994</year>
          .
          <article-title>The translanguage English database TED</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>ICSLP</given-names>
          </string-name>
          , Yokohama, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Lander</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>CSLU: Foreign Accented English</article-title>
          .
          <source>Release 1.2. LDC2007S08. Linguistic Data Consortium</source>
          , Philadelphia, USA.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>La Rocca</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Tomei</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>West point Russian speech corpus</article-title>
          .
          <source>Tech. Rep</source>
          .,
          <string-name>
            <surname>LDC</surname>
          </string-name>
          , Philadelphia, Pennsylvania, USA.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          Ludwig Maximilian University of Munich.
          <year>1998</year>
          .
          <article-title>Bavarian Archive for Speech Signals</article-title>
          . [http://www.phonetik.uni-muenchen.de/Bas]
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Menzel</surname>
          </string-name>
          et al.
          <year>2000</year>
          .
          <article-title>The ISLE corpus of non-native spoken English</article-title>
          .
          <source>In LREC</source>
          , pp.
          <fpage>957</fpage>
          -
          <lpage>963</lpage>
          , Athens, Greece.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Minker</surname>
          </string-name>
          et al. (editors).
          <year>2014</year>
          .
          <article-title>Spoken dialogue systems</article-title>
          .
          <source>Technology and design</source>
          . Springer, Heidelberg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Morgan</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>West point heroico Spanish speech</article-title>
          .
          <source>Tech. Rep</source>
          .,
          <string-name>
            <surname>LDC</surname>
          </string-name>
          , Philadelphia, Pennsylvania, USA.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Nishina</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Development of Japanese speech database read by non-native speakers for constructing CALL system</article-title>
          .
          <source>In ICA</source>
          , pp.
          <fpage>561</fpage>
          -
          <lpage>564</lpage>
          , Kyoto, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Palermo</surname>
          </string-name>
          .
          <year>2009</year>
          . Percorsi e strategie di
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>sondaggi su ADIL2</source>
          . Guerra Edizioni, Perugia, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Pigeon</surname>
          </string-name>
          et al.
          <year>2007</year>
          .
          <article-title>Design and characterization of the non-native military air traffic communications database</article-title>
          . In ICSLP, Antwerp, Belgium.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>M. Raab</surname>
          </string-name>
          et al.
          <year>2007</year>
          .
          <article-title>Non-native speech databases</article-title>
          .
          <source>In Proc. of IEEE-ASRU</source>
          , pp.
          <fpage>413</fpage>
          -
          <lpage>418</lpage>
          , Kyoto, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Robinson</surname>
          </string-name>
          et al.
          <year>1995</year>
          .
          <article-title>WSJCAM0: A British-English speech corpus for large vocabulary continuous</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          81-
          <fpage>84</fpage>
          , Detroit, USA.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Savy</surname>
          </string-name>
          et al.
          <year>2012</year>
          . DILS - Dialoghi in Italiano Lingua Straniera. [http://www.parlaritaliano.it/index.php/it/corpora-diparlato/794-corpus
          <article-title>-dils-dialoghi-in-italiano-linguastraniera]</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Speechocean</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>King-ASR-L-190: Chinese English Speech Recognition Database</article-title>
          . [http://kingline.speechocean.com/exchange.php?id= 13873&amp;act=view]
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Spina</surname>
          </string-name>
          et al.
          <year>2006</year>
          .
          <article-title>Corpus Parlato di Italiano L2</article-title>
          . [http://elearning.unistrapg.it/osservatorio/corpus/fra mes-cqp.html]
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Vedovelli</surname>
          </string-name>
          .
          <year>2006</year>
          ,
          <article-title>LIPS - Lessico di frequenza dell'Italiano Parlato dagli Stranieri</article-title>
          . In C. Bardel, J. Nystedt (editors),
          <source>Progetto Dizionario ItalianoSvedese</source>
          , pp-
          <volume>55</volume>
          -
          <fpage>58</fpage>
          .
          <source>Acta Universitatis Stockholmiensis</source>
          <volume>22</volume>
          ,
          <string-name>
            <surname>Romanica</surname>
            <given-names>Stockholmiensia</given-names>
          </string-name>
          , Stockholm, Sweden.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Speech Accent Archive</article-title>
          . George Mason University. [http://accent.gmu.edu]
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Wottawa</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>M. Adda-Decker</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>French Learners Audio Corpus of German Speech (FLACGS)</article-title>
          .
          <source>In LREC</source>
          , pp.
          <fpage>3215</fpage>
          -
          <lpage>3219</lpage>
          , Portorož, Slovenia.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Žgank</surname>
          </string-name>
          et al.
          <year>2006</year>
          .
          <article-title>SINOD - Slovenian non-native speech database</article-title>
          .
          <source>In LREC</source>
          , pp.
          <fpage>1620</fpage>
          -
          <lpage>1623</lpage>
          , Genoa, Italy.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>