<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DialettiBot: a Telegram Bot for Crowdsourcing Recordings of Italian Dialects</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Federico Sangati</string-name>
          <email>fsangati@unior.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ekaterina Abramova</string-name>
          <email>e.abramova@ftr.ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johanna Monti</string-name>
          <email>jmonti@unior.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nijmegen University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University L'Orientale</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University L'Orientale</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>1049</fpage>
      <lpage>1054</lpage>
      <abstract>
        <p>English. In this paper we describe DialettiBot, a Telegram based chatbot for crowdsourcing geo-referenced voice recordings of Italian dialects. The system enables people to listen to previously recorded audio and encourages them to contribute to building a collective linguistic resource by sending voice recordings of their own spoken dialects. The project aims at collecting a large sample of voice recordings in order to promote knowledge of linguistic variation and preserve proverbs or idioms typical for different local dialects. Moreover, the collected data can contribute to several voice-based Natural Language Processing (NLP) applications in helping them understand utterances in non-standard Italian.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In questo articolo descriviamo
DialettiBot, un chatbot basato su Telegram
per raccogliere registrazioni audio
georeferenziate di dialetti italiani. Il sistema
permette alle persone di ascoltare le
registrazioni precedentemente inserite, e le
incoraggia a contribuire alla costruzione
di questa risorsa linguistica collettiva,
attraverso l’invio di registrazioni audio
nel proprio dialetto. Il progetto mira
a raccogliere una grande mole di
registrazioni che possono aiutare a
promuovere la conoscenza delle variazioni
linguistiche e la salvaguardia dei proverbi o
modi di dire tipici di ogni dialetto locale.
I dati raccolti possono inoltre contribuire
a diverse applicazioni del trattamento
automatico del linguaggio (TAL) che hanno
bisogno di essere adattate per
comprendere espressioni dialettali.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>It is commonly known that Italy has an abundance
of different dialects, such as Florentine, Venetian,
and Neapolitan. These dialects are not only
characterized by simple phonetic variation as it is
usually meant by this term, but they are proper
Romance languages, with a fully developed grammar
and lexicon. As Repetti puts it:</p>
      <p>The Italian ‘dialects’ [...] are daughter
languages of Latin and sister languages
of each other, of standard Italian, and of
other Romance languages, and they may
be as different from each other and from
standard Italian as French is from
Portuguese. (Repetti, 2000)</p>
      <p>This dialectical variety is a resource that
deserves to be studied and preserved for both
cultural and applied reasons. The former, because
it is quickly disappearing with less and less
people who regularly use dialect at home and in
public places. According to UNESCO “Atlas of the
World’s Languages in Danger”,1 there are about
2,500 endangered languages worldwide. In Italy,
thirty dialects are at risk of extinction, such as
friulano, ladino and veneciano.2 The applied
motivation is that in recent years we have witnessed a
significant growth in the number of voice-based NLP
applications (such as virtual assistants), which are
currently not trained on local dialects and
therefore perform poorly with a number of Italian
speakers.</p>
      <p>In this paper we present a freely available tool
that enables geo-referenced recording of Italian
dialects: DialettiBot, a Telegram based chatbot,
whose aim is to collect a large sample of voice
recordings, promoting preservation of linguistic</p>
      <sec id="sec-2-1">
        <title>1http://www.unesco.org/languages-atlas</title>
        <p>2http://www.culturaitalia.it/opencms/en/contenuti/focus/
UNESCO_warns_that_thirty_Italian_dialects_are_at_risk_
of_extinction.html?language=en
variation and its use in NLP applications. The rest
of the paper is organized as follows: in section 2
we describe related work, in section 3 the
implemented system and in section 4 the collected data.
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Related work</title>
      <p>
        There has been an extensive linguistic research of
Italian dialects
        <xref ref-type="bibr" rid="ref1 ref2 ref9">(Lepschy and Lepschy, 1992;
Belletti, 1993; D’Alessandro et al., 2010)</xref>
        . Here we
summarize a number of projects that relate to the
idea of gathering linguistic recordings for
producing a map of dialects. We also point out their
limitations that inspire our project.
      </p>
      <p>
        VIVALDI project the “Vivaio Acustico delle
Lingue e dei Dialetti d’Italia” is a
collection of recordings and transcriptions of fixed
phrases in the dialects of different cities from
all regions in Italy
        <xref ref-type="bibr" rid="ref7">(Kattenbusch et al., 1998)</xref>
        .
Unfortunately, the project is no longer active
and has mainly focused on a finite set of
chosen sentences, as opposed to spontaneous
utterances.
      </p>
      <p>
        LOCALINGUAL A web application for
crowdsourcing recordings from around the world.
This project is the one that most closely
relates to ours. The main difference is that it is
not restricted to a specific country, does not
use geo-locations and works via a web
application, which makes it difficult to be used on
mobile devices or in case of poor data
connection.3
ALF Atlas Linguistique de la France: an
influential dialect atlas of Romance varieties
in France published in 13 volumes between
1902 and 1910
        <xref ref-type="bibr" rid="ref3">(Gilliéron and Edmont, 1902)</xref>
        .
An example of more recent work of this type
is
        <xref ref-type="bibr" rid="ref5">Hall, Damien (2012</xref>
        ).4
ALD Linguistic Atlas of Dolomitic Ladinian and
neighbouring Dialects (Skubic, 2000). The
project studies the linguistic variation
between dialects of the region which covers the
Grisons and Friuli region.5
IDEA The International Dialects of English
Archive was created in 1998 as the
internet’s first archive of primary-source
recordings of English-language dialects and accents
      </p>
      <sec id="sec-3-1">
        <title>3https://localingual.com</title>
        <p>4http://cartodialect.imag.fr/cartoDialect/accueil
5https://www.micura.it/en/activities/ald-linguistic-atlas
as heard around the world. With roughly
1,400 samples from 120 countries and
territories, and more than 170 hours of recordings,
IDEA is now the largest archive of its kind.6
MICROCONTACT aims at developing a theory
of syntactic change by observing the
evolution of the dialects spoken by Italians who
have migrated to North and South America
during the 20th century.7</p>
        <sec id="sec-3-1-1">
          <title>SPEAKUNIQUE and VOCALID are two sim</title>
          <p>ilar projects that aim at collecting English
voice sample from different regions for
creating personalized digital voices for
communication text to speech devices.8</p>
          <p>
            Our project aims to be an updated and
continuously evolving initiative that can capture
spontaneous (living) dialectical variation over the whole
Italian territory by being freely accessible and easy
to use for a variety of non-specialists. As such,
the project follows methodological practices
similar to other citizen-science projects
            <xref ref-type="bibr" rid="ref11 ref4 ref6">(Gurevych and
Zesch, 2013; Simpson et al., 2014; Hosseini et al.,
2014)</xref>
            , it incorporates a GWAP9 feature
            <xref ref-type="bibr" rid="ref8">(Lafourcade et al., 2015)</xref>
            , and fits within the line of
‘explicit crowdsourcing’ as defined by the
EnetCollect10 COST11 action.
3
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>System description</title>
      <p>In order to crowdsource recordings from Italian
dialects, we have built a Telegram chatbot:
DialettiBot.12 As shown in the screenshot in figure 1,
the user can interact with the bot via a standard
dialogue chat interface in a Telegram application
which is freely available for all mobile or desktop
operating systems.13 Apart from textual input, the
interface provides a small keyboard of buttons that
changes during the dialogue flow to simplify the
interaction. In addition, the bot is able to accept
vocal recordings and GPS locations.</p>
      <p>The bot gives the possibility to the user to listen
to approved recordings or to add new ones.</p>
      <p>
        In the listening mode, it is possible to search
for recording based on location or view the list
6https://www.dialectsarchive.com
7https://microcontact.sites.uu.nl/project
8https://www.speakunique.org, https://www.vocalid.co
9Game with a purpose.
10http://enetcollect.eurac.edu
11European Cooperation in Science and Technology.
12https://t.me/dialettibot
13https://telegram.org/apps
of the most recent recordings. As an element of
gamification
        <xref ref-type="bibr" rid="ref8">(Lafourcade et al., 2015)</xref>
        , there is the
possibility to ask for a random recording and try to
guess its location. The user would then receive a
feedback about the distance between the guessed
location and the correct one. With this simple
game we gather valuable data that would enable
us to plot a type of confusability matrix between
dialects, i.e., how much a dialect of place A
resembles a dialect of place B.
      </p>
      <p>In the recording mode, the user is asked to
submit a freely chosen vocal recording of a sentence,
that can be a simple phrase or a proverb, typical for
their dialect. In addition, the user is asked to
indicate the place where the dialect comes from (either
by sending a GPS location or inputting the name of
the place – in case the user is not currently located
in the place associated with the dialect), and
optionally the translation of the recording in Italian.
As soon as the recording is submitted, the
administrator of the system receives a notification (via
the bot) with the new recording and is asked to
approve or reject the contribution. Typical causes of
rejection are too much background noise and
explicitly offensive utterances. In case of approval,
the recording is inserted in the database and
becomes readily available to other users in the
listening mode.14</p>
      <p>In addition to the bot application, we developed
a web application15 (see figure 2) for visualizing
the approved recordings in a map and giving the
possibility to click on each of them to listen to the
audio and read the translation.
3.1</p>
      <sec id="sec-4-1">
        <title>Technical Specification</title>
        <p>The bot is implemented in Python using the
telegram bot API.16 We chose to deploy the system via
a chatbot (as opposed to a mobile app or web
application) because it is much faster to build and to
maintain since all the major functionalities (voice
recordings, GPS location) are already embedded
in the chat application and immediately
accessible via simple API calls. Moreover, the system
works on all mobile and desktop platforms
without the need to build system-specific versions.
Fi14In the future, there is a possibility to implement an
additional validation step where other users or experts might flag
some contribution as not being representative of a dialect.
15http://dialectbot.appspot.com/audiomap/mappa.html
16https://core.telegram.org/bots/api
nally, the simplified interface of a chatbot is
particularly suitable to elderly people which are one
of the most valuable target groups of the project,
and can be easily used for recording other people
while traveling also in case of no data connection
(recordings are saved locally and uploaded to the
server when data connection is again available).</p>
        <p>The server behind DialettiBot is hosted by the
Google Application Engine (GAE) framework and
the data is stored in the integration datastore. The
GAE technology guarantees full scalability up to
an unrestricted number of users which could
enable producing a significantly large volume of
recordings. The same system also serves the web
application with the map of the recordings
illustrated in figure 2, which has been implemented in
javascript using the Leaflet17 library.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Collected data</title>
      <p>The first version of DialettiBot has been deployed
in January 2016. Since then, 1,886 users have
interacted with the system and have submitted 255
voice recordings out of which 220 have been
approved.18 About 14% of users who interacted with
the system contributed a recording.</p>
      <p>Figure 3 shows the bar chart with the
distribution of the approved recordings over time. The
plot shows that the number of contributions in
2017 (31) has been significantly lower than in
2016 (117) , whereas in 2018 the number is
increasing again (72 in the first 3 quarters of the
year).</p>
      <p>Figure 4 shows the distribution of the approved
recordings on the map of Italy, with the counts
clustered by proximity (heat map). Campania is
the region with most recordings (38), followed
by Lazio (35), Trentino-South Tyrol and Sicily
(27), Puglia (22), Veneto (15), Piedmont and
Tuscany (12), Calabria and Lombardy (9), Basilicata
(5), Emilia-Romagna, Friuli-Venezia Giulia and
Marches (2), Abruzzo, Molise and Sardinia (1).
Currently we have no recordings from Liguria,
Umbria and Valle d’Aosta.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and future work</title>
      <p>We have presented DialettiBot, a chatbot
system based on Telegram for crowdsourcing
georeferenced recordings of Italian dialects.
17https://leafletjs.com
18As of 31st of September 2018.
19Created via https://mapmakerapp.com.
Preliminary tests show that the system can be
easily used by anyone who wishes to collect data
in the field as well as the dialect speakers
themselves. The recording quality is good and the data
is easily exportable to be used for further
processing in the service of linguistic research or NLP
applications. At the same time, the current state of
the project suffers from a number of limitations
that need to be addressed in future work and that
we discuss next.</p>
      <p>First, the preliminary tests have not been
informed by a detailed linguistic study of dialectical
variation nor have we implemented a
methodology for data collection. This is because the tests
have been carried out as a proof-of-concept for
the technology used to collect linguistic resources
rather than a full-fledged linguistic project. Future
tests will require a more careful consideration for
dialect characteristics in the Italian language, the
type of data that would be most valuable
(spontaneous speech vs a set of set sentences etc.) and a
construction of precise, reproducible instructions
for the contributors.</p>
      <p>Second, as described in section 3, we make use
of a centralized validation procedure to approve a
subset of recordings. However, since we have no
complete knowledge of all Italian dialects we may
end up accepting recordings which are not mapped
to the correct location. In the future, we would like
to decentralize the procedure, by delegating the
approval to a higher number of volunteers spread
out in all the regions, so that each new recording
will get validated by the closest volunteer.</p>
      <p>Finally, the number of users and recordings
collected so far is relatively modest. This is due to
the fact that no effort has been undertaken so far to
promote its use by researchers or the general
public. Accordingly, the current goal of the project
is to get support from cultural institutions (both at
a local and at a national level) to help us engage
the citizens in this crowdsourcing effort, as well
as academic partners to further refine the
methodology and extend the chatbot capabilities.</p>
      <p>We believe this project could contribute to help
safeguard the Italian dialectic richness and collect
useful resources for NLP applications, as we
intend to make all recordings openly available for
other researchers to use.20</p>
      <p>20We are planning to upload the data to the Common
Language Resources Infrastructure (CLARIN).</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We kindly acknowledge all users who have so
far contributed to the project by providing audio
recordings of their dialects, and the three
anonymous reviewers for their useful feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Belletti</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Syntactic Theory and the Dialects of Italy. Volume 9 of Linguistica (Turin, Italy)</article-title>
          .
          <source>Rosenberg &amp; Sellier.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Roberta D'Alessandro</surname>
            , Adam Ledgeway, Ian Roberts, and
            <given-names>Frank</given-names>
          </string-name>
          <string-name>
            <surname>Nuessel</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Syntactic Variation: The Dialects of Italy</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Jules</given-names>
            <surname>Gilliéron</surname>
          </string-name>
          and Ed. Edmont.
          <year>1902</year>
          . Atlas linguistique de la France,. H. Champion„ Paris,.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          and
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Zesch</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Collective intelligence and language resources: Introduction to the special issue on collaboratively constructed language resources</article-title>
          .
          <source>Lang</source>
          . Resour. Eval.,
          <volume>47</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Hall</surname>
          </string-name>
          , Damien.
          <year>2012</year>
          .
          <article-title>Vers un nouvel atlas linguistique de la france</article-title>
          .
          <source>SHS Web of Conferences</source>
          ,
          <volume>1</volume>
          :
          <fpage>2171</fpage>
          -
          <lpage>2189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Mahmood</given-names>
            <surname>Hosseini</surname>
          </string-name>
          , Keith Phalp, Jacqui Taylor, and
          <string-name>
            <given-names>Raian</given-names>
            <surname>Ali</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The four pillars of crowdsourcing: a reference model</article-title>
          . IEEE Eighth International Conference on Research Challenges in Information Science.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Dieter</given-names>
            <surname>Kattenbusch</surname>
          </string-name>
          , Carola Köhler, Marcel Lucas Müller, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Tosques</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>VIVALDI project: Vivaio acustico delle lingue e di dialetti d'italia</article-title>
          . https://www2.hu-berlin.de/vivaldi.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Lafourcade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joubert</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.L.</given-names>
            <surname>Brun</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Games with a Purpose (GWAPS)</article-title>
          .
          <source>Focus Series in Cognitive Science and Knowledge Management</source>
          . Wiley.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>A.L.</given-names>
            <surname>Lepschy</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.C.</given-names>
            <surname>Lepschy</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>The Italian Language Today</article-title>
          . Hutchinson university library. Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Lori</given-names>
            <surname>Repetti</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Phonological Theory and the Dialects of Italy</article-title>
          . John Benjamins Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Robert</given-names>
            <surname>Simpson</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kevin R. Page</surname>
          </string-name>
          , and David De Roure.
          <year>2014</year>
          .
          <article-title>Zooniverse: Observing the world's largest citizen science platform</article-title>
          .
          <source>In Proceedings of the Companion Publication of the 23rd International Conference on World Mitja Skubic</source>
          .
          <year>2000</year>
          .
          <article-title>Ladinia linguistica in una monumentale opera: Atlante linguistico del ladino dolomitico e dei dialetti limitrofi - ald1, dr</article-title>
          . ludwig reichert verlag, wiesbaden
          <year>1998</year>
          . Linguistica,
          <volume>40</volume>
          (
          <issue>1</issue>
          ):
          <fpage>188</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>