<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A little bit of bella pianura: Detecting Code-Mixing in Historical English Travel Writing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rachele Sprugnoli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Tonelli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Moretti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Menini</string-name>
          <email>meninig@fbk.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fondazione Bruno Kessler</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trento</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita` di Trento</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Code-mixing is the alternation between two or more languages in the same text. This phenomenon is very relevant in the travel domain, since it can provide new insight in the way foreign cultures are perceived and described to the readers. In this paper, we analyse EnglishItalian code-mixing in historical English travel writings about Italy. We retrain and compare two existing systems for the automatic detection of code-mixing, and analyse the semantic categories mostly connected to Italian. Besides, we release the domain corpus used in our experiments and the output of the extraction.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Il code-mixing e` l’alternanza di
lingue diverse nello stesso testo. Questo
fenomeno e` particolarmente importante
nel dominio dei viaggi, poiche´ aiuta a
comprendere meglio il modo in cui
vengono percepite e descritte culture diverse
da quella dell’autore. In questo lavoro,
analizziamo il code-mixing tra inglese ed
italiano nei testi di viaggio scritti in
inglese e aventi come soggetto l’Italia. A
questo scopo confrontiamo due sistemi
esistenti per il riconoscimento automatico
del code-mixing dopo averli ri-addestrati
e analizziamo le categorie semantiche
connesse alle parole/espressioni italiane.
Inoltre, rilasciamo il corpus e il risultato
dell’estrazione.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Code-mixing is the alternation between two or
more languages that can occur between sentences
(inter-sentential), within the same utterance
(intrasentential), or even inside a single token
(mixing of morphemes). This phenomenon has been
widely studied from the linguistic,
psycholinguistic, and sociolinguistic point of view
        <xref ref-type="bibr" rid="ref11 ref12">(GardnerChloros, 1995; Grosjean, 1995; Ho, 2007)</xref>
        but
there is no consensus on the terminology to be
adopted. In this paper code-mixing is used as an
umbrella term to indicate a manifestation of
language contact subsuming other expressions such
as code-switching, languaging, borrowing,
language crossing
        <xref ref-type="bibr" rid="ref18">(Muysken, 2000)</xref>
        .
      </p>
      <p>
        Code-mixing characterizes communication of
post-colonial, migrant and multilingual
communities
        <xref ref-type="bibr" rid="ref20 ref8">(Papalexakis et al., 2014; Frey et al., 2016)</xref>
        and it emerges in different types of documents,
for example parliamentary debates, interviews
and social media posts
        <xref ref-type="bibr" rid="ref2 ref22 ref6">(Carpuat, 2014; Das and
Gamba¨ck, 2015; Piergallini et al., 2016)</xref>
        . Travel
writings (e.g. guidebooks, travelogues, diaries,
blogs, travel articles in magazines) are affected as
well by this phenomenon that has been studied in
particular by analyzing small corpora of
contemporary tourism discourse through manual
inspection
        <xref ref-type="bibr" rid="ref5">(Dann, 1996)</xref>
        . Even if code-mixing occurs in
less than 1% of the cases
        <xref ref-type="bibr" rid="ref1">(Cappelli, 2013)</xref>
        , it has
several important functions in the travel domain:
it gives a “linguistic sense of place”
        <xref ref-type="bibr" rid="ref4">(Cortese and
Hymes, 2001)</xref>
        , it adds authenticity to a narration, it
provides translation of cultural-specific words and
it is a mean to define social identity (“us” tourists
versus “they” locals)
        <xref ref-type="bibr" rid="ref14">(Jaworski et al., 2003)</xref>
        .
      </p>
      <p>In this work, we investigate the phenomenon
of code-mixing in travel writings, but differently
from previous works we shift the focus of
analysis from contemporary to historical data and from
manual to automatic information extraction. As
for the first point, we present a corpus of more than
3.5 millions words of English travel writings
published between the end of the XIX Century and
the beginning of the XX Century, which we have
retrieved from freely available sources and we
release in a cleaned format. As for automatic
information extraction, we retrain two state-of-the-art
tools to identify English-Italian code-mixing and
evaluate them on a sample of our dataset. We
further launch the best system on the whole dataset
and then we perform a semi-automatic refinement
of the automatic annotation. The corpus, the
training and test data and the outcome of the extraction
are available online1.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        Automatic language identification of monolingual
documents has a long tradition in Natural
Language Processing
        <xref ref-type="bibr" rid="ref13 ref16 ref19">(Hughes et al., 2006; Lui and
Baldwin, 2012)</xref>
        . More recently a new hot topic of
research has emerged, that is the detection of
language at word level in code-mixing texts.
Dedicated workshops and evaluation exercises have
been organized on this task dealing with
different pairs of languages and with social media data
        <xref ref-type="bibr" rid="ref17 ref27 ref3">(Choudhury et al., 2014; Solorio et al., 2014;
Molina et al., 2016)</xref>
        . The most common approach
of the proposed systems is based on Conditional
Random Fields (CRFs) but there are also
implementations of Logistic Regression and deep
learning algorithms.
      </p>
      <p>
        To the best our knowledge, there is no
previous work on the automatic identification of
codemixing in travel writing. Cappelli (2013) and
Gandin (2014) have studied the phenomenon, but
they have mainly used standard corpus
linguistics tools, i.e. WordSmith
        <xref ref-type="bibr" rid="ref26">(Scott, 2008)</xref>
        , to
analyse language contact in English guidebooks, travel
blogs written by expatriates and travel articles
from 2002-2012.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Corpus Description</title>
      <p>
        Differently from the works cited in the
previous Section, we focus on historical texts. To
this end, we collect from Project Gutenberg2 a
corpus of travel writings about Italy written by
English native authors and published between
the country unification and the beginning of the
30’s. We choose this period because in the
second half of the XIX Century the tradition of the
Grand Tour declined and leisure-oriented travels
emerged. This radical transformation was
enabled by technological, economic and
sociological, factors, such as the development of
steampowered ships and of the railway network, the
1https://dh.fbk.eu/technologies/
code-mixing
2https://www.gutenberg.org/
growth of Anglo-American economy and a greater
emancipation of women with more female
travelers
        <xref ref-type="bibr" rid="ref24">(Schriber, 1995)</xref>
        . Moreover, after unification,
new routes to Southern Italy and the islands were
opened, so that travelers’ attention was no longer
limited to the classic destinations in the North and
Central Italy, such as Venice, Florence and Rome
        <xref ref-type="bibr" rid="ref16 ref19">(Ouditt and Polezzi, 2012)</xref>
        .
      </p>
      <p>
        The corpus is made by 57 texts3, divided into
travel narratives (reports, diaries, collections of
letters) and guidebooks, for a total of 3,630,781
tokens. We distinguish between these two types
of text, following a standard classification of
documents in the travel domain. However, the
distinction was not so clear-cut in the period we take
into account as it is now, since reports on
personal travel experiences were often mixed with
practical recommendations and long disquisitions
on art and history. Therefore, we adopt as a rule
of thumb the distinction suggested in
        <xref ref-type="bibr" rid="ref23">(Santulli,
2007)</xref>
        : travel narratives are those told in the first
person, while guidebooks are written in
impersonal form.
      </p>
      <p>The authors of the selected texts belong to
different nationalities (UK, US, Ireland, Australia)
and are both male and female. Some books dwell
on specific cities or regions, others cover different
parts of Italy or even several countries: in the
latter case we extracted only the chapters related to
Italy. Although we made an effort to have a
diverse, well-balanced corpus in terms of content,
author’s gender and nationality, this was only
partially possible because of the limited availability of
online travel books whose text is freely available
and cleaned from OCR errors. The distribution
of tokens according to the year of publication and
type of text is shown in Fig. 1. Details about
authors are given in a spreadsheet provided together
with the corpus.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Code-Mixing Detection</title>
      <p>In this Section we describe the experiments on
code-mixing, comparing the performance of two
available systems in different configurations. We
also detail the post-processing step introduced to
refine the output of the best performing system.</p>
      <p>
        3Thirty of these texts are also available in TEI-XML
format on the website https://sites.google.com/
view/travelwritingsonitaly.
In order to automatically extract Italian words,
expressions and sentences from the corpus
described in Section 3, we train and test two systems
whose source code is available on the web. The
first one (henceforth, langid) is based on
character n-grams (n = 1 to 5) and adopts a weakly
supervised approach, i.e. training data are
monolingual texts of few thousand tokens
        <xref ref-type="bibr" rid="ref15">(King and
Abney, 2013)</xref>
        . This system includes four
classification algorithms: Conditional Random Field
(CRF), Hidden Markov Model (HMM) and
Maximum Entropy Model with and without
generalized expectation criteria (MaxEnt-GE and
MaxEnt). langid has been successfully evaluated on
documents containing English texts mixed with
30 different minority languages such as Zulu and
Chippewa4.
      </p>
      <p>For our experiments, we retrain langid using
a collection of about 300,000 tokens taken from
monolingual Italian and English books, of
different genres, published in the same period of our
corpus5.</p>
      <p>The second system (henceforth,
CodeSwitching), has been developed to detect languages in
texts mixing Latin and Middle English (Schulz</p>
      <sec id="sec-5-1">
        <title>4http://www-personal.umich.edu/</title>
        <p>˜be5FnokrinItgal/iarne:
s“oLuercAevsve/nltuarnegdiidP_irneoclcehaios”eb.ytaCr..Cgozllodi, “Una donna” by S. Aleramo, “Il Valdarno da Firenze al
mare” by G. Carocci, “La vita operosa” by M. Bontempelli,
“Dopo il divorzio” by G. Deledda, “Novelle umoristiche” by
A. Albertazzi, “Lezioni e Racconti per i bambini” by I.
Baccini. For English: “The Adventures of Tom Sawyer” by M.
Twain, “Pioneers of the Old Southwest” by C. L. Skinner,
“The Happy Prince, and Other Tales” by O. Wilde, “Vanished
Arizona” by M. Summerhayes, “The Tale of Peter Rabbit” by
B. Potter, “The Strange Case of Dr. Jekyll and Mr. Hyde” by
R. L. Stevenson.
and Keller, 2016). It implements a CRF
classifier with features generated from TreeTagger
models and word lists of both languages6. Differently
from langid that classifies words as belonging to
one language rather than the other, this latter
system performs a fine-grained annotation by
distinguishing five classes (see below). Since this
system is fully supervised, we create a training set by
manually annotating 3,900 tokens from 4 samples
extracted from our corpus, a size in line with the
training data used in the original paper. The
training data were annotated with 5 different classes:
Italian tokens (i), English tokens (e), punctuation
(p), named entities (NEs) (n), and ambiguous
tokens that belong to the dictionary of both
languages (a).</p>
        <p>
          Both langid and CodeSwitching were evaluated
on the same test set, i.e. two samples of texts
(one from a travel narrative and one from a
guidebook) of 1,623 tokens. The test set was
annotated by assigning to each token a label for English
or Italian, as required by langid, and also
marking punctuation, NEs and ambiguous tokens,
following CodeSwitching scheme. Since the
performance of CodeSwitching is sensitive to the length
of the input file, we split the test set in batches of
40 sentences, replicating the experimental setting
presented in
          <xref ref-type="bibr" rid="ref25 ref8">(Schulz and Keller, 2016)</xref>
          .
4.2
        </p>
        <sec id="sec-5-1-1">
          <title>Evaluation</title>
          <p>Table 1 presents the performances of langid on the
test set: contrary to the results achieved by King
and Abney (2013), HMM – not CRF – proved
to be the best approach. This is likely due to
the greater sparseness of the code-mixing
phenomenon in our dataset with respect to what was
registered in the original corpus, where languages
different from English cover the 56% of the
overall number of tokens.</p>
          <p>Table 2 reports Precision, Recall and F-measure
of the retrained CodeSwitching system. Even if the
overall performance is slightly better than the one
obtained with HMM in langid, the scores for the
detection of Italian tokens (i) are lower (0.82
versus 0.90 in terms of F-measure). Punctuation (i)
and ambiguous tokens (a) are generally detected
with a good performance, while NEs (e)
represent the most challenging class. Given that we are
mainly interested in recognising English and
Ital</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>6https://github.com/sarschu/</title>
        <p>CodeSwitching
ian terms, and that on this task langid performs
better, we run this tool on the whole corpus.
4.3</p>
        <sec id="sec-5-2-1">
          <title>Post-processing</title>
          <p>
            In order to refine the output of langid (see Figure
2), we perform three post-processing steps. First
of all, we check whether tokens tagged as Italian
are included in Morph-it, an Italian lexicon of
inflected forms
            <xref ref-type="bibr" rid="ref29">(Zanchetta and Baroni, 2005)</xref>
            : in this
way we are able to detect false positives. Then, we
run the Polyglot Python module on the corpus to
find out if the processed documents contain other
languages beside English and Italian7. Indeed 27
books result to have a high probability of
including text written also in Latin, French, Germany or
Greek. These books are likely to be problematic
given that langid recognizes only English and
Italian. Information obtained in these two steps are
then used to manually check the outcome of langid
extraction and correct it semi-automatically.
Furthermore, we employ the USAS Italian semantic
tagger
            <xref ref-type="bibr" rid="ref21">(Piao et al., 2015)</xref>
            to obtain a
categorization of the terms tagged as Italian. Based on the
21 semantic classes recognised by USAS, we are
able to understand in which cases and why
writers used to switch their narration from English to
Italian.
5
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>The classification performed with the USAS
tagger shows that Italian is adopted to express
con</p>
      <sec id="sec-6-1">
        <title>7http://polyglot.readthedocs.io/en/</title>
        <p>
          latest/Installation.html
cepts covered by 20 semantic classes, both in
guidebooks and in travel narratives. Only one
USAS class, the one related to “Science and
technology”, is not found in the corpus. Table 5 shows
frequency and examples for each detected class.
As in contemporary travel writings
          <xref ref-type="bibr" rid="ref7">(Francesconi,
2007)</xref>
          , food is well represented: traditional dishes,
drinks and products (e.g. polenta, Chianti,
mortadella) appear together with fruits, vegetables
(e.g. mandarini, finocchio) and also eating
establishments (e.g. osteria, trattoria, locanda). The
attention for Italian art and architecture manifests
itself through the use of many specialized terms
(cassettoni, gotico, giallo antico). The semantic
areas of emotions and psychological processes are
not recorded in previous work on contemporary
texts but are frequent especially in travel reports
(e.g. addolorata, trionfo, simpatico). As for NEs,
city names reveal an increasing interest for towns
in Central regions (for example, Perugia has a high
frequency of occurrence in both genres).
Moreover, following Italy unification, travellers
discovered several locations in the South (e.g. Ragusa,
Catanzaro). Among the most mentioned
people, there are representatives of past Italian
politics (e.g. Lorenzo and Cosimo de Medici), artists
(e.g. Giotto, Dante) and religious figures (e.g.
Madonna, San Michele).
        </p>
        <p>In many cases, the use of Italian is not limited to
single words or multi-token expressions (e.g.
appartamento signorile) but longer utterances are
reported. Texts of both genres contain proverbs (e.g.
chi tardi arriva mal alloggia) and citations, not
only from the canon of Italian literature, such as
Leopardi’s poems, but also from the popular
tradition, such as Tuscan songs (O rosa O rosa O rosa
gentillina). The main difference between travel
narratives and guidebooks is the greater presence
in the former of dialogues or expressions heard by
the author during his/her stay in Italy (voi siete un</p>
        <sec id="sec-6-1-1">
          <title>GUIDEBOOKS</title>
          <p>SEMANTIC CLASS #
names &amp; grammar 29,927
architecture 3,070
movement 2,294
social elements 1,590
materials &amp; objects 717
environment 713
general/abstract terms 580
measurement 340
arts &amp; crafts 231
time 225
life 222
body 211
public domain 205
psyche 198
food &amp; farming 162
entertainment 141
emotion 137
communication 131
money &amp; commerce 127
education 22</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>TRAVEL NARRATIVES</title>
          <p>SEMANTIC CLASS # EXAMPLES
names &amp; grammar 28,694 Donatello
social elements 3,134 popolo
architecture 3,065 palazzo
environment 1,311 lago
movement 1,207 vetturino
materials &amp; objects 965 rosso
general/abstract terms 943 fare
food &amp; farming 665 trattoria
life 479 fiore
measurement 464 grande
time 379 primavera
body 350 braccio
psyche 330 vedere
entertainment 319 marionetta
money &amp; commerce 269 dazio
communication 268 dire
public domain 260 carabiniere
arts &amp; crafts 206 arte
emotion 176 evviva
education 135 maestro
In this work, we presented the first automated
analysis of code-mixing in historical travel
writings. In particular, we focus on English
documents about Italy, and we compare guidebooks
and travel narratives, analysing the semantic
categories mostly related to code-mixing.</p>
          <p>
            In the future, we plan to investigate how
codemixing phenomena relate to content types in travel
writings
            <xref ref-type="bibr" rid="ref28">(Sprugnoli et al., 2017)</xref>
            . Besides, we are
planning to implement an algorithm to
automatically link code-mixing quotations to their original
source text. Finally, we would like to extend our
experiments to recognise code-mixing in
multiple languages, and compare the semantic domains
specific to each language.
          </p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Gloria</given-names>
            <surname>Cappelli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Travelling words: Languaging in english tourism discourse</article-title>
          .
          <source>Travels and translations</source>
          , pages
          <fpage>353</fpage>
          -
          <lpage>374</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marine</given-names>
            <surname>Carpuat</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Mixed-language and codeswitching in the canadian hansard</article-title>
          .
          <source>In Proceedings of EMNLP</source>
          <year>2014</year>
          , page 107.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Monojit</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , Gokul Chittaranjan, Parth Gupta, and
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Overview of fire 2014 track on transliterated search</article-title>
          .
          <source>In Proceedings of FIRE.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppina</given-names>
            <surname>Cortese</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dell</given-names>
            <surname>Hymes</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Languaging in and across human groups. Perspectives on difference and asymmetry</article-title>
          .
          <source>Textus. English Studies in Italy</source>
          ,
          <volume>14</volume>
          (
          <issue>2</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Graham MS Dann</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>The language of tourism: a sociolinguistic perspective</article-title>
          . Cab International.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          and
          <article-title>Bjo¨rn Gamba¨ck</article-title>
          .
          <year>2015</year>
          .
          <article-title>Code-mixing in social media text: the last language identification frontier? Revue TAL</article-title>
          , pages
          <fpage>41</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Sabrina</given-names>
            <surname>Francesconi</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Italian borrowings from the semantic fields of food and drink in English tourism texts</article-title>
          .
          <source>The Languages of Tourism: turismo e mediazione</source>
          , Milano: Unicopli, page
          <volume>129</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Jennifer-Carmen</surname>
            <given-names>Frey</given-names>
          </string-name>
          , Aivars Glaznieks, and Egon W Stemle.
          <year>2016</year>
          .
          <article-title>The DiDi Corpus of South Tyrolean CMC Data: A Multilingual Corpus of Facebook Texts</article-title>
          .
          <source>In Proceedings of CLIC-it.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Stefania</given-names>
            <surname>Gandin</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Investigating loan words and expressions in tourism discourse: A corpus driven analysis on the bbctravel corpus</article-title>
          .
          <source>European Scientific Journal</source>
          ,
          <volume>10</volume>
          (
          <issue>2</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Penelope</given-names>
            <surname>Gardner-Chloros</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Code-switching in community, regional and national repertoires: the myth of the discreteness of linguistic systems. One speaker, two languages: Cross-disciplinary perspectives on code-switching</article-title>
          , pages
          <fpage>68</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>Franc¸ois Grosjean</source>
          .
          <year>1995</year>
          .
          <article-title>A psycholinguistic approach to code-switching: The recognition of guest words by bilinguals</article-title>
          .
          <source>One speaker, two languages: Crossdisciplinary perspectives on code-switching</source>
          , pages
          <fpage>259</fpage>
          -
          <lpage>275</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Judy</given-names>
            <surname>Woon Yee Ho</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Code-mixing: Linguistic form and socio-cultural meaning</article-title>
          .
          <source>The International Journal of Language Society and Culture</source>
          ,
          <volume>21</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Baden</given-names>
            <surname>Hughes</surname>
          </string-name>
          , Timothy Baldwin, Steven Bird, Jeremy Nicholson, and Andrew MacKinlay.
          <year>2006</year>
          .
          <article-title>Reconsidering language identification for written language resources</article-title>
          .
          <source>In Proc. International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>485</fpage>
          -
          <lpage>488</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Jaworski</surname>
          </string-name>
          , Crispin Thurlow, Sarah Lawson, and Virpi Yla¨
          <fpage>nne</fpage>
          -
          <lpage>McEwen</lpage>
          .
          <year>2003</year>
          .
          <article-title>The uses and representations of local languages in tourist destinations: A view from British TV holiday programmes</article-title>
          .
          <source>Language Awareness</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Ben</given-names>
            <surname>King and Steven P Abney</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Labeling the languages of words in mixed-language documents using weakly supervised methods</article-title>
          .
          <source>In Proceedings of HLT-NAACL</source>
          , pages
          <fpage>1110</fpage>
          -
          <lpage>1119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Lui</surname>
          </string-name>
          and
          <string-name>
            <given-names>Timothy</given-names>
            <surname>Baldwin</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>langid. py: An off-the-shelf language identification tool</article-title>
          .
          <source>In Proceedings of the ACL 2012 system demonstrations</source>
          , pages
          <fpage>25</fpage>
          -
          <lpage>30</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Molina</surname>
          </string-name>
          , Nicolas Rey-Villamizar, Thamar Solorio,
          <string-name>
            <surname>Fahad</surname>
            <given-names>AlGhamdi</given-names>
          </string-name>
          , Mahmoud Ghoneim, Abdelati Hawwari, and
          <string-name>
            <given-names>Mona</given-names>
            <surname>Diab</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview for the second shared task on language identification in code-switched data</article-title>
          .
          <source>In Proceedings of EMNLP 2016</source>
          , pages
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Pieter</given-names>
            <surname>Muysken</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Bilingual speech: A typology of code-mixing</article-title>
          , volume
          <volume>11</volume>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Sharon</given-names>
            <surname>Ouditt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Loredana</given-names>
            <surname>Polezzi</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Introduction: Italy as place and space</article-title>
          .
          <source>Studies in Travel Writing</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <fpage>97</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Evangelos</given-names>
            <surname>Papalexakis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dong-Phuong Nguyen</surname>
            , and
            <given-names>A Seza</given-names>
          </string-name>
          <string-name>
            <surname>Dog</surname>
          </string-name>
          <article-title>˘ruo¨z</article-title>
          .
          <year>2014</year>
          .
          <article-title>Predicting code-switching in multilingual communication for immigrant communities</article-title>
          .
          <source>In Proceedings of The First Workshop on Computational Approaches</source>
          to Code Switching.
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Scott</given-names>
            <surname>Piao</surname>
          </string-name>
          , Francesca Bianchi, Carmen Dayrell,
          <string-name>
            <surname>Angela D'egidio</surname>
          </string-name>
          , and Paul Rayson.
          <year>2015</year>
          .
          <article-title>Development of the multilingual semantic annotation system</article-title>
          .
          <source>Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Mario</given-names>
            <surname>Piergallini</surname>
          </string-name>
          , Rouzbeh Shirvani, Gauri S Gautam, and
          <string-name>
            <given-names>Mohamed</given-names>
            <surname>Chouikha</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Word-level language identification and predicting codeswitching points in swahili-english language data</article-title>
          .
          <source>In Proceedings of EMNLP</source>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Santulli</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Le parole ei luoghi: descrizione e racconto</article-title>
          . Antelmi, Donelli/Held, Gudrun/Santulli, Francesca, pages
          <fpage>81</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Mary</given-names>
            <surname>Suzanne Schriber</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Women's place in travel texts</article-title>
          . Prospects,
          <volume>20</volume>
          :
          <fpage>161179</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Sarah</given-names>
            <surname>Schulz</surname>
          </string-name>
          and Mareike Keller.
          <year>2016</year>
          .
          <article-title>Codeswitching ubique est - Language identification and part-of-speech tagging for historical mixed text</article-title>
          .
          <source>In Proceedings of LaTeCH Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Mike</given-names>
            <surname>Scott</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>WordSmith tools version 5</article-title>
          .
          <source>Liverpool: Lexical Analysis Software</source>
          ,
          <volume>122</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Thamar</given-names>
            <surname>Solorio</surname>
          </string-name>
          , Elizabeth Blair, Suraj Maharjan, Steven Bethard, Mona Diab, Mahmoud Ghoneim, Abdelati Hawwari,
          <string-name>
            <surname>Fahad</surname>
            <given-names>AlGhamdi</given-names>
          </string-name>
          , Julia Hirschberg,
          <string-name>
            <given-names>Alison</given-names>
            <surname>Chang</surname>
          </string-name>
          , et al.
          <year>2014</year>
          .
          <article-title>Overview for the first shared task on language identification in code-switched data</article-title>
          .
          <source>In Proceedings of the First Workshop on Computational Approaches</source>
          to Code Switching, pages
          <fpage>62</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Rachele</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          , Tommaso Caselli, Sara Tonelli, and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The Content Types Dataset: a new Resource to Explore semantic and functional Characteristics of Texts</article-title>
          .
          <source>In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>2</volume>
          ,
          <string-name>
            <surname>Short</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pages
          <fpage>260</fpage>
          -
          <lpage>266</lpage>
          , Valencia, Spain, April. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Eros</given-names>
            <surname>Zanchetta</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Morph-it! A free corpus-based morphological resource for the Italian language</article-title>
          .
          <source>Corpus Linguistics</source>
          <year>2005</year>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>