<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Analysis of Twitter Corpora and the Di erences between Formal and Colloquial Tweets</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Meritxell Gonzalez Oxford University Press, Oxford, United Kingdom Universitat Politecnica de Catalunya</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work reviews recent publications addressing the Twitter translation task, and highlights the lack of appropriate corpora that represents the colloquial language used in Twitter. It also discusses the most well-know issues in the Twitter genre: the use of hashtags and the amount of OOVs, with especial focus in comparing the di erences between formal and colloquial texts. Resumen: Este trabajo resume las publicaciones recientes en el area de la traduccion automatica de tweets, destacando la falta de un corpus que represente el lenguaje coloquial presente en Twitter. Tambien se tratan los problemas mas conocidos del genero de Twitter: el uso de hashtags i la gran cantidad de palabras OOV, con especial enfoque en las diferencias entre tweets formales y coloquiales.</p>
      </abstract>
      <kwd-group>
        <kwd>/Palabras clave</kwd>
        <kwd>corpus</kwd>
        <kwd>tweets</kwd>
        <kwd>hashtags</kwd>
        <kwd>and OOV</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The success and increasing popularity of
microblogging has raised the need to analyse
and process its content. Traditional methods
for natural language processing fail when
applied over these texts. The reason
is not circumscribed to few nor simple
issues. Roughly, microblogs documents do
not follow the traditional structure of a
formal text or document, they use a number
of language variants, styles and registers
among other linguistic phenomena, and can
even include multimedia content as a way of
communication
        <xref ref-type="bibr" rid="ref10 ref10 ref11 ref13 ref2 ref3 ref3 ref4 ref4 ref5 ref8">(Jehl, 2010; Fabrizio Gotti
and Phillippe Langlais and Atefeh Farzindar,
2014; Kaufmann, Max and Kalita, Jugal,
2010; Bertoldi, Nicola and Cettolo, Mauro
and Federico, Marcello, 2010)</xref>
        .
      </p>
      <p>Machine Translation (MT) is a hard task
within the natural language processing eld.
It has received considerable attention during
the last decades, and it is still an active eld
with many research challenges. As in other
natural language processing tasks, it counts
among its di culties the ambiguity of the
language, and the need of corpora and a gold
standard. The former can be addressed by
analysing the context in which a sentence
occur, while the second has been typically
addressed by combining large amounts of
general purpose data and smaller subsets of
domain speci c datasets. The creation of
a gold standard in MT requires the use of
parallel data that helps to assess the quality
of the output.</p>
      <p>
        When addressing the automatic
translation within the microblogging
genre, one has to deal with the additional
di culty of having little or no context and
the fact that microblogs exhibit eeting
domains. Twitter is not di erent from
other microblogs, and has, in addition, its
own particularities. As described in
        <xref ref-type="bibr" rid="ref8">(Jehl,
2010)</xref>
        , tweets actually share the spontaneity
and expressiveness of the spoken language,
but limited to 140 characters. Due this
constraint, tweets have usually a very
simple syntax. However, they are mined
of ungrammaticalities, misspellings and an
unlimited number of lexical variants created
out of the human imaginary and the common
ground of part of the audience.
      </p>
      <p>In this document, Section 2 summarises
recent studies in this eld and di erent
approaches followed to address these
phenomena. Next, Sections 3 to 5 give a
numerical analysis of 6 di erent corpora
of tweets written in Basque, Catalan, and
Spanish. The goal of this analysis is to
sketch the content of the Twitter messages
(tweets), highlight which are their principal
characteristics and discuss the di erences
between formal and colloquial tweets.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Recent Work on Twitter Translation</title>
      <p>The automatic translation of tweets, in
general, is more di cult than regular MT.
Although the MT community has already
addressed the translation of tweets, there are
still few works in this area, mainly because
of the lack of corpora, and especially those
showing a fair representation of colloquial
texts. The number of authors publishing
content in multiple languages is not small,
but their messages tend to be correct and
well structured, in contrast to those posted
by the gross of the users.
2.1</p>
      <sec id="sec-2-1">
        <title>Twitter Corpora</title>
        <p>
          The availability of parallel corpora for
Twitter is growing but still scarce. The
following four works gathered parallel data
following diverse approaches, but them all
contain formal texts only.
          <xref ref-type="bibr" rid="ref7">(Gotti, Fabrizio
and Langlais, Philippe and Farzindar,
Atefeh, 2013)</xref>
          gathered data from Canadian
Government Agencies, written in French
and English. This work describes an MT
system that uses in-domain parallel data
crawled from the links appearing in the
tweets. Hence, tuning was conducted with
documents from the same domain. The
corpus built in
          <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Ling, Wang and Marujo,
Luis and Dyer, Chris and Black, Alan W
and Trancoso, Isabel, 2014)</xref>
          contains tweets
written in Chinese and English. This work
describes a tool and a methodology to help
users to identify parallel excerpts in the
messages and to annotate their boundaries.
The data obtained with this method was
fairly cheap (crowd-sourced) and it resulted
to have a high degree of quality.
          <xref ref-type="bibr" rid="ref6 ref9">(Jehl, Laura
and Hieber, Felix and Riezler, Stefan, 2012)</xref>
          used a corpus of Arabic sentences that were
manually translated into English. The data
was crawled by ltering the topic (Arabic
Spring) and was cleaned and pruned, also
by means of crowd-sourcing. Finally, the
shared task described in
          <xref ref-type="bibr" rid="ref1">(Alegria et al., 2015)</xref>
          distributed a collection of parallel corpora
in the languages spoken in the Iberian
peninsula. These corpora have been used in
this study and they are detailed in Section 3.
        </p>
        <p>
          In contrast, the following four works
deal with the noisy input from colloquial
texts, but either they do not belong to
the Twitter genre or they do not contain
parallel data.
          <xref ref-type="bibr" rid="ref10 ref3 ref4">(Kaufmann, Max and Kalita,
Jugal, 2010)</xref>
          describes an MT system able
to translate from colloquial English into
standard English. The rationale is that
traditional NLP techniques can be applied
over standardised text. Their methodology
includes the use of aligned data from a
corpus of SMSs that contains most common
acronyms and short forms.
          <xref ref-type="bibr" rid="ref10 ref3 ref4">(Bertoldi, Nicola
and Cettolo, Mauro and Federico, Marcello,
2010)</xref>
          and
          <xref ref-type="bibr" rid="ref6 ref9">(Formiga, Llu s and Fonollosa,
Jose A. R., 2012)</xref>
          address the problem
of translating noisy input. The former
by trying to simulate and generate noisy
input automatically; the latter by adding a
preprocessing layer to convert the input into
clean text. Finally, the corpus described
in
          <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Alegria, In~aki and Aranberri, Nora
and Comas, Pere R and Fresno, V ctor
and Gamallo, Pablo and Padro, Lluis and
San Vicente, In~aki and Turmo, Jordi and
Zubiaga, Arkaitz, 2014)</xref>
          was distributed to
the participants of the TweetNorm shared
task. This is a monolingual corpus of Spanish
tweets. Since this corpus has been used in
this study it is further detailed in Section 3.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Linguistic Phenomena</title>
        <p>
          Although the previous works addressed
di erent problems, they share a common
ground on the principal di culties of the
Twitter genre. First, the translation of
hashtags is an open issue that includes its
segmentation, identi cation and analysis of
its role in sentences
          <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Fabrizio Gotti and
Phillippe Langlais and Atefeh Farzindar,
2014)</xref>
          . Second, the correct tokenisation of
the text is essential but di cult due the
extreme noisiness of the text. Also, making
the translation t in 140 characters can harm
the quality of the output, although
          <xref ref-type="bibr" rid="ref8">(Jehl,
2010)</xref>
          addressed this issue in her thesis and
reported good results.
        </p>
        <p>
          The increasing interest in the eld has
promoted the design of tools to create
especialised corpora. However, the human
translation of tweets also raises open
questions
          <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Subert and Bojar, 2014)</xref>
          . For
instance, how to translate idioms and slang,
out-of-vocabulary words, onomatopoeias,
emphasises (jajaaaaa), or irony. But also,
how to approach the translation of hashtags
and symbols (such as emoticons), how to
interpret wrong syntax, nd the translated
version of a link, and t the nal translation
into 140 characters, among others.
        </p>
        <p>
          All in all, the creation of synthetic corpus
to simulate these phenomena seem a feasible
approach
          <xref ref-type="bibr" rid="ref10 ref3 ref4">(Bertoldi, Nicola and Cettolo,
Mauro and Federico, Marcello, 2010)</xref>
          , yet out
of the scope of this study. Last, but not least,
an appropriate methodology and measures
to assess the quality of Twitter translations
including its particular characteristics has
not been addressed so far.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Description of the Used Corpora</title>
      <p>
        The next sections analyse six datasets of
tweets from the Tweet-Norm
        <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Alegria, In~aki
and Aranberri, Nora and Comas, Pere R
and Fresno, V ctor and Gamallo, Pablo and
Padro, Lluis and San Vicente, In~aki and
Turmo, Jordi and Zubiaga, Arkaitz, 2014)</xref>
        ,
Tweet-MT
        <xref ref-type="bibr" rid="ref1">(Alegria et al., 2015)</xref>
        and Social
Media
        <xref ref-type="bibr" rid="ref12">(Roser Saur , 2013)</xref>
        corpora. The
goal is to discuss a few of the phenomena
mentioned in the previous section.
      </p>
      <p>A set of four datasets was obtained
from the Tweet-MT corpora. It consists
of 2 bitexts for Catalan{Spanish and
Basque{Spanish language pairs. The four
datasets contain both, the development and
the test sets for each language: CAES.ca,
CAES.es, EUES.eu and EUES.es. The
tweets in these datasets were obtained from
a sample of manually selected accounts
of authors that tend to tweet in various
languages, being namely public organisations
and personalities. Hence, the content of the
messages is mainly formal, i.e., they do not
contain misspellings and do not abuse of the
use of symbols.</p>
      <p>The fth dataset, TNORM, was obtained
from the Tweet-Norm corpus that gathered a
random selection of geolocated tweets within
the Iberian peninsula, excluding multilingual
areas where other languages than Spanish
are spoken. The corpus was processed
to identify and annotate out-of-vocabulary
words. Hence, it contains not only correct
messages, but also colloquial ones. The
dataset used in this work contains the two
development sets and the test provided in the
workshop.</p>
      <p>The last dataset used in this work
is TSM. It is a portion of the Social
Media Corpus, and in particular the corpus
of tweets in Spanish. It contains a
EUES.eu</p>
      <p>EUES.es
# tweets
# tokens
avg. tokens/tweet
# tweets
# tokens
avg. tokens/tweet
# tweets
# tokens
avg. tokens/tweet
general domain set of tweets randomly
selected. So similarly to TNORM, it
contains both formal and colloquial tweets.
They were manually processed to classify
them according to the language of the
tweet and annotate di erent layers such
as communication function, polarity, target,
and topic. This process included some clean
up of the twitter mark-up for privacy reasons.
Hence, the author id and user mentions,
hashtags and URLs were substituted with
the labels @USER, #HASHTAG and [URL],
respectively.</p>
      <p>
        The six datasets were processed to
have similar characteristics: the tokens
that correspond to the author id and RT
(re-tweet) were removed when present, and
they were tokenised using an adaptation
to Spanish and Catalan languages of the
Twokenize tool
        <xref ref-type="bibr" rid="ref10 ref3 ref4">(Brendan O'Connor and
Michel Krieger and David Ahn, 2010)</xref>
        .
Table 1 shows the number of tweets, the
number of tokens and the average number of
tokens per tweet in each corpus. Regardless
the di erences in nature of the datasets and
their size, they show a similar number of
tokens per tweet, being CAES.ca the dataset
with longer ones and EUES.eu the shortest.
The messages in the two colloquial corpora
TNORM and TSM seem to have slightly
shorter posts compared with their formal
ones in the same language CAES.es and
EUES.es.
      </p>
      <p>Although tweets are similar in length,
a deeper analysis of their content shows
remarkable di erences between the formal
and the colloquial corpora. This section
analyses the use of user mentions and
URLs whereas Section 4 analyses the use
of hashtags. Although dealing with user
EUES.eu</p>
      <p>EUES.es
# URLs
avg. URLs/tweet
% URLs wrt. tokens
# URLs
avg. URLs/tweet
% URLs wrt. tokens
# URLs
avg. URLs/tweet
% URLs wrt. tokens</p>
      <p>
        In contrast, the use of URLs seems to be
consistent across the two types of datasets.
The four bitexts contain almost the same
number of URLs, and we can nd almost
one URLs in each tweet. In return, TNORM
and TSM contain a remarkable small number
of URLs, less than 0:1% per tweet. Out of
curiosity, the majority of URLs in the bitexts
link to documents in the same language
as the tweet. Given that the selected
authors post multilingual messages, it seems
reasonable that they also link to the right
URL when available.
# hashtags
# hashtag types
# avg. hashtags/tweet
% hashtags wrt. tokens
# tweets &gt; 1 hashtag
# hashtags
# hashtag types
# avg. hashtags/tweet
% hashtags wrt. tokens
# tweets &gt; 1 hashtag
# hashtags
# hashtag types
# avg. hashtags/tweet
% hashtags wrt. tokens
# tweets &gt; 1 hashtag
This section analyses the use of hashtags in
the datasets. This study and the next one in
Section 5 follow the procedure in
        <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Fabrizio
Gotti and Phillippe Langlais and Atefeh
Farzindar, 2014)</xref>
        that resulted very clear and
appropriate to this end. Table 3 shows some
statistics on the occurrences of hashtags.
The di erent number of hashtags between
formal and colloquial datasets is noticeable.
The former contains more than one hashtag
per tweet, whereas the latter contains a
remarkable low number of them.1 It seems
to indicate that formal tweets tend to use
hashtags to categorise its topic and, maybe,
create a trend. This is also re ected in
Figure 1: the most of the formal tweets,
in the bitexts, contain one or two hashtag,
whereas the most of the colloquial ones have
none.
      </p>
      <p>A more interesting issue is the translation
of hashtags. In terms of the number of
occurrences, each side of the bitexts contain
a similar amount. However, the number
of hashtag types in CAES.ca is much lower
than the ones in CAES.es. A peer review
of the hashtag sets reveals that the Spanish
versions contain more written variants than
their counterparts in Catalan. For instance,
the hashtag \#revistapremsa" (Catalan) has
four variants in the Spanish text: \#revista",
1The number of hashtag types in TSM is 1 because
the corpus contains only the #HASHTAG label.
TSM
TNORM
EUES.es
EUES.eu
CAES.es
CAES.ca
% tweets with a prologue
% tweets with an epilogue
% of # in a prologue
% of # in an epilogue
% tweets with a prologue
% tweets with an epilogue
% of # in a prologue
% of # in and epilogue
% tweets with a prologue
% tweets with an epilogue
% of # in a prologues
% of # in a epilogues
EUES.eu</p>
      <p>EUES.es
\#revistadeprensa", \#revistaprensa", and
\#revistaprensa".</p>
      <p>
        According to
        <xref ref-type="bibr" rid="ref11 ref13 ref2 ref5">(Fabrizio Gotti and
Phillippe Langlais and Atefeh Farzindar,
2014)</xref>
        , hashtags can be classi ed by the
role they play in the text. They distinguish
between hashtags that appear at the
beginning of the text (prologue), in the text
(inline) and at the end of the text (epilogue).
Correctly identifying this role is important
since a number of hashtags may have a
syntactic function inside the text (inline), or
can help to identify the domain of the text
(prologue and epilogue). A simple heuristic
was used to split the tweets into these three
parts, and the results shown are in line with
the mentioned study. We can observe, in
Table 4, how the hashtag role within the
text varies in each corpus. Although in
di erent proportion, the gross of hashtags in
the formal datasets appear in the epilogue,
which indicates there is a common practice
to add any hashtag at the end of the tweet.
In contrast, the colloquial datasets have a
very few proportion of tweets with either
a prologue or an epilogue, but a higher
proportion of them appear in the prologues
(in comparison to the formal tweets).
This behaviour may simply indicate that
colloquial tweets do not follow necessarily
any common practice. All datasets actually
exhibit a low rate of tweets having a prologue,
although the EUES bitext show a remarkable
higher number in comparison to the rest.
Finally, it is worth to note that, although
the number of hashtags is lower in the
colloquial texts, roughly half of them appear
inline, and hence, they play a syntactic role
in the message. This is important since
they may contain an essential part of the
semantics and thus worth to deal with them.
Unfortunately, hashtags contains mainly of
out-of-vocabulary words, as discussed next
in Section 5.
5
      </p>
      <p>On the OOV words in Twitter
The use of out-of-vocabulary (OOV) words
in Twitter has been claimed to be a hard
issue. The reason is not only the high number
of misspellings, symbols and orthographic
errors, that could be partially tackled by
using spell-checkers, but also the use of
speci c lexica and lexical variants. For
instance, the use of word combinations (e.g.,
in hashtags), the combination of di erent
languages (especially in multilingual regions,
but also English terms) and the unlimited
ability of the microblogging sphere to invent
new terms.</p>
      <p>This section gives a numerical analysis
of OOVs that occur in Twitter. In order
to conduct this analysis, the datasets were
processed to remove the user mentions and
URLs, since them all are tokens that do
not need to be translated. Some variants
of the datasets were built. First, only the
CAES bitext was used due the lack of a
Language Model (LM) for Basque. Then,
since the TNORM annotations provide the
corrected forms for some OOV tokens (only
spelling variants), they were used to build
a new dataset TNORM-S were OOVs were
substituted with the correct word when
available. In addition, two di erent versions
TNORM</p>
      <p>TNORM-S
were created out of each dataset. In the
rst one (clean data), the hashtags were kept
(the # symbol was removed) since they play
an important role in the text, carry part
of the semantics of the message and need
to be translated in most of the cases. In
the second dataset, all the hashtags were
removed. The purpose of this second version
is to highlight the impact of hashtags in the
perplexity estimation of the texts.</p>
      <p>Table 5 shows the results of this analysis.
As expected, colloquial datasets contain a
higher number of OOVs. The TNORM-S
contains slightly a lower number of them in
comparison to the non-normalised version,
which indicates that the use of spell-checkers
and the substituion of lexical variants in
not enough to deal with OOVs. This is
re ected in the gures on the perplexity of
the datasets. The perplexity is high across
all the datasets, and it slightly decreases
after removing the hashtags from the data,
indicating that the language used in the text
is notable di erent from the LM. This can
be ascribed to the fact that the LM was
build using an out-of-domain corpus. In
turn, removing the hashtags from the data
decreases the amount of OOVs, and seems
to have an impact only in the formal dataset,
where half of the OOVs occur in the hashtags.
However, their proportion is smaller when
compared with the colloquial datasets.</p>
      <p>For the sake of comparison, the same
calculation was carried on using a LM trained
on TNORM corpus, the only corpus publicly
# OOV - clean data
# OOV - no hashtags
ppl - clean data
ppl - no hashtags
11:08%
10:30%
591
591
11:26%
7:51%
735
669
available out of the two colloquial ones. The
new LM was used to obtain the % of OOVs
and perplexity estimations on CAES.es and
TSM datasets. The results are shown in
Table 6. The % of OOVs is higher in both
cases, most probably due the small size of the
corpus. However, the perplexity of the TSM
dataset has decreased. This seems to indicate
that the LM was able to capture a high
proportion of the particular characteristics
of colloquial tweets, and that these may be
recurrent in the colloquial genre and do not
appear in formal texts.
6</p>
      <p>Conclusions and Further Work
Twitter has its own particularities that
makes it a hard genre to deal with. This
work reviews recent publications that address
the problem of Twitter translation. The
number of works in this eld is still scarce
due the lack of corpora, but also because
of the lack of a gold standard and speci c
evaluation methodologies that can help to
assess the quality of a tweet translation.
This work also discusses the most well-know
issues in the Twitter genre: the use of
hashtags and the amount of OOVs, with
especial focus on comparing the di erences
between formal and colloquial texts. The
results obtained are preliminary, but they
clearly show that these two registers are
di erent not only from a linguistic point of
view, but also in terms of tweet structure
and content. Further work has to be done
to align the hashtags and the OOVs in
bitexts corpora and analyse the way their
are translated. Also, the annotation layers
of the TSM corpus enables the possibility
to ne-grain the study, for instance, by
focusing in the di erences between tweets
with di erent communication functions. To
conclude, no major di erences were found
between languages, but this may be ascribed
to the fact that the datasets were obtained
from bitexts corpora.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alegria</surname>
            , Iaki, Nora Aranberri, Cristina Espaa-Bonet, Pablo Gamallo, Hugo G. Oliveira, Eva Mart nez, Iaki San Vicente, Antonio Toral, and
            <given-names>Arkaitz</given-names>
          </string-name>
          <string-name>
            <surname>Zubiaga</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Overview of TweetMT: A Shared Task on Machine Translation of Tweets at SEPLN 2015</article-title>
          .
          <source>In Proceedings of the Tweet Translation Workshop co-located with 31th Conference of the Spanish Society for Natural Language Processing</source>
          , Alacant, Spain, September.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Alegria</surname>
          </string-name>
          , In~aki and Aranberri, Nora and Comas,
          <string-name>
            <surname>Pere</surname>
            <given-names>R</given-names>
          </string-name>
          and
          <string-name>
            <surname>Fresno</surname>
          </string-name>
          , V ctor and Gamallo, Pablo and Padro, Lluis and San Vicente, In~aki and Turmo, Jordi and Zubiaga, Arkaitz.
          <year>2014</year>
          .
          <article-title>TweetNorm es Corpus: an Annotated Corpus for Spanish Microtext Normalization</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bertoldi</surname>
          </string-name>
          , Nicola and Cettolo, Mauro and Federico, Marcello.
          <year>2010</year>
          .
          <article-title>Statistical Machine Translation of Texts with Misspelled Words</article-title>
          .
          <source>In Proceedings of the 2010 Annual Conference of the North American Chapter of the ACL</source>
          , pages
          <volume>412</volume>
          {
          <fpage>419</fpage>
          . ACL.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Brendan O'Connor</surname>
            and
            <given-names>Michel</given-names>
          </string-name>
          <string-name>
            <surname>Krieger</surname>
            and
            <given-names>David</given-names>
          </string-name>
          <string-name>
            <surname>Ahn</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>TweetMotif: Exploratory Search and Topic Summarization for Twitter</article-title>
          .
          <source>In Proceedings of the International Conference on Web and Social Media (ICWSM)</source>
          .
          <source>The AAAI Press.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Gotti</surname>
          </string-name>
          and
          <string-name>
            <given-names>Phillippe</given-names>
            <surname>Langlais</surname>
          </string-name>
          and
          <string-name>
            <given-names>Atefeh</given-names>
            <surname>Farzindar</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Hashtag Occurrences, Layout and Translation: A Corpus-driven Analysis of Tweets Published by the Canadian Government</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , Reykjavik, Iceland, may. ELRA.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Formiga</surname>
          </string-name>
          ,
          <article-title>Llu s and Fonollosa, Jose</article-title>
          <string-name>
            <surname>A. R.</surname>
          </string-name>
          <year>2012</year>
          .
          <article-title>Dealing with Input Noise in Statistical Machine Translation</article-title>
          .
          <source>In Proceedings of COLING 2012: Posters</source>
          , pages
          <volume>319</volume>
          {
          <fpage>328</fpage>
          , Mumbai, India, December.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Gotti</surname>
          </string-name>
          , Fabrizio and Langlais, Philippe and Farzindar, Atefeh.
          <year>2013</year>
          .
          <article-title>Translating Government Agencies' Tweet Feeds: Speci cities, Problems and (a few) Solutions</article-title>
          .
          <source>In Proceedings of the Workshop on Language Analysis in Social Media</source>
          , pages
          <volume>80</volume>
          {
          <fpage>89</fpage>
          ,
          <string-name>
            <surname>Atlanta</surname>
          </string-name>
          , Georgia, June.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Jehl</surname>
          </string-name>
          , Laura.
          <year>2010</year>
          .
          <article-title>Machine Translation for Twitter</article-title>
          .
          <source>Master's thesis</source>
          , University of Edimburgh, United Kingdom.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Jehl</surname>
          </string-name>
          , Laura and Hieber, Felix and Riezler, Stefan.
          <year>2012</year>
          .
          <article-title>Twitter Translation Using Translation-based Cross-lingual Retrieval</article-title>
          .
          <source>In Proceedings of the Seventh Workshop on Statistical Machine Translation, WMT '12</source>
          , pages
          <fpage>410</fpage>
          {
          <fpage>421</fpage>
          ,
          <string-name>
            <surname>Stroudsburg</surname>
          </string-name>
          , PA, USA. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , Max and Kalita, Jugal.
          <year>2010</year>
          .
          <article-title>Syntactic normalization of Twitter messages</article-title>
          .
          <source>In Proceedings of the International Conference on Natural Language Processing</source>
          , Kharagpur, India.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ling</surname>
          </string-name>
          , Wang and Marujo, Luis and Dyer, Chris and Black, Alan W and Trancoso, Isabel.
          <year>2014</year>
          .
          <article-title>Crowdsourcing High-Quality Parallel Data Extraction from Twitter</article-title>
          .
          <source>In Proceedings of the Ninth Workshop on Statistical Machine Translation</source>
          , pages
          <volume>426</volume>
          {
          <fpage>436</fpage>
          ,
          <string-name>
            <surname>Baltimore</surname>
          </string-name>
          , Maryland, USA, June. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Roser</given-names>
            <surname>Saur</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Corpus de Dominio Generico y Espec cos (Ingles, Espan~ol, Catalan y Portugues)</article-title>
          .
          <source>Technical report</source>
          , Social Media.
          <article-title>Metodos y Tecnolog as para los medios sociales</article-title>
          .
          <source>Programa CENIT 2010 (CEN-20101037).</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Subert</surname>
            , Eduard and
            <given-names>Ondrej</given-names>
          </string-name>
          <string-name>
            <surname>Bojar</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Twitter crowd translation { design and objectives</article-title>
          .
          <source>In Translating and the Computer</source>
          <volume>36</volume>
          , pages
          <fpage>217</fpage>
          {
          <fpage>227</fpage>
          ,
          <string-name>
            <surname>Geneva</surname>
          </string-name>
          , Switzerland. AsLing, The International Association for Advancement in Language Technology, Editions Tradulex; AsLing.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>