<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Nordic Tweet Stream: A dynamic real-time monitor corpus of big and rich language data</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Linnaeus University</institution>
          ,
          <addr-line>Universitetsplatsen 1, 35195 Växjö</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Eastern Finland</institution>
          ,
          <addr-line>Agora 235, 80101 Joensuu</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article presents the Nordic Tweet Stream (NTS), a cross-disciplinary corpus project of computer scientists and a group of sociolinguists interested in language variability and in the global spread of English. Our research integrates two types of empirical data: We not only rely on traditional structured corpus data but also use unstructured data sources that are often big and rich in metadata, such as Twitter streams. The NTS downloads tweets and associated metadata from Denmark, Finland, Iceland, Norway and Sweden. We first introduce some technical aspects in creating a dynamic real-time monitor corpus, and the following case study illustrates how the corpus could be used as empirical evidence in sociolinguistic studies focusing on the global spread of English to multilingual settings. The results show that English is the most frequently used language, accounting for almost a third. These results can be used to assess how widespread English use is in the Nordic region and offer a big data perspective that complement previous small-scale studies. The future objectives include annotating the material, making it available for the scholarly community, and expanding the geographic scope of the data stream outside Nordic region.</p>
      </abstract>
      <kwd-group>
        <kwd>Twitter</kwd>
        <kwd>corpus linguistics</kwd>
        <kwd>language choice</kwd>
        <kwd>English as a lingua franca</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper introduces a new real-time monitor text corpus of tweets from the Nordic
countries. This corpus is a result of cross-disciplinary collaboration between computer
scientists and a group of sociolinguists interested in language variability in general and
English as a lingua franca (ELF) in particular. ELF is understood as second language
use outside educational settings
        <xref ref-type="bibr" rid="ref20">(cf. Mauranen et al. 2015)</xref>
        . Our collaboration aims at
better methodological accuracy in collecting new types of ELF data in multilingual
settings, and we show how new sources of real-time data can lead to new insights in
linguistics that are not only related to language learning
        <xref ref-type="bibr" rid="ref4">(Bradley, 2016)</xref>
        and language
assessment
        <xref ref-type="bibr" rid="ref1 ref27">(García Laborda et al., 2016)</xref>
        , but increasingly also to variationist
sociolinguistics and its applications
        <xref ref-type="bibr" rid="ref15">(Laitinen et al. 2017)</xref>
        .
      </p>
      <p>
        As is generally known, big and rich data from social media such as blogs, Facebook
and Twitter have turned the web into a user-generated repository of information in
everincreasing numbers of areas. For instance, Twitter data have been used in social
sciences to study the Arab spring
        <xref ref-type="bibr" rid="ref5">(Campbell, 2011)</xref>
        , to predict political campaigns
        <xref ref-type="bibr" rid="ref11 ref25">(Gayo
Avello et al., 2011; Tumasjan et al., 2010)</xref>
        and to predict stock markets
        <xref ref-type="bibr" rid="ref2">(Bollen et al.,
2011)</xref>
        , and to model the geographic diffusion of new lexis
        <xref ref-type="bibr" rid="ref9">(Eisenstein et al., 2014)</xref>
        .
Recent attempts also include incorporating data from various sources for applied
purposes, such as the modelling of the impact of social networks in purchase intentions
        <xref ref-type="bibr" rid="ref27">(Wang et al., 2016)</xref>
        . Recently in linguistics, there have been various successful attempts
to build both mono-lingual
        <xref ref-type="bibr" rid="ref23">(Scheffler, 2014)</xref>
        , and multilingual
        <xref ref-type="bibr" rid="ref1">(Barbaresi, 2016)</xref>
        text
corpora of tweets. In our field, corpus-based English sociolinguistics, more attention
has been put on social media discussion fora
        <xref ref-type="bibr" rid="ref18">(Mair, 2013)</xref>
        , but tweets have also been
explored
        <xref ref-type="bibr" rid="ref13 ref14 ref27 ref6 ref7">(Knight et al., 2014; Huang et al., 2016; Coats 2016, 2017)</xref>
        . At the same time,
it has been argued that using real-time data streams is still in its infancy
        <xref ref-type="bibr" rid="ref8">(Davies, 2015)</xref>
        .
      </p>
      <p>The Nordic Tweet Stream (NTS) initiative was started in April 2016, and it
downloads tweets in real time from the Nordic region. The overarching goal is to tackle the
role of social media and big language data in the global expansion and diversification
of English. We explore the prospects of using Twitter data as a diagnostic tool in
evaluating the changing role of English in lingua franca contexts and, as opposed to much
of previous research in the field, our objective is to integrate two types of empirical
data. We not only rely on traditional structured corpus data as is commonly done but
also use unstructured data sources that are often big and rich in metadata, such as
Twitter streams.</p>
      <p>
        While we focus on a single geographic setting, the objective is to create a scalable
tool that can be implemented in any location. The restriction to the five Nordic countries
is justified in view of their many similarities. The countries constitute a geographically
restricted region, and the main languages spoken there are largely related (Finno-Ugric
Finnish being the exception to North Germanic Danish, Icelandic, Norwegian and
Swedish), and English has a strong, though largely unofficial role
        <xref ref-type="bibr" rid="ref3 ref5">(see, e.g., Leppänen et
al., 2011; Bolton and Meierkord, 2013)</xref>
        in spite of there being no previous colonial ties
between Britain and the Nordic region. English is the first foreign language taught in
schools from an early age, and it is being used increasingly in research and higher
education, business and the media.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Tapping into Twitter to create the NTS</title>
      <p>Twitter is a microblogging platform that enables users to exchange short messages,
tweets (www.twitter.com). Since its launch in 2006, it has expanded rapidly, and in
November 2017 it was ranked as one of the most popular websites in the world by the
Alexa ranking (http://www.alexa.com/topsites) with an estimated 310 million users
publishing over half a billion tweets each day
(www.internetlivestats.com/twitter-statistics/). Each Twitter user has a unique username (prefixed with @, e.g.,
@ThisIsHarryPotter) that can be used both as a signature and a reference. Each user has a number
of friends (users they follow) and followers (users following them). Users can group
posts together by topic or type by use of hashtags (words or phrases prefixed with a #
sign, e.g., #EurovisionSongContest). To repost a message from another Twitter user
and share it with their own followers, a user can retweet the message.</p>
      <p>In addition to the actual message, each tweet comes with a rich set of metadata (a
selection is illustrated in Table 1), enabling researchers to make use of various metadata
attributes, both user-generated and service-provided ones.
place of residence (country/city/ etc.)
name of place of residence
name of country of residence
[GPS Coordinates]</p>
      <sec id="sec-2-1">
        <title>User-related info</title>
        <p>Name
screen_name
Location
Description
verified*
followers_count*
friends_count*
account_identifier*
tweets_issued*
created_at*
time_zone
lang</p>
      </sec>
      <sec id="sec-2-2">
        <title>Place-related info</title>
        <p>place_type*
place_name*
country_code*
geo_location*</p>
      </sec>
      <sec id="sec-2-3">
        <title>Tweet-specific info</title>
        <p>Date*
Time*
Weekday*
Lang*
Tweet
2016-07-03
00:00:31
Sunday
En
Why does Davos seem to be the only one around Stannis
with his head on right? &lt;HT&gt;#emeliewatchesgot&lt;/HT&gt;
&lt;HT&gt;#got&lt;/HT&gt; &lt;HT&gt;#GameofThrones&lt;/HT&gt;
NB: * Indicates that these pieces of information are automatically generated as opposed
to being user-provided information that can be misleading and inaccurate.</p>
        <p>
          One of the advantages for sociolinguists is the fact that Twitter tries to assign a
geolocation to each tweet (cf. the groundbreaking work done by Huang et al. (2016) for
instance). We are here not referring to the user provided home location which often is
misleading or missing
          <xref ref-type="bibr" rid="ref12">(e.g. “Mars”, see Graham et al. 2013)</xref>
          , we refer to the geolocation
information provided by Twitter. Depending on user’s privacy settings and the
geolocation method used, tweets either have an exact location specified as a pair of latitude
and longitude coordinates or an approximate location specified as a rectangular
bounding box. Alternatively, no location at all is specified. This type of geographic
information (‘device location’) represents the location of the machine or device on which a
user sent a Twitter message. The data are derived either from the user’s device itself
(using the GPS) or by detecting the location of the user’s Internet Protocol (IP) address
(GeoIP). The primary source for locating an IP address is the regional Internet registries
allocating and distributing IP addresses among organizations located in their respective
service regions (e.g. RIPE NCC at www.ripe.net handles the European IP addresses).
Exact coordinates are almost certainly from devices with built-in GPS receivers (e.g.
phones and tablets).
        </p>
        <p>Another reason why Twitter has been tapped into in various scientific projects is that
it comes with an open policy allowing third-party tools or users to retrieve at most a
1% sample of all tweets. This service, the Twitter Streaming API, enables programmers
to connect to the Twitter server and to download tweets in real time. The Streaming
API provides three parameters – keywords, hashtags and geographical boundaries –
which can be used to delimit the scope of tweets to be downloaded. Once the number
of tweets matching the request starts to reach 1% of all available tweets, Twitter will
begin to sample the data returned to the user.</p>
        <p>Indeed, the drawback of using the Streaming API is the 1% limitation and the fact
that Twitter is secretive about the sampling mechanism used. Morstatter et al. (2013)
compare the sample provided by the Streaming API with the expensive Firehose API
allowing access to 100% of all public tweets. Their investigation shows that the sample
provided by the Streaming API was a sufficient random sample when the stream was
filtered using geographical boundaries, and that the sample contained 43.5% of all
tweets when the geographical boundary was a rectangle large enough to enclose the
entire country of Syria. The sample received when filtering on keywords and especially
hashtags was not as good. Their attempts to replicate the actual top-100 lists of most
used hashtags using the Streaming API gave mixed results “indicating that the
Streaming data may not be good for finding the top hashtags.”</p>
        <p>
          The NTS data collection makes uses the free Twitter Streaming API. As our
downloading mechanism we use hbc (https://github.com/twitter/hbc), which is the default
Twitter client when programming is done in Java. To collect tweets, we first specify a
geographic region covering the five Nordic countries. A second filtering is added to
select only the tweets tagged with a Nordic country code (DK, FI, IS, NO or SE). This
second filtering is necessary to exclude tweets from neighboring countries (e.g.,
Germany and Russia) located within the chosen geographic boundary. Hence, NTS uses
the geolocation information in each tweet to identify Nordic tweets and consequently,
Twitter users who do not want to share their location are not included. Previous studies
suggest that the proportion of geolocated tweets is low, between 0.7% and 2.9%
depending on geographic contexts
          <xref ref-type="bibr" rid="ref1">(e.g. Barbaresi, 2016)</xref>
          . In an in-house experiment to
test the accuracy of our data (during a 10-day period in August 2017), we compared the
number of NTS tweets in which the language tag was Swedish (n=53,614) with tweets
taken from another download project that tries to capture all tweets language-tagged as
Swedish independent of geolocation or not (n=1,880,844). This results indicates that
only 2.8% of all tweets, tagged as Swedish, are geolocated to one of the Nordic
countries. It should also be noted that GeoIP based device location can easily be tricked by
using proxy gateways, allowing a user anywhere in world to “appear” to be located at
a certain GeoIP address. The use of proxy gateways to hide the location is probably
most common among user with a malicious intent. Figure 1 visualizes the corpus
creation pipeline.
        </p>
        <p>In order to test the coverage of the Streaming API we set up an automatic tweet
generator publishing one tweet per hour with Sweden as the country code. In 67 days
this generator published 1,608 tweets, 1,606 of which were captured by the NTS. We
have also identified three other users from Finland, Norway and Sweden that publish
tweets at regular intervals. We tracked them for 80 days and found that on average
98.9% of their tweets were captured by NTS. It thus seems likely that this way of
downloading tweets includes a large majority of all geolocated tweets in the region.</p>
        <p>
          As is common in studies that use geolocated tweets, the raw data also include
tweets that are generated by automated bots (i.e. non-personal and
organization-initiated machines), which often skews sampling
          <xref ref-type="bibr" rid="ref13 ref27">(cf. Huang et al. 2016)</xref>
          . In a spin-off
project, we are using machine learning algorithms to recognize suspected bot accounts and
use the method developed by Lundberg et al. (2018). This algorithm recognizes
botgenerated tweets written in English and in Swedish, and we currently expanding the
algorithm to the other main languages in the region. Since writing this article, we have
implemented the bot-filtering algorithm, but the initial results reported here are based
on data that contain English and Swedish bots.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Basic statistics</title>
      <p>As of April 30, 2017, the NTS has downloaded 12,443,696 tweets from 273,648 user
accounts and over 0.7 billion points of metadata. Figure 2 presents the number of tweets
per day in the material.
The average count per day is just below 36,805 tweets per day, and during the time
period visualized here, there were two days during when the downloading system
crashed. Some of the peaks in the frequencies of tweets are connected to events covered
by the media. For instance, four of the five highest spikes in the data overall occurred
on the 10, 12, 14 and 15 of May, and the Eurovision Song Contest, one of the most
watched TV programs every year in the Nordic countries, took place on May 14,
spilling over into May 15. The peak on June 27 is largely due to Iceland unexpectedly
defeating England in the Euro 2016 football tournament. In 47,000+ tweets there were
more than 5,000 occurrences or hashtags with the names of the two countries (e.g.,
Island till kvartsfinal!!!! (‘Iceland to the quarterfinals’ (Swedish)) and I love you
Iceland. (tweeted by a Norwegian)). Immediately after the game, more than ten Nordic
tweets per second were registered that discussed the game. Other days when the activity
was greater include the day after the Brexit vote in Britain in the end of June, and the
day after the presidential elections in the U.S. in November 2016.</p>
      <p>In the near future, we plan on making parts of the data available for academic
purposes via an intuitive search interface, which is possible according to Twitter’s Terms
of Service (Twitter TOS). This work is currently in progress, but it builds on the idea
that the interface should be designed to be suitable for scholars with limited
competences in data mining.</p>
    </sec>
    <sec id="sec-4">
      <title>Case study: Some observations of language choice in the NTS</title>
      <p>
        The empirical part demonstrates the potential uses of this corpus. Our approach is
variationist sociolinguistic, meaning that we see language use as variable, in which a
speaker/writer makes choices between alternative forms that are drawn from the pool
of resources available
        <xref ref-type="bibr" rid="ref24">(Tagliamonte, 2012: 3)</xref>
        . We provide broad cross-country data on
language choice, making use of the automatically generated metadata parameter of a
tweet language in the data (Table 1 above). While we recognize that automated
language identification methods are not entire accurate, the agreement between human
coders and Twitter’s language recognition system is fairly high for languages written
in the Latin alphabet
        <xref ref-type="bibr" rid="ref12">(Graham et al. 2013)</xref>
        .
      </p>
      <p>
        The results provide empirical evidence to two theoretically relevant questions related
to ELF. Firstly, in many accounts ELF is still seen to be restricted to domain-specific
uses such as academia and international business/law
        <xref ref-type="bibr" rid="ref18">(cf. Mair 2013)</xref>
        . Note that such a
restricted view is not always shared among ELF scholars, and Pietikäinen (2017) has
shown that ELF is used in the family settings for instance. Our big social media data
approach enables testing this empirically and estimate the extent to which ELF is used
in the Nordic setting. Secondly, we add a big data perspective that complements
traditional survey studies; we use a sample of over 200,000 informants. This figure can be
contrasted with the samples used in a few previous studies. For instance, the results of
a traditional mail-in survey in Finland in 2007 were based on a stratified sample of
1,495 respondents (Leppänen et al., 2011). Similarly, an exploratory interview study of
the role of English Sweden drew data from 28 respondents
        <xref ref-type="bibr" rid="ref3">(Bolton and Meierkord,
2013)</xref>
        . Naturally, the amount of information extracted through a carefully-designed
survey or an interview study can be extensive, but we wish to highlight the need to
combine methods from both traditional methodologies and studies making use of big and
rich data.
      </p>
      <p>
        Figure 3 shows the language distribution in our data. English is the main language,
and its share is 32.9%. This figure is slightly smaller than the share of English in the
Austrian monitor tweet corpus
        <xref ref-type="bibr" rid="ref1">(Barbaresi, 2016)</xref>
        , in which the share was 42.2%. The
main languages in the region are the next most frequent: Swedish (25.9%), Finnish
(11.5%), Norwegian (5.5%), Danish (4.9%), and Icelandic (1.9%).
      </p>
      <p>Some brief comments are needed about the some of the other classifications. The
category “und(efined)”, which represents about 7%, to a great extent consists of
messages from two different categories: on the one hand promotions of hashtags or URLs,
either to individual friends or to all followers, and on the other instances where there is
too little linguistic material to identify the language (e.g., simply ? or ? wtf ? &lt;URL&gt;
… &lt;/URL&gt;”). The unexpectedly large share of Tagalog (code=tl) stems from laughter
such as hahhah erroneously being coded as tl.</p>
      <p>Apart from English and the main/national languages, the shares of other languages
are small. Most of them are European languages, but also immigrant languages in
Europe are among most frequently used ones, i.e. Arabic, Turkish, Russian, Indonesian,
and Thai.</p>
      <p>The distributions of the languages show both regional similarities and differences.
With regards to similarities, when we divide the data according to the five countries,
English is among the top two languages in every country (Table 2). Its share varies
between the lowest share (26%) in Finland and the highest (46%) in the Danish data.
Table 2 shows the five most frequently used languages. It excludes the tweets with
unidentified language codes.</p>
      <sec id="sec-4-1">
        <title>Rank</title>
        <p>1
2
3
4
5</p>
      </sec>
      <sec id="sec-4-2">
        <title>Denmark</title>
        <p>English
(46%)
Danish
(30%)
Spanish
(2%)
Norwegian
(2%)
Swedish
(2%)</p>
      </sec>
      <sec id="sec-4-3">
        <title>Finland</title>
        <p>Finnish
(55%)
English
(26%)
Estonian
(2%)
Russian
(2%)
Swedish
(1%)</p>
      </sec>
      <sec id="sec-4-4">
        <title>Iceland</title>
        <p>Icelandic
(46%)
English
(36%)
Spanish
(2%)</p>
        <sec id="sec-4-4-1">
          <title>German (1%)</title>
        </sec>
        <sec id="sec-4-4-2">
          <title>French (1%)</title>
        </sec>
        <sec id="sec-4-4-3">
          <title>Spanish (2%)</title>
        </sec>
      </sec>
      <sec id="sec-4-5">
        <title>Norway</title>
        <p>English
(37%)
Norwegian
(31%)</p>
        <sec id="sec-4-5-1">
          <title>Danish (5%)</title>
        </sec>
        <sec id="sec-4-5-2">
          <title>Swedish (2%)</title>
        </sec>
      </sec>
      <sec id="sec-4-6">
        <title>Sweden</title>
        <p>Swedish
(52%)
English
(29%)
Spanish
(1%)
Arabic
(1%)
Turkish
(1%)</p>
        <p>If we add up the proportions of English and the main/national language of each
country in Table 3, the two most frequent languages account for over 80% of the languages
used in Iceland (82%), in Sweden (81%) and Finland (81%). The total shares in
Denmark (76%) and Norway (68%) are substantially lower. A notable fact is that immigrant
languages primarily appear among the most frequent language in Sweden (Arabic and
Turkish) and in Finland (Estonian and Russian). We do not yet know if these
differences are reflections of real variability or whether they have been brought about by
technical factors.</p>
        <p>Figure 4 below shows a pattern over the average day in the Nordic region. It
visualizes the hourly averages of all the material captured by the stream from April 2016 to
April 2017 and the share of English (lang=en) material. As for a measure of variability,
the vertical bars for the standard error of the mean help to estimate how significant the
hourly shifts are.</p>
        <p>
          The figure presents a clear pattern of the distribution of tweets. On the one hand, the
temporal distribution follows the daily patterns of most people. The pattern is highly
similar to the one in the German Twitter snapshot
          <xref ref-type="bibr" rid="ref23">(Scheffler, 2014)</xref>
          , but there are also
noticeable differences. Firstly, as could be expected, people tweet the least in the early
morning, reflected in the dramatic drop in the hourly activity after midnight, and from
then on the frequency increases steadily. The normal office hours see a constant
increase in the activity, and the activity peaks at 9–11 PM from where it decreases. The
high frequencies late in the evening support the finding that tweeting to a great extent
is connected to late evening leisure activities. The main difference between our data
and the data presented in Scheffler (2014) is that the peak in her German material occurs
around 8 or 9 PM.
        </p>
        <p>The proportion of English is the highest when Twitter activity is at its lowest at 5
in the morning, reaching almost 50% of all tweets, and at its lowest when the Twitter
activity is at its highest, dropping just below the 30% mark. The high proportion of
English can partly be explained by the low overall proportion of tweets, since some
tweets that are automatically generated around the clock, such as weather reports, are
produced in English.</p>
        <p>
          Lastly, we are currently testing various options of visualizing the material. Figure
5 below illustrates how StanceXplore
          <xref ref-type="bibr" rid="ref19">(Martins et al. 2017)</xref>
          could be used to visualize
language choice. This tool was originally developed as an interactive tool for modelling
stance taking in social media platforms. Stance and sentiment refer to the ways speakers
position themselves in relation to their own or other people’s beliefs, opinions, and
statements in communicative situations with others.
        </p>
        <p>The tool makes it possible to capture a more fine-grained view of language choice
in specific regional contexts, as is the case with the counties in Sweden here. This
snapshot shows how 114,597 tweets in May 2016 are distributed regionally and
chronologically. The tool visualizes the number of messages in various languages and the user
can select individual languages and see their regional frequencies and the actual
messages keyed in by the user.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This article has introduced a new multilingual Twitter corpus covering five countries
in the Nordic region. The corpus is a real-time monitor corpus that is both big in size
and rich in metadata. The article has presented some of the early observations in the
first months of the streaming process, which started in spring 2016. The objective is to
continue the streaming for at least a year, thus updating the corpus with nearly 37,000
tweets per day. The data collection is taking place on a two-layered model in which we
limit ourselves geotagged tweets in a specified geographic region, and we hope to
expand the method to new regions.</p>
      <p>We are currently working on annotating parts of the material. This work consists of
lemmatizing the tweets and running parts-of-speech tagging to the material. The work
has started from the English and the Finnish material. Another future objective is to
build an intuitive search interface for the material.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barbaresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Collection and indexation of Tweets with a geographical focus</article-title>
          .
          <source>Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ),
          <article-title>May 2016</article-title>
          .
          <source>In: Proceedings of the 4th Workshop on Challenges in the Management of Large Corpora (CMLC)</source>
          ,
          <fpage>24</fpage>
          -
          <lpage>27</lpage>
          . &lt;hal-
          <fpage>01323274</fpage>
          &gt;. (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bollen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Twitter mood predicts the stock market</article-title>
          .
          <source>Journal of Computational Science</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bolton</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Meierkord</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>English in contemporary Sweden: Perceptions, policies, and narrated practices</article-title>
          .
          <source>Journal of Sociolinguistics</source>
          <volume>17</volume>
          (
          <issue>1</issue>
          ),
          <fpage>93</fpage>
          -
          <lpage>117</lpage>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bradley</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The Mobile Language Learner - Use of Technology in Language Learning</article-title>
          .
          <source>Journal of Universal Computer Science</source>
          <volume>21</volume>
          (
          <issue>10</issue>
          ),
          <fpage>1269</fpage>
          -
          <lpage>1282</lpage>
          . (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .3217/jucs-021- 10-1269.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>D. G.</given-names>
          </string-name>
          :
          <article-title>Egypt Unsh@ckled: Using Social Media to @ # :) the System: how 140 Characters Can Remove a Dictator in 18 Days</article-title>
          . Llyfrau Cambria/Cambria Books. (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Coats</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Grammatical feature frequencies of English on Twitter in Finland</article-title>
          . In: Squires,
          <string-name>
            <surname>L.</surname>
          </string-name>
          , English in Computer-mediated Communication: Variation, Representation, and Change,
          <volume>179</volume>
          -
          <fpage>210</fpage>
          . Berlin: De Gruyter. (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Coats</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Gender and lexical type frequencies in Finland Twitter English</article-title>
          . In: Hiltunen,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>McVeigh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Säily</surname>
          </string-name>
          , T. (eds.),
          <article-title>Big and Rich Data in English Corpus Linguistics: Methods and Explorations. (Studies in Variation, Contacts and Change in English 19)</article-title>
          . http://www.helsinki.fi/varieng/series/volumes/19/.
          <source>(Accessed 15 Jan</source>
          <year>2018</year>
          ).
          <article-title>(</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Davies</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Corpora: an introduction</article-title>
          . In: Biber,
          <string-name>
            <given-names>D.</given-names>
            and
            <surname>Reppen</surname>
          </string-name>
          , R. (eds.),
          <source>The Cambridge Handbook of English Corpus Linguistics</source>
          ,
          <fpage>11</fpage>
          -
          <lpage>31</lpage>
          . Cambridge: Cambridge University Press. (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Eisenstein</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A</given-names>
          </string-name>
          and
          <string-name>
            <surname>Xing</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          :
          <year>2014</year>
          .
          <article-title>Diffusion of lexical change in social media</article-title>
          .
          <source>PLoS ONE</source>
          <volume>9</volume>
          (
          <issue>11</issue>
          ). doi:
          <volume>10</volume>
          .1371/journal.pone.0113114
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>García</given-names>
            <surname>Laborda</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          , Magal Royo,
          <string-name>
            <given-names>T.</given-names>
            and
            <surname>Bakieva</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <year>2015</year>
          .
          <article-title>Looking towards the Future of Language Assessment: Usability of Tablet PCs in Language Testing</article-title>
          .
          <source>Journal of Universal Computer Science</source>
          <volume>21</volume>
          (
          <issue>10</issue>
          ),
          <fpage>114</fpage>
          -
          <lpage>123</lpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Gayo</given-names>
            <surname>Avello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Metaxas</surname>
          </string-name>
          , P. T. and
          <string-name>
            <surname>Mustafaraj</surname>
          </string-name>
          , E:.
          <article-title>Limits of electoral predictions using twitter</article-title>
          .
          <source>In: Proceedings of the Fifth International AAAI Conference on Weblogs and Social Media. Association for the Advancement of Artificial Intelligence</source>
          . (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Graham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hale</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gaffney</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Where in the world are you? Geolocation and language identification in twitter</article-title>
          .
          <source>The Professional Geographer</source>
          <volume>66</volume>
          ,
          <fpage>568</fpage>
          -
          <lpage>578</lpage>
          . (
          <year>2013</year>
          ). doi 10.1080/00330124.
          <year>2014</year>
          .907699
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Kasakoff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Grieve</surname>
          </string-name>
          , J.:
          <article-title>Understanding US regional linguistic variation with Twitter data analysis</article-title>
          .
          <source>In: Computers, Environment and Urban Systems</source>
          <volume>59</volume>
          (
          <year>2016</year>
          )
          <fpage>244</fpage>
          -
          <lpage>255</lpage>
          .(
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1016/j.compenvurbsys.
          <year>2015</year>
          .
          <volume>12</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adolphs</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Carter</surname>
          </string-name>
          , R.: CANELC:
          <article-title>Constructing an e-language corpus</article-title>
          .
          <source>Corpora</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <fpage>29</fpage>
          -
          <lpage>56</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Laitinen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lundberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levin</surname>
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lakaw</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Revisiting weak ties: using presentday social media data in variationist studies</article-title>
          . In: Säily,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Palander-Collin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Nurmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            and
            <surname>Auer</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.),
          <source>Exploring Future Paths for Historical Sociolinguistics</source>
          ,
          <fpage>303</fpage>
          -
          <lpage>325</lpage>
          . Amsterdam: John Benjamins. (
          <year>2017</year>
          .).
          <source>doi 10</source>
          .1075/ahs.7.12lai.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Leppänen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pitkänen-Huhta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikula</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kytölä</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Törmäkangas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nissinen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kääntä</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Räisänen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laitinen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koskela</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lähdesmäki</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Jousmäki</surname>
          </string-name>
          , H..
          <article-title>National Survey on the English Language in Finland: Uses, Meanings and Attitudes</article-title>
          . Available at: &lt;http://www.helsinki.fi/varieng/series/volumes/05/&gt; (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lundberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nordqvist</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Matosevic</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>On-the-fly Detection of Autogenerated Tweets, arXix preprint (</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mair</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The World System of Englishes: Accounting for the Transnational Importance of Mobile and Mediated Vernaculars</article-title>
          .
          <source>English World-Wide</source>
          <volume>34</volume>
          ,
          <fpage>253</fpage>
          -
          <lpage>278</lpage>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Martins</surname>
            ,
            <given-names>R. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simaki</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kucher</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paradis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kerren</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>StanceXplore: Visualization for the Interactive Exploration of Stance in Social Media</article-title>
          .
          <source>In: 2nd Workshop on Visualization for the Digital Humanities (VIS4DH'17)</source>
          ,
          <year>October 2017</year>
          , Phoenix, Arizona, USA. (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Mauranen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carey</surname>
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ranta</surname>
          </string-name>
          . E.:
          <article-title>New answers to familiar questions: English as a lingua franca</article-title>
          . In: Biber D. and
          <string-name>
            <surname>Reppen</surname>
            <given-names>R</given-names>
          </string-name>
          . (eds.), Cambridge Handbook of English Corpus Linguistics,
          <fpage>401</fpage>
          -
          <lpage>417</lpage>
          . Cambridge: Cambridge University Press. (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Morstatter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfeffer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>H.</given-names>
            , and
            <surname>Carley. K.M.</surname>
          </string-name>
          <article-title>: Is the sample good enough? Comparing data from Twitter's streaming API with Twitter's firehose</article-title>
          .
          <source>In: Association for the Advancement of Artificial Intelligence International Conference on Weblogs and Social Media</source>
          <volume>7</volume>
          :
          <fpage>400</fpage>
          -
          <lpage>408</lpage>
          . https://www.aaai.org/ocs/index.php/ICWSM/ICWSM13/paper/view/6071. (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pietikäinen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>ELF in social contexts</article-title>
          . In: Jenkins,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Baker</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          , and Dewey M. (eds.)
          <article-title>The Routledge handbook of English as a lingua franca</article-title>
          ,
          <fpage>321</fpage>
          -
          <lpage>332</lpage>
          . Abingdon: Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Scheffler</surname>
          </string-name>
          , T.:
          <article-title>A German Twitter Snapshot</article-title>
          .
          <source>In: Proceedings of LREC</source>
          ,
          <fpage>2284</fpage>
          -
          <lpage>2289</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Tagliamonte</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Variationist Sociolinguistics: Change, Observation, Interpretation. London: Blackwell. (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Tumasjan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sprenger</surname>
            ,
            <given-names>T. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandner</surname>
            ,
            <given-names>P. G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Welpe</surname>
            ,
            <given-names>I. M.:</given-names>
          </string-name>
          <article-title>Predicting elections with twitter: What 140 characters reveal about political sentiment</article-title>
          .
          <source>In: ICWSM 10</source>
          , pp.
          <fpage>178</fpage>
          -
          <lpage>185</lpage>
          . (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Twitter</surname>
            <given-names>TOS</given-names>
          </string-name>
          ,
          <article-title>Twitter's Terms of Service</article-title>
          , https://twitter.com/tos,
          <source>last accessed 31 Sept</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H-W.</given-names>
          </string-name>
          et al.:
          <article-title>Exploring the Impacts of Social Networking on Brand Image and Purchase Intention in Cyberspace</article-title>
          .
          <source>Journal of Universal Computer Science</source>
          <volume>21</volume>
          (
          <issue>11</issue>
          )
          <fpage>1425</fpage>
          -
          <lpage>1438</lpage>
          , (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .3217/jucs-021-11-1425.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>