<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Expedition into the Swiss Twitter Corpus - Workshop Results at SwissText 2018</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ralf Grubenmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William Fallouh</string-name>
          <email>william.fallouh@isen.yncrea.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoforos Nalmpantis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Cieliebak</string-name>
          <email>ciel@zhaw.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>In: Mark Cieliebak, Don Tuggener and Fernando Benites (eds.): Proceedings of the 3rd Swiss Text Analytics Conference (Swiss- Text 2018)</institution>
          ,
          <addr-line>Winterthur</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Zurich University of Applied Sciences</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data Expeditions are short, collaborative events focusing on finding interesting patterns and insights in a dataset through interdisciplinary teams. This is a report of the Data Expedition into the Swiss Twitter Corpus expedition hosted by SpinningBytes AG at the SwissText 2018 conference. The aim was to research interesting topics related to Switzerland in the Swiss Twitter Corpus1. Two teams with a total of 11 participants were given 140'521 Switzerland-related Tweets with relevant metadata and analyzed topics of their choice during the 4 hour workshop. We explain how the data expedition was organized and discuss some of the results and lessons learned.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        More and more data is available to industry and to
researchers, which leads to more and more new
avenues of research being available all the time. For this
purpose, Data expeditions are a popular tool for
educational and research purposes that can quickly
produce interesting analyses from a dataset
        <xref ref-type="bibr" rid="ref1 ref1 ref2 ref4">(Radchenko
and Sakoyan, 2016; Ciociola and Reggi, 2015; Burov
et al., 2016)</xref>
        . They allow groups of people to quickly
try out new ideas and test hypotheses, leading to new
results that might not be found in a more traditional
research setting.
      </p>
      <p>Since visitors to the SwissText conference come
from a wide spectrum of industrial as well as research
backgrounds, we decided to go with this format to
give participants the possibility of hands-on
experience, giving them an opportunity to exchange ideas
with people outside their field and possibly
discovering new topics for future research. Since the focus
of the SwissText is on Natural Language Processing
in Switzerland, giving the participants Tweets from or
regarding Switzerland was the natural choice. These
Tweets were specifically curated for this workshop, as
detailed in Section 2.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Swiss Twitter Corpus</title>
      <p>The Swiss Twitter Corpus is a collection of over 3
million Tweets related to Switzerland which has been
collected since January 2018. Being related to
Switzerland, or ”Swissness”, is defined as either originating
in Switzerland, being written by an important Swiss
Twitter account or being about one of a number of
hand-curated keywords related to a Swiss topic.
Additionally, we look at the users profile location being
in Switzerland and whether the language of the user is
Swiss-German.</p>
      <p>For the expedition, a subsample of Tweets was
selected by selecting Tweets with Swiss
Geocoordinates, Tweets with at least two keywords present
and Tweets with a Swissness-Score of at least 3 or
more. The Swissness-Score counts how many of
the Swissness-Rules apply to a Tweet, for instance
a Tweet with two relevant keywords and
Geocoordinates in Switzerland would get a Swissness-Score of
3. This results in a sample of 140’521 Tweets that are
highly relevant to Switzerland.</p>
      <p>Each Corpus entry contains the Tweet text, the
name of the user, the date of the Tweet, the
Tweetlanguage2, the users country-code, the latitude and
longitude (if provided), the keywords found,
senti2as provided by Twitter
ment annotation (based on Deriu et al. (2017)) as well
as the Swissness-Score and why the Tweet was
included in the set.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data Expedition</title>
      <p>The Data Expedition followed the following format.
Participants were split into groups of 4-8 people (5
and 6 in our case), with the goal of forming diverse
groups as a mix of developers, researchers, designers
and storytellers. The participants were then handed
the data along with an explanation of the data format
and with an introduction to the Data Expedition. The
participants then had roughly 3.5 hours to decide on
one or more research topics and to analyze and
visualize the data and their results. The teams were then
able to present their findings to each other. A
summary of the findings were presented the following day
to the general audience of the conference.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Team 1</title>
        <p>This section details the results and findings of the two
teams.</p>
        <p>The first team focused on finding interesting patterns
related to Swiss celebrities, most of which were
included in the keyword set already. They looked into
the relative number of Tweets per celebrity to find
the most and least popular ones. The most popular
celebrity in the dataset was Roger Federer, a famous
Swiss Tennis player, and the least popular one was
Christoph Blocher, a Swiss politician. The ranking
can be seen in Figure 1.</p>
        <p>The group then focused on analyzing Tweets about
Roger Federer further. First, they looked at the
language and geographic location of relevant Tweets,
noticing that most Tweets originate in Switzerland,
but that Roger Federer is also a popular topic world
(See Figure 1). A majority of the Tweets was
written in English, followed by German and French, with
almost no Tweets being in Italian.</p>
        <p>To finish their analysis, they looked into common
words occurring together with Roger Federer as well
as Hashtags related to the topic. A lot of the
associations found were to be expected, like ”Wawrinka”,
an important opponent of Federer, though some were
surprising to the participants, like ”Rotterdam”, which
they couldn’t find an explanation for.
(a) Popularity of Swiss celebrities measured by number
of Tweets.
(b) Geographic location of Tweets about Roger Federer.
The second team decided to create a Twitter-based
tourism guide of Switzerland. Specifically, they
wanted to recommend top locations for a visitor to
Switzerland in the months of February to May. They
compared manually curated tourism guides with the
Twitter data. To this end, they performed case studies
for four different Swiss destinations.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Destination # of Tweets</title>
        <p>Geneva 7536
Lausanne 3379
Sion 1969
Zermatt 909</p>
        <p>Verbier 312</p>
        <p>Lake Geneva: Geneva was listed as the most
popular tourist destination in multiple guides. The task
was hampered by the existence of a town called ”Lake
Geneva” in the United States of America, which had
to be filtered out. The team created a ranking of
towns around lake Geneva by popularity (number of
Tweets), as seen in Table 1.</p>
        <p>Brugg: They then analyzed mentions of Brugg, a
relatively small town in Switzerland, due to a
number of participants coming from the FHNW university
situated in Brugg. Due to the small size of Brugg,
only 40 relevant Tweets were found in the dataset and
no conclusion could be drawn. Though participants
did find an amusing, sexually explicit Tweet that they
shared with the other workshop participants.</p>
        <p>Lucerne: Lucerne, another popular tourist
location, was mentioned in 5115 Tweets, with 227
mentioning the nearby Titlis mountain, a local tourist
attraction. The participants didn’t find any
interesting information regarding this town, though they
remarked on the Queen Victoria exhibition taking place
there, which was of interest to the British team
member.</p>
        <p>Lugano: Next, the second team wanted to see if the
mo¯untains around the city of Lugano were mentioned
in the dataset, since those are purported popular tourist
locations. Surprisingly, the mountains were only
mentioned a total of 14 times, even though Lugano itself
was mentioned 2792 times.
% of positive Tweets
61.2
65.9
71.4
72.6
86.7
89.8
90.5
94.4</p>
        <p>To round off their analysis, the team members
looked at the distribution of sentiment annotations for
mentions of Swiss cities (see Table 2). They couldn’t
find any overwhelmingly negative Swiss cities, but
noticed that in general, the Italian and French part of
Switzerland is more happy than the German one.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>We organized and executed a data expedition into
Swiss Twitter data with a group of 11 people. The
participants were very motivated and interested in the
topic at hand and discovered several new and
surprising insights from the data. Even though the total
time available for the analysis was only 3 hours, the
teams quickly settled on a topic to study and produced
the first results. The workshop itself was praised by
several participants and received positive feedback in
general, pointing to data expeditions being a useful
and easily introduced tool in education and research.</p>
      <p>In the future, it might be useful to let participants
chose their role in advance, to ease team formation.
Producing general statistics about the data in advance
and adding scaffolding code for participants to use
might help participants finding a suitable topic and
speed up development, at the risk of biasing
participants towards certain avenues of exploration.</p>
      <p>Overall, the expedition was successful and the
format will likely be repeated by us in the future.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We would like to thank all the participants in the
expedition (In no particular order): Stephen, Khalil,
Nathan, Jacky, Stefan, Ela, Alma Karalic, Christoph
Sess, Michael Sladoje, Alexandru Dimofte, Matthias
Sommer.</p>
      <p>We would also like to thank the organizers of the
SwissText conference for the opportunity to lead this
workshop.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>AV</given-names>
            <surname>Burov</surname>
          </string-name>
          ,
          <source>AV Baranov, and AV Tagaev</source>
          .
          <year>2016</year>
          .
          <article-title>Data expedition as an effective tool of creating a culture of working with open data of the future state and municipal officials</article-title>
          .
          <source>In Proceedings of the International Conference on Electronic Governance and Open Society: Challenges in Eurasia. ACM</source>
          , pages
          <fpage>167</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Ciociola</surname>
          </string-name>
          and
          <string-name>
            <given-names>Luigi</given-names>
            <surname>Reggi</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A scuola di opencoesione: From open data to civic engagement. Open Data as Open Educational Resources page 26</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Jan</given-names>
            <surname>Deriu</surname>
          </string-name>
          , Aurelien Lucchi, Valeria De Luca, Aliaksei Severyn, Simon Mu¨ller, Mark Cieliebak, Thomas Hofmann, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Jaggi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification</article-title>
          . In WWW 2017 - International World Wide Web Conference. Perth, Australia.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Irina</given-names>
            <surname>Radchenko</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Sakoyan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>On some russian educational projects in open data and data journalism</article-title>
          .
          <source>In Open Data for Education</source>
          , Springer, pages
          <fpage>153</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>