<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Serendipitous Browsing: Stumbling through Wikipedia</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Claudia Hauff and Geert-Jan Houben Web Information Systems Delft University of Technology Delft</institution>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>While in the early years of the Web, searching for information and keeping in touch used to be the two main reasons for 'going online', today we turn to the Web in many di erent situations, including when we look for entertainment to pass the time or relax. A popular tool to facilitate the users' desire for entertainment is StumbleUpon, which allows users to \stumble" through the Web one (semi-random) page at a time. Interestingly to us, many StumbleUpon users appreciate being served Wikipedia articles, which are informative pieces of text that educate the reader about a particular concept. The leisure activity of stumbling can thus also incorporate a learning experience. Since life-long learning is an important characteristic of knowledge economies, it is crucial to understand the interplay between these two - at rst sight - opposing forces. We hypothesize that a greater understanding of what makes certain Wikipedia articles more attractive to the serendipitously browsing user than others, will enable us to develop adaptations that expose a greater amount of Wikipedia articles to the leisure seeking user.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>In the early years of the Web, searching for information
and keeping in touch used to be the two main reasons for
'going online'. Today, we rely on the Web in increasingly
diverse situations including shopping, consultations and
learning. While these examples are all directed towards a
particular goal the user has, we also turn to the Web at times when
we simply want to be entertained to pass the time or relax.
The possibilities for entertaining yourself on the Web are
manifold, one can play games, listen to music, watch movies
or simply browse through the Web in the hope of nding
entertaining pages. Due to the sheer size of the Web though,
random browsing is not e ective for discovering pages that
may b interesting to the individual user. For this reason,
a number of services have become popular that recommend
web pages to users based on their interests. One popular tool
to facilitate the users' desire for entertainment by
serendipPresented at Searching4Fun workshop at ECIR2012. Copyright c 2012 for
the individual papers by the papers’ authors. Copying permitted only for
private and academic purposes. This volume is published and copyrighted
by its editors.
itous browsing is StumbleUpon1 (SU), which allows users
to \stumble" through the Web one (semi-random) page at
a time. Interestingly to us, many SU users appreciate
being shown Wikipedia2 articles, which are informative pieces
of text that educate the reader about a particular concept.
The leisure activity of stumbling thus can also incorporate
a learning experience, which might contribute to the
development of novel ideas and lead to creative insights. Since
life-long learning is an important characteristic of
knowledge economies, it is crucial to understand the interplay
between these two seemingly opposing forces (entertainment
vs. learning). We hypothesize that a greater understanding
of what makes certain Wikipedia articles more attractive to
the serendipitously browsing user than others, will enable
us to develop adaptations that expose a greater amount of
Wikipedia articles to the leisure seeking user.</p>
      <p>
        In this position paper we make an argument for the
importance of this task. We draw from a number of insights
gained in museum studies [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] where the question of how
learning can be facilitated in leisure settings (the museum
visit) has been investigated for many years. While we do
not consider the SU pages to be similar to museum objects,
we do nd a number of parallels.
      </p>
      <p>A rst experiment on the stumbled Wikipedia pages
revealed that, just as in museums not all objects are equally
attractive to visitors, not all articles are interesting to the
average StumbleUpon user. In fact, only a very small
number of Wikipedia articles gather a large number of views by
SU users, most articles are rarely viewed. While we have no
answer yet to the question of how to automatically classify
articles according to their attractiveness to the
serendipitously browsing user, we have developed a number of
hypotheses which are outlined in Section 3.2.</p>
      <p>If we assume for a moment that we are indeed able to
develop such an approach, a number of application scenarios
can be envisioned:</p>
      <p>A qualitative study of the features that play a role in
to trickling the interest of users who do not have an
information need, will enable Wikipedia contributors
to write their articles in a way that is more accessible
to such users.</p>
      <p>Wikipedia is available in many di erent languages and
such a prediction method would allow us to bootstrap a
recommender like StumbleUpon in di erent languages
by adding an initial set of interesting, high quality
pages before the critical mass of users is reached.
1http://www.stumbleupon.com/
2http://www.wikipedia.org/
Outliers (articles with many 'Likes' but a low
probability of being attractive) can be manually investigated
to reduce spam. Or conversely, undiscovered articles
are obtained and can be injected into the index.
The passages that trigger the surprise or the
attractiveness of an article can be identi ed and highlighted
to the browsing user. This may help to keep those
serendipitously browsing users engaged that initially
only quickly scan the article.</p>
      <p>E-learning applications can also bene t, as articles which
are interesting to the casual reader can be found this
way.</p>
      <p>The rest of the paper is organized as follows: related work
is presented in Section 2, followed by a preliminary analysis
of stumbled Wikipedia pages (Section 3) and the conclusiosn
(Section 4).</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>For this work, we draw inspirations from two areas. On
the one hand we consider research into so-called educational
leisure settings and free-choice learning which is a
multidisciplinary eld that includes aspects from sociology,
psychology and education. On the other hand, our work is also
strongly related to serendipity.</p>
      <p>
        Education leisure settings can be found in a wide range
of institutions including museums [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], national parks, zoos,
science centers [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], etc. As the name suggests, these
institutions serve two purposes: to educate the public as well as
to provide an entertaining experience to the visitors.
Education leisure settings can be characterized by a number of
commonalities with respect to the visitors and their learning
experience [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
        ]: (i) the visitors gain direct experience,
(ii) they decide what and whether at all to learn, (iii) the
learning process is guided by their interests, (iv) learning
is in uenced by the visitors' social interactions and (iv) the
visitors are a highly diverse group, with di erent educational
backgrounds and prior knowledge. Since learning in this
setting is voluntary, the visitors' motivation plays an important
role: why did they come?
      </p>
      <p>
        Serendipity, the act of encountering information nuggets
unexpectedly, has mostly been investigated in the context
of education [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and work-related discoveries after
serendipitious moments. One of the works outside of this realm is [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
where tools were developed to help people reminisce in their
own digital collections. In goal-directed Web search the
potential for serendipitous encounters has also been recently
investigated [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], while [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] o ers an insightful discussion of
serendipity and how it is used, exploited and induced in
computer science.
      </p>
      <p>
        Finally we note that di erent aspects of Wikipedia
articles have also been investigated in the past, though not
from a perspective of serendipitously browsing users. For
instance, in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] it was found that the writing style
distinguishes so-called featured articles in Wikipedia3 from
unfeatured articles. Classifying Wikipedia articles according
to their quality, as de ned by Wikipedia contributors, was
also investigated in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], where network motifs and graph
patterns in the editor-article graph were exploited.
3. STUMBLEUPON
3Featured Wikipedia articles are of particularly high quality
and chosen by Wikipedia editors.
userdiscovery
userbrowsing
      </p>
      <p>
        The usage of StumbleUpon is depicted in Figure 1. A user
\stumbles" pages with a simple click of the 'Stumble!' button
in his browser toolbar. In response, the user is presented
with a random page from the Web, biased according to his
user pro le or his friends' 'Likes'. The simplicity of the
system protects the user from information overload [
        <xref ref-type="bibr" rid="ref4 ref8">8, 4</xref>
        ], a
user has only two choices when faced with a stumbled page:
either to start reading or to continue stumbling. Users can
also contribute pages to the SU index: whenever a SU user
discover a web page that is not yet in the index and that he
likes, he can add it by means of the 'Like' button. Finally, for
each page in the SU index, there is a SU page which contains
meta-data, including the number of users who viewed/liked
the page, the category the user who discovered the page
placed it in and the comments users left about the page.
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Wikipedia Articles in StumbleUpon</title>
      <p>In all experiments we report here, we utilize the English
Wikipedia dump enwiki-20111007 from October 2011. In a
pre-processing step, we selected all Wikipedia articles that
are neither redirects to other articles, nor new articles or
explicit disambiguation pages and have a length of at least
500 characters (to remove stubs). In total, 3; 552; 059
articles remained.</p>
      <p>In order to determine the popularity of Wikipedia
articles in StumbleUpon, we randomly selected half of these
Wikipedia articles and queried the StumbleUpon API for
their number of views by SU users. Since SU is a
recommendation engine, we can safely assume that the highly
viewed pages are also highly popular and liked. We note,
that the number of 'Likes' a page has received is not
accessible through the StumbleUpon API. The information is
accessible though at the SU meta-data page, which we
manually checked for the results reported in Table 1.</p>
      <p>Among the evaluated 1; 776; 029 articles, we found 267; 958
(15:13%) of them to be contained in the SU index. In our
initial investigation, we also considered French and
German Wikipedia which are two of the largest non-English
Wikipedia repositories. However, we only found a very
limited number of their articles in the SU index (in both cases
less than 1%) and thus did not consider them further. Thus,
an application scenario as proposed in the introduction (to
bootstrap a recommender for a new language) is highly
desirable.</p>
      <p>Let us now focus on those articles that were submitted
by Stumblers to the index. Figure 2 shows a scatter plot of
the number of views versus the number of Wikipedia articles
in the index. As can be expected, most articles have very
few views (the median number of views is 10) while a small
number of articles have gathered more than half a million
views.</p>
      <p>To give an impression of the type of articles that have
gathered few or many views, Table 1 contains the ten most
viewed Wikipedia articles in our data set as well as ten
random examples of articles that were viewed one hundred
times. We chose these two settings as they represent two
extremes: on the one hand, articles that were viewed and also
liked by a large number of people and on the other hand
articles, that were shown a number of times but less well
received by the SU users.</p>
      <p>
        It should also be noted that the SU category Bizarre &amp;
Oddities, which dominates the list of the ten most viewed
articles is not as prevalent when considering a larger set of
articles. In fact, the top 100 viewed articles in our data set
belong to 59 di erent SU categories: Bizarre &amp; Oddities occurs
12 times, followed by the Writing category (5 times) and a
number of categories with three occurrences, including Arts,
Science and Linguistics. Only one of the top 100 articles was
a so-called featured article (indicating that previous work on
featured article prediction, e.g. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], might not be applicable
here), while seven were semi-protected articles due to
previous vandalism activities. Notable is also the fact that 12
out of the 100 articles are of the form List of X where X =
falgorithms; legendary creatures; band name etymologiesg to
name three examples.
      </p>
      <p>While for a human reader it is usually not di cult to
quickly judge whether an article is potentially interesting to
him or not, it is a challenge to derive a method that
automatically classi es articles accordingly. What exactly makes one
article more interesting to the general public than another?
In order to get get a rst understanding of what users think
about the most viewed articles and possibly also why they
like them, we analysed the comments that were posted on
the SU info page for each of the ten most viewed Wikipedia
articles. This analysis is very cursory, as compared to the
number of views, very few users actually comment on an
article, as commenting distracts from the 'stumbling'
experience. For example, the article Wrap rage with 0.86 million
views and forty-thousand likes has a 41 comments. In total,
we analysed 479 comments and identi ed four broad
categories:
(A) Comments expressing surprise
\There's a name for this?"
\I'd never heard of this before (go StumbleUpon!).</p>
      <p>Very cool."
(B) Comments expressing admiration, sadness, sorrow, etc.
\That's so sad"
\No one should go through life afraid to take a
walk."
\don't know what to say actually.."
(C) Comments about the usefulness of the knowledge
\Simple, but helpful for designers."
\An exceptional list of colours and their code,
invaluable to graphic designers, webmasters etc."
(D) Comments expressing negative sentiments towards the
article
\Fake."
\Why stumble everyday wikipedia articles?"
3.2</p>
    </sec>
    <sec id="sec-4">
      <title>Working Hypotheses</title>
      <p>Based on the preliminary qualitative insights gained, we
developed three intuitions that we believe will enable us to
predict to what a Wikipedia article is likely to be bene cial
to the average SU user.</p>
      <p>Intuition A. Articles that contain unexpected nuggets of
information can be identi ed by considering how semantically
related the article is to the other articles it contains links to.
For instance, the List of unusual deaths Wikipedia article
has, among others, outgoing links to the following diverse
articles: Common g, Malvasia (wine), Eddystone Lighthouse,
Hawaii, and Chimney. We hypothesize that nding such
seemingly unrelated articles can be used as a measure of the
likelihood of the article being of interest.</p>
      <p>Intuition B. Articles that evoke emotional feelings can be
discovered through a form of sentiment analysis. Although
Wikipedia articles are written in a neutral style, some topics
are bound to evoke emotions and those emotional topics can
be identi ed.</p>
      <p>Intuition C. Articles that contain useful knowledge may be
identi ed indirectly, when considering their Talk pages, the
amount of discussions that are ongoing and the style of the
discussions. Articles about practically useful information
are not likely to be emotionally charged, unlike discussions
for instance about politicians, religious topics, etc.</p>
      <p>We emphasize, that these are hypotheses that need to be
veri ed in future work.
4.</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS</title>
      <p>In this position paper we have proposed to investigate
what makes certain Wikipedia articles interesting to users
who are browsing the Web without a goal in order to pass
the time or relax. Since such articles are education to some
degree, the leisure activity of browsing (stumbling) can thus
also incorporate a learning experience. Since life-long
learning is an important characteristic of knowledge economies,
it is crucial to understand the interplay between these two
Most viewed articles
List of unusual deaths
Flying Spaghetti Monster
Wrap rage
Shigeru Miyamoto
Benjaman Kyle
One red paperclip
List of colors
Do not stand at my grave and weep
Fuel cell
Raymond Robinson (Green Man)</p>
      <p>SU Category
Bizarre/Oddities</p>
      <p>Satire
Bizarre/Oddities</p>
      <p>Video Games
Bizarre/Oddities
Bizarre/Oddities</p>
      <p>Arts
Poetry</p>
      <p>Science
Bizarre/Oddities</p>
      <p>Example articles viewed 100 times
Biblioscape
Edge of chaos
Gottfried Wilhelm Leibniz Prize
Mario Buda
Proto-Indo-European language
Cisco Adler
Biofeedback
Ovipositor
Concealer
Winklepickers
forces. We argue that a greater understanding of features
are indicative of an article's attractiveness to the average
user (stumbler) will enable us to develop adaptations that
expose a greater amount of Wikipedia articles to the leisure
seeking user.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Andre</surname>
          </string-name>
          , m. schraefel, J. Teevan, and
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Dumais</surname>
          </string-name>
          .
          <article-title>Discovery is never by chance: designing for (un)serendipity</article-title>
          . In C&amp;C '09, pages
          <fpage>305</fpage>
          {
          <fpage>314</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Andre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Dumais</surname>
          </string-name>
          .
          <article-title>From x-rays to silly putty via uranus: serendipity and its role in web search</article-title>
          .
          <source>In CHI '09</source>
          , pages
          <year>2033</year>
          {
          <year>2036</year>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bjo</surname>
          </string-name>
          <article-title>rneborn. Design dimensions enabling divergent behaviour across physical, digital, and social library interfaces</article-title>
          .
          <source>In Persuasive Technology</source>
          , volume
          <volume>6137</volume>
          , pages
          <fpage>143</fpage>
          {
          <fpage>149</fpage>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bollen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Graus</surname>
          </string-name>
          .
          <article-title>Understanding choice overload in recommender systems</article-title>
          .
          <source>In RecSys '10</source>
          , pages
          <fpage>63</fpage>
          {
          <fpage>70</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Falk</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Storksdieck</surname>
          </string-name>
          .
          <article-title>Science learning in a leisure setting</article-title>
          .
          <source>Journal of Research in Science Teaching</source>
          ,
          <volume>47</volume>
          (
          <issue>2</issue>
          ),
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Helmes</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. O'Hara</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Vilar</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Taylor</surname>
          </string-name>
          .
          <article-title>Meerkat and tuba: Design alternatives for randomness, surprise and serendipity in reminiscing</article-title>
          .
          <source>In Human-Computer Interaction - INTERACT</source>
          <year>2011</year>
          , volume
          <volume>6947</volume>
          , pages
          <fpage>376</fpage>
          {
          <fpage>391</fpage>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Lipka</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>Identifying featured articles in wikipedia: writing style matters</article-title>
          .
          <source>In WWW '10</source>
          ,
          <year>2010</year>
          , pages
          <fpage>1147</fpage>
          {
          <fpage>1148</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oulasvirta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Hukkinen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          .
          <article-title>When more is less: the paradox of choice in search engine use</article-title>
          .
          <source>In SIGIR '09</source>
          , pages
          <fpage>516</fpage>
          {
          <fpage>523</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Packer</surname>
          </string-name>
          .
          <article-title>Learning for fun: The unique contribution of educational leisure experiences</article-title>
          .
          <source>Curator: The Museum Journal</source>
          ,
          <volume>49</volume>
          (
          <issue>3</issue>
          ):
          <volume>329</volume>
          {
          <fpage>344</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Packer</surname>
          </string-name>
          .
          <article-title>Beyond learning: Exploring visitors' perceptions of the value and bene ts of museum experiences</article-title>
          .
          <source>Curator: The Museum Journal</source>
          ,
          <volume>51</volume>
          (
          <issue>1</issue>
          ):
          <volume>33</volume>
          {
          <fpage>54</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Packer</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ballantyne</surname>
          </string-name>
          .
          <article-title>Motivational factors and the visitor experience: A comparison of three sites</article-title>
          .
          <source>Curator: The Museum Journal</source>
          ,
          <volume>45</volume>
          (
          <issue>3</issue>
          ):
          <volume>183</volume>
          {
          <fpage>198</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Packer</surname>
          </string-name>
          .
          <article-title>Motivational factors and the experience of learning in educational leisure settings</article-title>
          .
          <source>PhD thesis</source>
          , Queensland University of Technology,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Harrigan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>Characterizing wikipedia pages using edit network motif pro les</article-title>
          .
          <source>In SMUC '11</source>
          , pages
          <fpage>45</fpage>
          {
          <fpage>52</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>