<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Google Web Searches and Wikipedia Results: a Measurement Study?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vittoria Cozza</string-name>
          <email>vittoria.cozza@poliba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Van Tien Hoang</string-name>
          <email>vantien.hoang@imtlucca.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marinella Petrocchi</string-name>
          <email>m.petrocchi@iit.cnr.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Electrical &amp; DEI, Polytechnic University of Bari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IMT Institute for Advanced Studies</institution>
          ,
          <addr-line>Lucca</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Informatics and Telematics (IIT-CNR)</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>How are users exposed to Wikipedia results, in return to their web searches? Where are such results positioned on the screen? In this study, we experimentally measure the ranking of Wikipedia pages on Google Italia.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        As observed by a recent article of Nature News [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], \Wikipedia is among the
most frequently visited websites in the world and one of the most popular places
to tap into the world's scienti c and medical information". One of the seventh
most visited websites4, the online encyclopaedia is a dominant source of Internet
knowledge. Remarkably, a 2012 study assessed that Wikipedia pages appeared
in 96% of the results all the searches through Google UK [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In his popular
work on Filter Bubbles [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Pariser was one among the rst ones to theorize the
phenomenon according to which users are unknowingly trapped into protective
bubbles, created by search engines and social platforms to automatically lter
contents. Given that users usually focus on the rst few results of their web
searches [
        <xref ref-type="bibr" rid="ref1 ref4">1, 4</xref>
        ], the exclusive privilege of Wikipedia at the very rst positions
could bias the informative content reachable on the Internet.
      </p>
      <p>In this study, we measure the ranking of Wikipedia pages on Google Italia.
To the best of our knowledge, it is the rst study of this kind for the Italian
language. The procedure is as follows. For our web searches, we concentrate
on Italian keywords (and set of keywords). To collect the most popular search
keywords, they have been chosen from Google Trend,from the Google Display
Planner5, and from the Italian trending words on Twitter. Google Trend gives
the most searched terms in a year, as well as trending searches in the past 24
hours6 and trending searches in the recent past7. Google Display Planner is a
? Work partially funded by the Registro.it project MIB (My Information Bubble)
4 http://www.alexa.com/topsites (7th March, 2016)
5 https://support.google.com/adwords/answer/3056115?hl=en
6 https://www.google.com/trends/home/all/IT
7 https://www.google.com/trends/hottrends#pn=p27
tool providing a series of websites linked to speci c categories. It has also been
used, to look for suggested keywords tied to particular categories, like, e.g., Sport
and Vehicles. We have performed all the searches on Google Italia, from Italy. To
avoid personalised results, we have used newly created browser instances and we
have simulated users not logged into Google. The default settings for searches
were the Google default settings. Among the obtained results, we considered
only organic results and not sponsored one (i.e., Advertisements, Google News,
and so on). In the following, we present the experiments and the results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Experiments and results</title>
      <p>Experiment settings. Our reference date is April 7th, 2016. We have collected
the top search terms on Google Italia from 2011 to 2015. Also, we have extracted
the trending news for the reference date and the hot trends stories of the ten
days prior to that date. As a whole, we obtained 1169 unique search terms.
For a wider view, we have further gathered the top Italian trending keywords
from Twitter (updated three times and within three hours, on April 7th, 2016),
leading to 40 unique terms from Twitter.</p>
      <p>
        For the experiments, we have extended AdFisher[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], an automated tool for
information ow experiments, freely available at GitHub8. In our work,
AdFisher runs browser-based experiments that emulate search queries and store
the results. AdFisher interacts with Selenium, a web browser automation tool.
Selenium allows to run a unique instance of Firefox creating a fresh pro le, with
new associated cookies.
      </p>
      <p>
        For each keyword (or keywords set), a new browser instance has been launched
and we have searched such keyword(s). The browser instance has been destroyed
after saving the query results. New browser instances have been used to avoid
the so called carry-over e ects [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which would lead to results for the current
search being in uenced by the previous searches.
      </p>
      <p>Examples of keywords that we have searched for are \Elezioni presidenziali
negli Stati Uniti d'America del 2016" \Credito Valtellinese", \Una lama di luce".
The complete list of keywords (and keywords set) is at https://goo.gl/9KasJc.
Results. Experiments were performed with more than 1,200 keywords,
spanning 33 categories, with an average of 21 keywords per category. In order to
evaluate the search results, we have considered only those keywords for which
the corresponding Wikipedia link appeared in the rst page of the result list
this corresponds to 708 keywords. Searching for those keywords, we found that
Wikipedia pages in the result lists are ranked 1st 41.1% and 2nd 19.9%. Overall,
they appear in the rst ve positions 78.8 times over 100. Figure 1 shows the
occurrence of Wikipedia pages according to their position.</p>
      <p>Searching keywords belonging to \Topic Emergenti" and \Mostre d'Arte"
always yield Wikipedia links as the rst result. Searching keywords belonging to
\Assicurazioni", \O erte di Lavoro" and \Economia e Finanza" yield Wikipedia
links as rst result for less than 10% of such keywords, see Figure 2.
8 https://github.com/tadatitam/info-flow-experiments
Amongst all, \Mostre d'Arte" is a top category: all the keywords belonging
to that category always lead to results with a Wikipedia link located within the
rst three Google Italia results. Examples of keywords are \Picasso - Milano",
\Renoir - Pavia" and \Leonardo - Venaria". \Scienze" and \Benessere" are top
categories too. More than 50% of searches of related keywords lead to Wikipedia
results in the top three positions, see Figure 3. Instead, only 2% of keywords
in the \News" category lead to results in the top three positions, while results
for keywords related to \Politica", \Economia e Finanza", and \Hobbies" are
ranked in the top three positions less than 10% of times (Figure 3). Figure 2 and
Figure 3 show only the categories with the largest number of keywords.</p>
      <p>
        Work in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] performed a similar analysis on Google UK, focusing on
encyclopaedic subjects, like scienti c and natural sciences. Wikipedia scored
extremely well, being its links in the top two result positions. The keywords
belonging to the \Scienze" category (the Italian word for \Science") obtained more
than 60% of Wikipedia links in the top three result positions.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>This preliminary study measures the position of Wikipedia links, resulting from
searching set of keywords over Google Italia. The outcome provides evidence
that Wikipedia is dominant on Google Italia search results ranking: more than
78% of the times, there is one Wikipedia page within the rst ve search results.
A closer look shows that keywords related to \Mostre d'Arte" category always
have an associated Wikipedia page in the top three results, while those related
to \News" are less than 10%. We have focused on quantifying the phenomenon,
without investigating the semantics motivation behind the di erence rankings
for categories. We leave this study as a future work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Edward</given-names>
            <surname>Cutrell</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zhiwei</given-names>
            <surname>Guan</surname>
          </string-name>
          .
          <article-title>What are you looking for?: An eye-tracking study of information usage in web search</article-title>
          .
          <source>In Human Factors in Computing Systems, CHI '07</source>
          , pages
          <fpage>407</fpage>
          {
          <fpage>416</fpage>
          . ACM,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Amit</given-names>
            <surname>Datta</surname>
          </string-name>
          , Michael Carl Tschantz, and
          <string-name>
            <given-names>Anupam</given-names>
            <surname>Datta</surname>
          </string-name>
          .
          <article-title>Automated experiments on ad privacy settings: A tale of opacity, choice, and discrimination</article-title>
          .
          <source>CoRR, abs/1408.6491</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Aniko</given-names>
            <surname>Hannak</surname>
          </string-name>
          , Piotr Sapiezynski, Arash Molavi Kakhki, Balachander Krishnamurthy, David Lazer,
          <string-name>
            <given-names>Alan</given-names>
            <surname>Mislove</surname>
          </string-name>
          , and Christo Wilson.
          <article-title>Measuring personalization of web search</article-title>
          .
          <source>In 22nd World Wide Web, WWW '13</source>
          , pages
          <fpage>527</fpage>
          {
          <fpage>538</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Nadine</given-names>
            <surname>Ho</surname>
          </string-name>
          <article-title>chstotter and Dirk Lewandowski. What users see - structures in search engine results pages</article-title>
          .
          <source>Inf. Sci.</source>
          ,
          <volume>179</volume>
          (
          <issue>12</issue>
          ):
          <volume>1796</volume>
          {
          <year>1812</year>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Richard</given-names>
            <surname>Hodson</surname>
          </string-name>
          .
          <article-title>Wikipedians reach out to academics</article-title>
          .
          <source>Nature News, Sept</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Eli</given-names>
            <surname>Pariser</surname>
          </string-name>
          .
          <article-title>The Filter Bubble: What the Internet Is Hiding from You</article-title>
          . Penguin Group , The,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Sam</given-names>
            <surname>Silverwood-Cope</surname>
          </string-name>
          .
          <article-title>Wikipedia: Page one of Google UK for 99% of searches. pi-datametrics</article-title>
          .
          <source>com Blog</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>