<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Type of Wiki Citation Tag Citation Count Distribution(%)
{{cite web}}</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>{{citation needed}}: Filling in Wikipedia's Citation Shaped Holes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kris Jack</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo López-García</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maya Hristakeva</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roman Kern</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Know-Center GmbH - Inffeldgasse 13/6</institution>
          ,
          <addr-line>A-8010 Graz</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mendeley Ltd - 144a Clerkenwell Road</institution>
          ,
          <addr-line>London EC1R 5DF</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <volume>4</volume>
      <issue>796</issue>
      <abstract>
        <p>Wikipedia authors cite external references to support claims made in their articles in order to increase their validity. A large number of claims, however, do not have supporting citations, putting them in question. In this paper, we describe a study in which we attempt to retrieve relevant citations for claims using a variety of information retrieval algorithms. These algorithms are inspired by bibliometric and altmetric insights that exploit readership data from Mendeley's community and rerank results using a Bradfordising approach. The results of the small scale study indicate that both of these approaches can improve upon basic keyword-based search, typically used in digital libraries, in order to return relevant documents for unsupported claims.</p>
      </abstract>
      <kwd-group>
        <kwd>Digital Libraries</kwd>
        <kwd>Bibliometrics</kwd>
        <kwd>Altmetrics</kwd>
        <kwd>Wikipedia</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Wikipedia has shown to be of extreme value not only for the general public
but also in the academic world. Research has shown an increasing number of
scholarly publications citing Wikipedia and that academic institutions are one
of Wikipedia’s major consumers [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Given the crowdsourced nature of Wikipedia
and its associated coordination problems, however, Wikipedia is often sceptically
viewed by experts with respect to information quality [
        <xref ref-type="bibr" rid="ref1 ref2">2, 1</xref>
        ].
      </p>
      <p>Similar to scholarly literature, Wikipedia authors are encouraged to cite
reliable and verifiable1 information sources that back up their claims and therefore
strengthen the information quality of articles. This practice has been widely
adopted, resulting in a rich set of Wikipedia articles with numerous references.</p>
      <p>Experts, however, might still find two serious concerns regarding
information quality when it comes to Wikipedia claims and their associated references.
On the one hand, despite the large number of already existing references, it is
still commonplace for co-authors or readers to encounter unsupported claims
in Wikipedia articles. These claims needing corresponding evidence are marked
with a {{citation needed}} tag2. It is then up to Wikipedia contributors to find
the corresponding references that support such claims.</p>
      <sec id="sec-1-1">
        <title>1 http://en.wikipedia.org/wiki/Wikipedia:Verifiability</title>
      </sec>
      <sec id="sec-1-2">
        <title>2 http://en.wikipedia.org/wiki/Wikipedia:Citation_needed</title>
        <p>In order to increase the number of cited claims in Wikipedia, it would be
useful to develop a tool that can automatically suggest articles that back them
up. Wikipedia contributors could then check if the articles indeed back up the
claims and choose to insert relevant ones into Wikipedia pages.</p>
        <p>
          In this paper, (i) we investigate the distribution of citation sources in Wikipedia
to better understand what authors are currently using as verifiable evidence and
the extent to which citations are missing, and (ii) we explore if findings in the
field of bibliometrics can be exploited in developing a system that can
automatically retrieve articles that back up claims. The second study is particularly
focussed on comparing the well established technique of Bradfordising [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] to
a new technique of biasing retrieval ranking based on signals from digital
communities of researchers. As this study was carried out in Mendeley, the digital
community chosen is that of Mendeley’s social network, whose readership data
has previously shown to be useful in understanding research trends [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ].
2
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Wikipedia Citations</title>
      <p>3 Wikipedia English Official Offline Edition (version 20130805) [Xprt] - http://
academictorrents.com/details/30ac2ef27829b1b5a7d0644097f55f335ca5241b</p>
      <sec id="sec-2-1">
        <title>4 https://code.google.com/p/wikixmlj/</title>
      </sec>
      <sec id="sec-2-2">
        <title>5 http://en.wikipedia.org/wiki/Wikipedia:Citation_templates</title>
        <p>claims that require citations. The number of missing citations in Wikipedia
articles indicates the need for a tool that can help people to retrieve articles that
back up claims.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approaches to Finding Citations</title>
      <p>
        In this study, we were interested in applying some insights from bibliometrics
and altmetrics to inform the design of a tool that can help retrieve articles
that support natural language claims. The behaviour of three algorithms was
investigated. The first is Bradfordising, a technique that has been shown to
improve the ranking of research article results in search engines [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The second is
to bias search results based on how often they are read in Mendeley’s community.
The third is to combine both approaches.
      </p>
      <p>All algorithms were investigated using the popular search engine Lucene6.
The article metadata (e.g. title, authors, year of publication, abstract) of
Mendeley’s 100 million research articles were indexed. For each claim, the text of the
claim plus the title of the Wikipedia page were entered into the search engine
and the results were reanked based on either Bradfordising, readership or a
combination of Bradfording and readership.</p>
      <p>The standard Bradfordising approach was followed, applying it to the first
100 results. That is, the first 100 results were reranked so that the articles from
the most frequent publication venue, appearing in the first 100 results, were
ranked above the articles appearing in less frequent publication venues, from the
first 100 results.</p>
      <p>In order to exploit Mendeley’s readership information, a new query handler
was written in Lucene, that extended the basic keyword-based search with a
weighted boost. The weighted boost is based on a logarithmic function of the
number of readers that an article has, as follows;
score log10(number_of _readers + 1)</p>
      <p>The final score given to each result was based on its original keyword-based
score plus the boosting given by the logarithmic function. As a result, articles
that had more readers should have higher ranked positions.</p>
      <p>Finally, the third algorithm combines both approaches, first applying the
readership bias and then Bradfordising.</p>
      <sec id="sec-3-1">
        <title>6 https://lucene.apache.org/</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>As a case study, a small scale gold standard data set was manually created
by selecting 10 Wikipedia pages with claims citing scholarly articles in them.
A claim was randomly chosen from each Wikipedia page and associated with
the scholarly article that it cited. It is recognised that more than one scholarly
article may support the claim made. This approach, despite this limitation, is
common practice in the information retrieval community, as it provides enough
information to fairly compare different algorithms and can be fully automated
for large scale testing. The task of the algorithms is, given a claim, to retrieve
articles that can be used to back it up. We employed two baseline systems: (i)
Google Scholar and (ii) Mendeley’s catalogue search. Both cover a broad range of
research disciplines and are two of the world’s largest research collection
repositories. These baselines were compared to versions of Mendeley’s search engine
enhanced separately using Bradfordising and readership biases, as described in
the previous section.</p>
      <p>Five of the 10 selected claims are provided as examples (Table 2). These
claims are made using natural language sentences that paraphrase and/or
summarise findings from research articles. The 10 claims cross multiple disciplines
of research, just as Google Scholar and Mendeley’s collections do.</p>
      <p>Based on the claim and the title of the Wikipedia article page, a query
was constructed. The query contained all the words that appeared in the claim
with the Wikipedia article page’s title concatenated to it. This query was used
to evaluate each of the approaches: (i) Google Scholar (Google Sch.), (ii)
Basic Mendeley Keyword Search (Men), (iii) Mendeley + Readers (Men+R), (iv)
Mendeley + Bradfordising (Men+B), and (v) Mendeley + Readers +
Bradfordising (Men+R+B). The first 100 results lists from each tool were gathered and
the position of the cited article in each list was recorded. The closer the article’s
position to the start of the results list, the better the approach performs.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>Mendeley’s basic keyword search outperformed Google Scholar in retrieving
the correct target document (i.e. the document actually cited in the Wikipedia
article) in the first 100 results in 8 cases. Google Scholar failed to retrieve the
target documents in the first 100 results for all queries. When considering only
the top 10 results returned, 2 of the sample queries retrieved the correct target
citations in the top 10 results using Mendeley’s basic keyword search.</p>
      <p>Three algorithms were tested based on bibliometrics and altmetrics insights.
The use of readership counts ranked the target articles higher, on average, than
the use of Bradfordising and the combination of readership and Bradfordising,
resulting in 5, 4, and 3 top 10 hits respectively. When considering the number
of cases in which an algorithm ranked the target document highest, Mendeley
+ Readership provided the highest ranked results in 5 of the cases. Mendeley’s
basic keyword search ranked the target article higher in 2 cases compared to
Mendeley + Readership + Bradfordising’s single case. The results suggest the
combination of the 2 algorithmic enhancements, readership boosting and
Bradfordising, appear to produce worse results than using either of these algorithms
alone.</p>
      <p>There were 2 cases when all algorithms tested failed to retrieve the target
article in the top 100 results. In both of these cases, the claims did not contain
the keywords present in the metadata of the articles.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>
        There is evidence that scholarly articles are increasingly citing Wikipedia. One
study showed that Wikipedia had been cited 3,679 times within a reference data
set taken from the Web of Science (WoS) and Elsevier’s Scopus databases [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Regarding the information quality of Wikipedia articles, in 2007 Nielsen studied
the relationship between a journal citation in Wikipedia and the impact factor
of the journal, and a correlation between them could be observed, especially for
high impact journals [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In 2012 Priem et al. sampled a number of scholarly
articles and found that about 5% were cited by the English Wikipedia [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. These
results suggest that Bradfordising also applies in Wikipedia articles.
      </p>
      <p>
        When it comes to the influence of readership, incorporating the readership
count or popularity into a information retrieval system has been studied by many
research groups. Researchers proposed that the readership count can be seen as
an indicator for the quality of the retrieved articles and the to rerank the results
accordingly. They found that among multiple quality metrics, the popularity
contributed significantly to the improvement of the results [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], in good
agreement with our findings. In our case, using keyword-based search with Mendeley
and readership boosting retrieved the target citation in the top 10 results in 5 out
of the 10 queries ran. Comparisons of download and citation data from Scopus
with readership data from Mendeley have shown a medium to high correlation
between downloads and readership and downloads and citations, while there
is a medium-sized correlation between readership and citations. These results
suggest some difference between the different usage features [
        <xref ref-type="bibr" rid="ref10 ref11">11, 10</xref>
        ].
      </p>
      <p>
        None of the algorithms tested managed to retrieve the target article in their
top 100 results in 2 of the tests. In considering the 2 queries, it appears that they
do not share enough keywords in common with the target article’s metadata.
This points to the need for an alternative representation of articles beyond
metadata and possibly an alternative representation of the query itself. Including the
full text could possibly prove beneficial in such scenarios as suggested by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Furthermore, a deeper linguistic representation such as the semantics revealed
through topic modelling is worth considering.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and Future Work</title>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>The presented work was developed within the EEXCESS project funded by the
European Union Seventh Framework Programme FP7/2007-2013 under grant
agreement number 600601. The Know-Center GmbH is funded within the
Austrian COMET Program Competence Centers for Excellent Technologies of the
Austrian Federal Ministry of Transport, Innovation and Technology, the
Austrian Federal Ministry of Economy, Family and Youth and by the State of Styria.
COMET is managed by the Austrian Research Promotion Agency (FFG).</p>
      <sec id="sec-8-1">
        <title>7 http://eexcess.eu/</title>
      </sec>
      <sec id="sec-8-2">
        <title>8 http://www.europeana.eu/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Kittur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kraut</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          :
          <article-title>Harnessing the wisdom of crowds in wikipedia: Quality through coordination</article-title>
          .
          <source>In: Proceedings of the 2008 ACM Conference on Computer Supported Cooperative Work</source>
          . pp.
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          . CSCW '08,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2008</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/ 1460563.1460572
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Kittur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pendleton</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chi</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          :
          <article-title>He says, she says: Conflict and coordination in wikipedia</article-title>
          .
          <source>In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>
          . pp.
          <fpage>453</fpage>
          -
          <lpage>462</lpage>
          . CHI '07,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2007</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1240624.1240698
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Kraker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Körner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jack</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Granitzer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Harnessing User Library Statistics for Research Evaluation and Knowledge Domain Visualization</article-title>
          , p.
          <fpage>1017</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2012</year>
          ), http://know-center.tugraz.at/download_extern/ papers/user_library_statistics.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Kraker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trattner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jack</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindstaedt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schloegl</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Head Start :
          <article-title>Improving Academic Literature Search with Overview Visualizations based on Readership Statistics (</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Mayr</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Relevance distributions across bradford zones: Can bradfordizing improve search? arXiv preprint</article-title>
          arXiv:
          <volume>1305</volume>
          .0357 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nielsen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Å</surname>
          </string-name>
          .:
          <article-title>Scientific citations in wikipedia</article-title>
          .
          <source>arXiv preprint arXiv:0705.2106</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The visibility of wikipedia in scholarly publications</article-title>
          .
          <source>First Monday</source>
          <volume>16</volume>
          (
          <issue>8</issue>
          ) (
          <year>2011</year>
          ), http://pear.accc.uic.edu/ojs/index.php/fm/article/ view/3492
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Peacock</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peacock</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>How well do structured abstracts reflect the articles they summarize</article-title>
          .
          <source>European Science Editing</source>
          <volume>35</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Priem</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piwowar</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hemminger</surname>
            ,
            <given-names>B.M.:</given-names>
          </string-name>
          <article-title>Altmetrics in the wild: Using social media to explore scholarly impact</article-title>
          .
          <source>arXiv preprint arXiv:1203.4745</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Schloegl</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gorraiz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gumpenberger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jack</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kraker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Are downloads and readership data a substitute for citations? The case of a scholarly journal</article-title>
          . In: LIDA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Schloegl</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gorraiz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gumpendorfer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jack</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kraker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Download vs</article-title>
          .
          <source>Citation vs. Readership Data: The Case of an Information Systems Journal</source>
          (
          <year>2013</year>
          ), http://know-center.tugraz.at/download_extern/papers/ issi2013_schloegletal.pdf
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>White</surname>
          </string-name>
          , H.D.:
          <article-title>'bradfordizing'search output: how it would help online users</article-title>
          .
          <source>Online Information Review</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <fpage>47</fpage>
          -
          <lpage>54</lpage>
          (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gauch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Incorporating quality metrics in centralized/distributed information retrieval on the world wide web</article-title>
          .
          <source>In: ACM SIGIR Conference</source>
          . pp.
          <fpage>288</fpage>
          -
          <lpage>295</lpage>
          . SIGIR '00,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2000</year>
          ), http://doi. acm.
          <source>org/10</source>
          .1145/345508.345602
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>