<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CiteTracked: A Longitudinal Dataset of Peer Reviews and Citations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Barbara Plank</string-name>
          <email>bplank@itu.dk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Reinard van Dalen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IT University of Copenhagen</institution>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Rijskuniversiteit Groningen</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Scienti c dissemination is of central importance for the scienti c process. This paper presents CiteTracked, a dataset of peer reviews and citation statistics covering scienti c papers from the machine learning community and spanning six years. We describe and analyze the data collection of over 3,000 published papers, their peer review texts and citation counts, and depict possible usage directions. The dataset aims at fertilizing novel interdisciplinary work between elds such as scientometrics, information retrieval, computational linguistics and natural language processing to study the scienti c publishing process.</p>
      </abstract>
      <kwd-group>
        <kwd>Peer reviews</kwd>
        <kwd>citations</kwd>
        <kwd>NLP for science</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Researchers around the globe continuously contribute invaluable information to
the world's knowledge by publishing their ndings in conferences and journals.
These ndings are subject to the scrutiny of the scienti c review process. Peer
reviewing is an essential component of the scienti c process. Leading conferences
and journals use peer reviewing to decide which manuscripts to include in their
proceedings and journals. The reviewing process is of vital importance, yet the
process itself is often subject to debate. For example, a recent experiment to
examine the consistency of the review process observed that reject/accept
decisions were following a disagreement rate of 26% on a random sample of a tenth
of the papers that went through the review process twice (e.g., [
        <xref ref-type="bibr" rid="ref3 ref4 ref6">6, 3, 4</xref>
        ]).
      </p>
      <p>
        Typically, reviews are accessible only to the authors of a manuscript, and
to the few selected individuals who organize the scienti c venue. Studies on the
qualitative and quantitative properties of peer reviews had been limited. Few
selected top-scienti c venues recently started to make the peer reviews publicly
accessibly. An example is NeurIPS (the conference on Neural Information
Processing Systems, previously named NIPS) and the OpenReview initiative. Such
initiatives contribute to opening up the largely covert process of peer reviewing.
This starts a recent surge of interest in the study of peer reviews, e.g., [
        <xref ref-type="bibr" rid="ref1 ref13">1, 13</xref>
        ].
      </p>
      <p>
        Once a paper is published, citations can be used to estimate the importance
of a paper, as it encodes the implicit judgement of the importance of a paper
by the community. Citation statistics hence can provide a valuable signal to
study the scienti c impact. E orts in understanding such signal typically resort
to modeling citation networks and hence typically refer to the past [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Some
work exists on predicting future impact, e.g., by correlating textual properties
of papers (such as content from its title or abstract) to citation statistics.
      </p>
      <p>In this paper, we present a novel corpus that provides a possible link between
these research strands and enables prediction of scienti c impact and the study
of peer reviews. Our corpus called CiteTracked contains over 3,000 papers and
over 12,000 reviews from the NeurIPS conference spanning the last six years.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>A Dataset of Peer Reviews and Citation Statistics</title>
      <sec id="sec-2-1">
        <title>Peer review collection and meta-data</title>
        <p>The dataset contains 12,260 peer reviews and meta-data for a total of 3,427
papers published in the NeurIPS proceedings (http://papers.nips.cc) from 2013
to 2018. The meta-data includes information such as author names, abstract,
title, a link to the original paper, and for a subset of editions also the paper
presentation type (event type: oral, spotlight or poster) and author feedback.</p>
        <p>An overview of the data set is provided in Table 1. First we notice that
the data follows the general trend of increasing publication volume at computer
science venues. There was an (almost) steady growth in papers, from 360 in 2013
to 679 in 2017, with a big step in 2018 (1,009 papers).3</p>
        <p>For all years except 2015 and 2016, for each paper an average of 3 reviews per
paper was solicited. In 2015, an average of 4 reviews per paper was implemented,
while in 2016 this amounted to 6 reviews per paper. There is no further numerical
scoring publicly available besides the review text, except for a single year (2016,
in which reviewer con dence scores are available as well).4
3 Note that we had to exclude three papers from the 2018 edition due to review pages
which were inaccessible.
4 We would like to note that beyond NeurIPS there are further venues such as ICLR
which make review data publicly available, e.g., on the OpenReview platform.
Col</p>
      </sec>
      <sec id="sec-2-2">
        <title>Collection of citations</title>
        <p>We embarked on a manual e ort to collect citation statistics over time.
CiteTracked is intended to be an on-going dataset collection e ort. The current
release contains citation counts for all papers published up to 2018, i.e., for a
total of all 3,427 scienti c papers citation counts of 5 time spans are available.</p>
        <p>
          The goal is to collect citation counts at di erent time intervals, with at least
one such iteration per year for every paper (collected in a short time span) which
was originally a manual e ort and has now been semi-automatized. The recording
of the citation scores started in April 2016 (citations1). The citation scores were
further recorded in November 2016 (citations2), June 2017 (citations3), March
2018 (citations4), and June 2019 (citations5). Citation counts were collected via
the bibliographic database provided freely by Google Scholar. In contrast to
subscription-based services such as Web of Science of Scopus, Google Scholar
provides higher coverage.5 Citation statistics collected in this way provide an
alternative to within- eld citation networks [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          lecting the reviews from this source and respective citation counts for papers is
currently beyond the scope of the current project. Recent work has started to collect
such peer review data, cf. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], including data from the Association of Computational
Linguistics [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
5 https://libraryguides.helsinki. /metrics/citations
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Analysis and Potential Use Cases</title>
      <p>In this section, we provide a rst quantitative and qualitative analysis of the
corpus. We showcase the potential of using CiteTracked for data-driven analysis
of linking peer reviews to scienti c impact.
3.1</p>
      <sec id="sec-3-1">
        <title>Analysis of most impactful papers and paper categories</title>
        <p>
          In this section, we highlight the papers that received the most citations per year.
We analyze whether the presentation type of a paper is linked to higher impact.
Datasets like CiteTracked or PeerRead [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] can be used in a variety of ways.
The language of peer reviews The analysis of peer reviews can for example
provide a more nuanced understanding of argumentation in the scienti c process.
For instance, [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] quanti ed how reviews recommending an oral presentation
differ from those recommending a poster. A very recent study provides a dataset
of ICLR reviews annotated for argumentation types [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] (Evaluation, Request,
Fact, Reference, or Quote), which we use to train a bilstm-CRF. We analyzed
the predicted argumentation types in reviews of top (and least cited) papers.
As shown in Figure 2, top cited papers get more evaluative reviews (in 3 out of
the 4 years). This analysis could be a starting point to analyze the stance of the
review (and whether it is a potential `advocate' for the paper).
        </p>
        <p>Peer Reviews for Citation Impact Another use case is to study whether reviews
are predictive of citation impact. This is an aspect that, to the best of our
knowledge, has not been studied yet. Therefore, in this paper we study whether
we can successfully predict citation impact from peer reviews, and to what extent
it is complementary to earlier work that relied on aspects of the paper itself.</p>
        <p>In order to predict the impact of the scienti c papers, we discretize
timenormalized citation statistics into low, medium and high impact papers based
on a boxplot and outlier analysis. We use the Upper Outlier Threshold (UOT).
UOT is de ned by adding up the Inter Quartile Range Rule (IQRR) to the third
quartile. Papers with a growth rate above the UOT are therefore de ned as high
impact papers. Papers with a growth rate that is below the UOT, are de ned as
low/medium impact papers, which were further split up into low and medium
impact papers based on an UOT analysis on the subsequent UOT analysis.</p>
        <p>
          We create baselines which consider meta-data from the paper itself, namely
the title and the abstract as done in earlier studies. We train a Support Vector
Machine with commonly used features extracted from the text and titles of
the paper, i.e., word n-grams and character n-grams, including special features
motivated by earlier work such as the average word length of the paper title,
paper title length. In particular, this speci c feature set includes the use of
question marks in paper titles [
          <xref ref-type="bibr" rid="ref10 ref8">8, 10</xref>
          ], the use of colons in paper titles [
          <xref ref-type="bibr" rid="ref10 ref8 ref9">9, 10,
8</xref>
          ], the length of the paper title [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and the number of authors of a paper [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
The dataset (2013-2017) was split into: 60% training, 20% development and 20%
test set. Model performance is reported in F1-score on the nal test set. To put
the results into perspective, we provide a random strati ed baseline as well as
models inspired by prior work, which only use title and abstract as indicators.
We use review texts which may include review summaries (whenever available).
        </p>
        <p>Table 4 shows the results. There are
several take-aways. First of all, paper
information is predictive of scienti c
impact. A classi er that only uses the pa- Table 4. Results of predicting impact
per title is able to achieve an average level of papers (F1 score).
performance of .81 F1-score. This out- low mid high avg
performs the random strati ed baseline title .91 .0 .26 .81
of .71 and con rms earlier ndings. A abstract .92 .0 .39 .83
closer look reveals that the model strug- reviews .93 .09 .49 .85
gles to predict the mid class. It falls all .92 .24 .48 .85
mostly back to the majority class (the
low impact papers). Adding the paper
abstract improves overall performance (from .81 to .83). Secondly and most
importantly, the results show the potential of review texts. Review texts are
predictive of scienti c impact. The performance of a model based on review texts
is higher than using only abstract or title, thereby con rming our hypothesis
that reviews constitute valuable information for scienti c impact prediction. A
model which uses all information (title, abstract and reviews, indicated as `all'
in Table 4) result in an overall similar performance to reviews alone, but it
improves prediction F1-score for the di cult mid class. This investigation shows
the potential of learning from peer review texts.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>This paper introduces CiteTracked, a corpus of peer reviews from the NeurIPS
conference enriched with citation statistics collected over several years. The
current corpus contains 3,427 papers and over 12,000 reviews. We outline corpus
collection, provide an initial analysis and discuss potential use cases that link
work on bibliographic indicators, peer reviews and scienti c publication impact.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We would like to thank Stijn Eikelboom for help with improving the citation
collection process and NVIDIA for supporting our research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ammar</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalvi</surname>
            , B., van Zuylen,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kohlmeier</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications</article-title>
          .
          <source>In Proceedings of NAACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthukrishnan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>The ACL anthology network corpus</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>47</volume>
          (
          <issue>4</issue>
          ),
          <fpage>919</fpage>
          -
          <lpage>944</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Greaves</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hannay</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Nature's trial of open peer review</article-title>
          .
          <source>Nature</source>
          ,
          <volume>444</volume>
          (
          <issue>7122</issue>
          ),
          <fpage>971</fpage>
          -
          <lpage>972</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tabibian</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muandet</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guyon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Von Luxburg</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Design and analysis of the NIPS 2016 review process</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1913</fpage>
          -
          <lpage>1946</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ruocco</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daraio</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Folli</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Leonetti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>Bibliometric indicators: the origin of their lognormal distribution and why they are not a reliable proxy for an individual scholar's talent</article-title>
          .
          <source>Palgrave Communications</source>
          .
          <volume>3</volume>
          :
          <fpage>17064</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Langford</surname>
          </string-name>
          , John, &amp; Mark
          <string-name>
            <surname>Guzdial</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>The arbitrariness of reviews, and advice for school administrators</article-title>
          .
          <source>Communications of the ACM 58</source>
          .4:
          <fpage>12</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Weihs</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Learning to predict citation-based impact measures</article-title>
          .
          <source>In Proceedings of the 17th ACM/IEEE Joint Conference on Digital Libraries</source>
          (pp.
          <fpage>49</fpage>
          -
          <lpage>58</lpage>
          ). IEEE Press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hudson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>An analysis of the titles of papers submitted to the uk ref in 2014: authors, disciplines, and stylistic details</article-title>
          .
          <source>Scientometrics</source>
          <volume>109</volume>
          (
          <issue>2</issue>
          ),
          <volume>871</volume>
          {
          <fpage>889</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jacques</surname>
            ,
            <given-names>T. S. &amp; N. J.</given-names>
          </string-name>
          <string-name>
            <surname>Sebire</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>The impact of article titles on citation hits: an analysis of general and specialist medical journals</article-title>
          .
          <source>JRSM short reports 1(1)</source>
          , 1{
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jamali</surname>
            ,
            <given-names>H. R.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>M. Nikzad</surname>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Article title type and its relation with the number of downloads and citations</article-title>
          .
          <source>Scientometrics</source>
          <volume>88</volume>
          (
          <issue>2</issue>
          ),
          <volume>653</volume>
          {
          <fpage>661</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Subotic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          &amp; B.
          <string-name>
            <surname>Mukherjee</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Short and amusing: The relationship between title characteristics, downloads, and citations in psychology articles</article-title>
          .
          <source>Journal of Information Science</source>
          <volume>40</volume>
          (
          <issue>1</issue>
          ),
          <volume>115</volume>
          {
          <fpage>124</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>E. S.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>J. A. Gomes</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Citations to scienti c articles: Its distribution and dependence on the article features</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>13</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hua</surname>
          </string-name>
          , Xinyu, Mitko Nikolov, Nikhil Badugu &amp; Lu
          <string-name>
            <surname>Wang</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Argument Mining for Understanding Peer Reviews</article-title>
          .
          <source>In Proceedings of NAACL.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>