<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HITS Hits Readersourcing: Validating Peer Review Alternatives Using Network Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>References</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Mathematics</institution>
          ,
          <addr-line>Computer Science</addr-line>
          ,
          <institution>and Physics. University of Udine</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Peer review is a well known mechanism exploited within the scholarly publishing process to ensure the quality of scienti c literature. Such a mechanism, despite being well established and reasonable, is not free from problems, and alternative approaches to peer review have been developed. Such approaches exploit the readers of scienti c publications and their opinions, and thus outsource the peer review activity to the scholar community; an example of this approach has been formalized in the Readersourcing model [5]. Our contribution is two-fold: (i) we propose a stochastic validation of the Readersourcing model, and (ii) we employ network analysis techniques to study the bias of the model, and in particular the interactions between readers and papers and their goodness and e ectiveness scores. Our results show that by using network analysis interesting model properties can be derived.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Peer review is an a priori mechanism exploited within the scholarly publishing
process to ensure the quality of scienti c literature; an article written by some
authors undergoes peer review when it is judged and rated by colleagues of the
same degree of competence. Such mechanism, despite being well established and
reasonable, is not free from problems; indeed, it is characterized by various issues
related to the process itself and the malicious behavior of some stakeholders [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In literature one can nd alternative approaches to peer review, which exploit
readers of scienti c publications and their opinions as a \review force", thereby
outsourcing the peer review activity itself to the community of readers. One of
these approaches has been proposed by Mizzaro [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and called Readersourcing,
as a portmanteau for \crowdsourcing" and \readers", and it is based on a model
proposed in a previous work [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Another similar model is TrueReview [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Soprano and Mizzaro [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] describe a general ecosystem called Readersourcing 2.0
which provides an implementation for such models.
      </p>
      <p>The aim of the Readersourcing model is to de ne a way to measure the overall
quality of a published article as well as the reputation of a scholar as a reader
/ assessor; moreover, from these measures it is possible to derive the reputation
of a scholar as an author. In other terms, the main issue to address is how the
numerical judgments given to publications should be aggregated into indexes
of quality and, from these indexes, how to compute indexes of reputation for
the readers and, eventually, indexes of how much an author is able to publish
papers which are positively rated by their readers. Therefore, to each entity (i.e.,
publications, authors, and readers) is assigned one or more scores which measure
how much good (skilled) it is.</p>
      <p>
        Network analysis is a discipline which studies features and properties of
(usually large) networks or graphs. Its algorithms can be can be quite general and,
therefore, applicable to di erent domains. Mizzaro and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] exploit link
analysis techniques such as the HITS algorithm proposed by Kleinberg [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to
address a research question related to the e ectiveness evaluation of Information
Retrieval (IR) systems. The evaluation of IR systems is performed within
different initiatives, such as TREC (Text REtrieval Conference). Before the actual
conference, TREC provides a test collection made of documents and topics (i.e.,
representations of information needs); such a test collection is used as a
benchmark to compare the performance of di erent IR systems. Participants use their
systems to retrieve, and submit to TREC, a list of documents for each topic.
System e ectiveness is then measured by well established metrics like Mean
Average Precision (MAP) and a nal ranking is built. Mizzaro and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
study the interactions between the di culty of topics and the nal rank of IR
systems. In particular, they investigate the correlation between topic ease and
the ability to predict system e ectiveness and they nd that to be e ective, a
system has to perform well on easy topics. Such nding is quite undesirable since
di cult topic are more useful to allow IR to evolve. Roitero et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] extend the
work of Mizzaro and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] by performing a more detailed analysis on
three di erent datasets: they con rm that the original result is valid and general
across datasets; they nd that when only the most e ective IR systems are
considered there is no evidence that the ranking is a ected only by easy topics; and
they prove that such results are robust across di erent e ectiveness metrics.
      </p>
      <p>
        In this paper we take advantage of the methodology proposed by Mizzaro
and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and extended by Roitero et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to address a similar research
question related to the Readersourcing model. More in detail, we intend to study
the interactions between the skill of a reader and the quality of a paper, where
such quantities are computed by Readersourcing models. This paper is structured
as follows. Section 2 details the related work; Section 4 describes the experiments
performed; Section 5 discusses the results. Finally, Section 6 concludes the paper.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        In an attempt to make this paper self contained, in this section we
summarize two major related work areas that we considered to do our analysis.
Section 2.1 summarizes the Readersourcing model proposed by Mizzaro [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], while
Section 2.2 summarizes the methodology proposed by Mizzaro and Robertson
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to investigate the correlation between topic ease and the ability to predict
system e ectiveness within the e ectiveness evaluation of IR systems activity.
2.1
      </p>
      <sec id="sec-2-1">
        <title>The Readersourcing Model</title>
        <p>In the Readersourcing model three entities are identi ed: papers, readers,
and authors. The score of an author is simply de ned as a weighted average of
his or her papers; we do not analyze it in detail in this paper, where we focus on
. . .
p1
r1 g(jr1;p1)
.
.
.
rm g(jrm;p1)
Sp Sp(p1)
p p(p1)
. . .</p>
        <p>pn Sr r
g(jr1;pn) Sr(r1) r(r1)</p>
        <p>... ...
g(jrm;pn) Sr(rm) r(rm)
Sp(pn)
p(pn)
(a) RP matrix with judgments (RPJ) (b) RP matrix with goodness values (RPG)
more readers and papers. A generic reader is asked to give a numerical judgment
to each paper he reads. Such judgments are used to compute a quality score for
each paper. Likewise, each reader is characterized by a score which measures its
skill/reputation. To each judgment is assigned a measure of its goodness with
respect to other judgments given to the same paper. Moreover, to papers and
readers is assigned a steadiness value which a ects the update of the scores; a
high (low) steadiness value leads to faster (slower) change of the score themselves.</p>
        <p>
          Scores are dynamic and they change depending on user behaviour. For
example, if an author with a low score publishes a paper positively rated by readers,
his score increases; if a reader expresses a judgment which is judged as untruthful
and/or biased because \distant" from other judgments (for a given paper) his
score decreases, and so on. Therefore, there is a temporal dimension to consider,
since the internal state of the model evolves as time passes. In the following, we
hypothesize to \freeze" such state at a xed timestamp, where no new judgments
can be expressed and no new papers can be added. Figure 1 shows a
representation of the model as a reader-paper matrix (RP) with m rows (i.e., readers) and
n columns (i.e., papers) which can be represented in two ways. In the former
(RPJ), each cell contains the numerical judgment given by reader r to paper p,
while in the latter (RPG) each cell contains a measure of the goodness of the
numerical judgment given by reader r to paper p. In both representations, each
reader (paper) has a related score and steadiness pair, which are represented
by the Sr and r column (and Sp and p row) vectors. These are computed
according to the formulas de ned by the Readersourcing model [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>The output of a TREC-like initiative can be represented as a system-topic
matrix (ST) with m rows (i.e., systems) and n columns (i.e., topics). Each cell
contains an e ectiveness measure of each system with respect to each topic
according to some metric such as Average Precision (AP). Each row is averaged
to compute Mean Average Precision (MAP), which is a measure of system e
ectiveness with respect to all topics. Each column is averaged to compute Average
Average Precision (AAP) which is a measure of topic ease.</p>
        <p>The ST matrix is then normalized in two ways. Let us call AAP and MAP
the AAP column and the MAP row of ST. In the former normalization, each
AP(si; tj ) value is transformed into a APA(si; tj ) value (Normalized AP
according to AAP) by subtracting AAP to ST. In the latter, each AP(si; tj ) value is
transformed into a APM(si; tj ) value (Normalized AP according to MAP) by
subtracting MAP to ST. The normalized matrices STA and STM are exploited
to study the interactions between topic ease and system e ectiveness. More in
detail, these two matrices can be merged into a single adjacency matrix which
represents a complete weighted bipartite system-topic graph.</p>
        <p>Each link s ! t with weight APM between a system s and a topic t of
systemtopic matrix (ST) represents how much s \thinks" that t is easy (or \un-easy",
i.e., di cult, with APM &lt; 0). Each link s t with weight APA represents how
much t thinks that s is e ective (or \un-e ective", with APA &lt; 0).</p>
        <p>
          Mizzaro and Robertson [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] exploit the complete weighted bipartite graph to
compute hubness and authority values by using an extended version of the HITS
algorithm proposed by Kleinberg [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] which allows to include negative values
for links weights. As explained by Mizzaro and Robertson [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], the authority
value of a topic t of the system-topic matrix (ST) represents its easiness; when
considered for a system s, it represents its e ectiveness. The hubness value of a
topic t represents its ability to recognize e ective systems; when considered for
a system s, it represents its ability to recognize easy topics.
        </p>
        <p>3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>HITS Hits Readersourcing</title>
      <p>
        We intend to study the interactions between reader skill and paper quality
where such quantities are computed by the Readersourcing model proposed by
Mizzaro [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (Section 2.1) by taking advantage of the methodology proposed by
Mizzaro and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (Section 2.2).
      </p>
      <p>The starting point is a slightly di erent version of the RPJ matrix shown in
Figure 1 (left), which is shown in Figure 2 (left). Let us consider the judgment
matrix RPJ*. The only di erence with respect to RPJ is that RPJ* has only one
additional row and column. The former is called MJp and its values are used to
normalize each column of RPJ* (like AAP in the original methodology), while
the latter is called MJr and its values are used to normalize each row of RPJ*
(like MAP). The goodness matrix RPG* is built similarly. This formalization is
useful since it allows to analyze di erent combinations of MJr and MJp (MGr
and MGp) with judgment (goodness) matrices.</p>
      <p>Once the set of MJr and MJp (MGr and MGp) have been computed, the
RPJ* matrix shown in Figure 2 (left) and the RPG* one are normalized in
two ways. In the former normalization, each jri;pj =g(jri;pj ) value is transformed
into a jari;pj =ga(jri;pj ) (Normalized Judgment/Goodness according to MJp) by
subtracting MJp (MGp) to RPJ* (RPG*). In the latter, each jri;pj =g(jri;pj )
value is transformed into a jmri;pj =gm(jri;pj ) (Normalized Judgment/Goodness
according to MJr) value by subtracting MJr (MGr) to RPJ* (RPG*).</p>
      <p>
        The normalized matrices RPJ*A and RPJ*M (RPG*A and RPG*M) are then
used to build a complete weighted bipartite reader-paper graph. Such a graph
represents relationships between readers and papers which depend on the chosen
set of MJp and MJr and it is used to compute hubness and authority values as
done by Mizzaro and Robertson [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
p1
r1 jr1;p1
...
rm jrm;p1
MJp MJp(p1)
...
      </p>
      <p>pn MJr
jr1;pn MJr(r1)</p>
      <p>...
jrm;pn MJr(rm)
MJp(pn)
...
...
rm
p1
.
.</p>
      <p>.
% pn</p>
      <p>T
RPA</p>
      <p>RPM
0
0
0
A VAL</p>
      <p>M VAL 0
(c) r</p>
      <p>p with weight A VAL
T
Fig. 3: (a) Construction of the adjacency matrix. RPA is the transpose of RPA.
(b-c) Relationships between readers and papers of RP matrix with weight
M VAL and A VAL.</p>
      <p>4</p>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>In our experiments we hypothesize a scenario in which there is a publishing
system; authors submit their papers to such a system and readers are able to
rate the papers. We run some stochastic simulation experiments, in which readers
express stochastic judgments on papers according to some prede ned setting, and
we measure the outcome. There are 5,000 readers, 10,000 papers, and 134,000
judgments. We simulate one month of activity. Readers are partitioned into ve
groups GRi of equal size. The members of each group rate a certain amount of
papers, as shown in Table 1 (left). Papers are partitioned into ve groups GPi
of di erent size to simulate the internal state of a publishing system.</p>
      <p>For each reader a sample of papers is picked, whose size depends on his
group. Every paper is simulated by a beta distribution de ned by two
parameters and ; its support is the [0; 1] interval. The beta distribution probability
density function can assume ve shapes which are represented in Figure 4,
depending on the chosen set of and parameters. The beta distribution allows us
to represent ve di erent distributions of user behavior across papers, thus
representing ve di erent kinds of paper. Each of the distribution shapes (shown
(b) r ! p with weight M VAL
4
3
2
1
0
GP1 - flat
GP2 - bell-shaped
GP3 - U-shaped
GP4 - J-shaped
GP5 - skewed-bell</p>
      <p>
        Group % Parameters
GR1
GR2
GR3
GR4
GR5
1 x 2 Weeks
1 x Week
2 x Week
1 x Day
3 x Day
in Figure 4) originates a di erent simulation of the judgment agreement over
the paper: the at distribution (GP1) simulates a completely random judging
behavior; the bell shaped distribution (GP2) simulates a judgment distribution
centered around a data point in the centre of the judgment scale, simulating
a case of high agreement; the U-shaped distribution (GP3) simulates the case
of maximum disagreement, where two-polarized behavior act in the opposite
boundaries of the judgment scale; the J-shaped distribution (GP4), and the
skewed bell distribution (GP5) simulate, as well as the bell-shaped distribution,
the case of high agreement distributed near to the scale boundaries. The usage of
the Beta distribution to capture and mimic di erent level of agreement, as well
as the relationships between agreement and scale boundaries has been formally
discussed in detail by Checco et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The beta distributions for each paper are
generated in the following way: the set of all papers is partitioned into ve groups
GPi (one for each con guration) where each group contains a xed percentage of
papers. To each of these papers an instance of the beta distribution is assigned,
whose and parameters depend on the paper group. Table 1 (right) shows
such paper groups and parameters. Therefore, the stochastic judgment given to
a paper by a reader is generated by sampling a value from the corresponding
beta distribution.
      </p>
      <p>The simulation produces a list of tuples ht; r; p; a; si: at timestamp t reader r
judges paper p written by author a with a score equal to s. Such a list is provided
as input data to an implementation of the Readersourcing model and from its
output the nal RPJ (RPG) matrix (Figure 1) is built. For each RPJ (RPG)
matrix the corresponding RPJ* (RPG*) matrix is built, where each of them is a
judgment (goodness) matrix characterized by a related set of MJp and MJr (MGp
MJp
MGp
Sp
p</p>
      <p>MJp</p>
      <p>MJr
0:11
0:11
0:01</p>
      <p>MGr
0:07
and MGr) values. The RP matrices are then normalized to build adjacency
matrices which are then used to compute hubness and authority values. In the
following section we will discuss the meaning of the resulting relations (i.e., the
links of the complete weighted bipartite graph) and hubness/authority values.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>We now detail the results of our experiments: Section 5.1 focuses on the
measures de ned in the Readersourcing model and analyzes the correlations between
them; Section 5.2 discusses the outcome of HITS applied to our simulations.
5.1</p>
      <sec id="sec-5-1">
        <title>Correlation Between Readersourcing Measures</title>
        <p>The Readersourcing model produces both score and steadiness values for both
readers and papers (i.e., Sr, r, Sp, and p). We also compute: the mean judgment
received by a paper (MJp) the mean goodness of the judgments received by a
paper (MGp), the mean judgment expressed by a reader (MJr), and the mean
goodness of the judgments expressed by a reader (MGr). Table 2 shows the
correlation values for the paper (left) and reader (right) scores, from which we
can draw several remarks.</p>
        <p>Let us focus on the mean judgment of a paper and the paper score (i.e., MJp
and Sp of the left table, highlighted with y), and between the mean goodness of
a reader and the reader score (i.e., MGr and Sr of the right table, highlighted
with z). The rst correlation highlights some potential bias in how we generate
the simulated data: there is lack of variance in the judgments of readers of a given
paper. In other words, for each paper the vast majority of the readers that rated
it present high agreement in their scores. If we look at Table 1 (right) we see
that the beta distributions that induce high agreement between readers (GP2,
GP4, and GP5) represent the 75% of the total scores. We leave for future work
the analysis of a di erent group distribution in the statistical simulation. The
second correlation strengthens, and is a consequence of, the previous remark:
there is a lack of variance in the quality of readers of a given paper; once a
reader expressed a judgment on a paper, all the other readers of the same paper
tend to express judgments of the similar quality.</p>
        <p>Conversely, when looking at the dual scenario (i.e., MGp and Sp of the left
table, highlighted with ?, and MJr and Sr of the right table, highlighted with ),
we see that a sort of dual symmetry is present: neither mean judgment of a
reader nor the mean goodness of a reader are correlated with respectively the
paper and the reader scores. This suggests that: (i) the readers vote using the
whole judgment scale, and (ii) the papers receive judgments that span across
all the judgment scale. This shows that the current simulation setting is able to
cover all the judgment scale.</p>
        <p>The correlations between all other measures are low and not interesting.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>HITS Algorithm and Hubness</title>
        <p>
          In this section we detail the results of the HITS algorithm when run on the
normalized RP matrices, both when considering the judgments (i.e., RPJ*)
and the goodness (i.e., RPG*). As detailed in previous work [
          <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
          ], the most
interesting index we obtain from running HITS is the hubness of readers and
papers; when we consider the judgment matrix, the hubness of a reader measures
its capability to recognize papers that tend to obtain high judgments, while the
hubness of a paper measures the paper capability to recognize reader that tend
to give high judgments (or, in other words, readers that are biased towards
giving high judgments). Symmetrically, when we consider the goodness matrix
the hubness of a paper measures its capability to recognize readers that tend to
give judgments that have a high quality (or, in other words, high quality readers),
while the hubness of a reader measures the reader capability to recognize papers
that tend to receive judgments of a high quality (i.e., papers that tend to be
judged from high quality readers).
        </p>
        <p>We start with the judgments, i.e., the RPJ* matrix. Figure 5 shows some
scatterplots. All the y-axes report the hubness computed by HITS. In the plots
on the left column, the x-axes report the model measures that refer to a paper,
while in the plots on the right column, the x-axes report the model measures
that refer to a reader; thus, in the scatterplots of the left each point is a paper,
while in the scatterplots of the right each point is a reader. Each scatterplot also
shows the respective Pearson's and Kendall's correlations. The meaning of
the correlations of each plot in the gure can be detailed as follows.
(a) The higher the score of a paper, the higher its capability of recognizing
readers that tend to give high judgments. This correlation is expected to
be high due to how the Readersourcing model is formalized; intuitively, if a
score of a paper is high then the paper will be good in recognizing readers
that tend to give high judgments.
(b) Since the correlation is really low, and close to zero, whatever the score of a
reader (high/low, i.e, high-/low-quality reader), he has the same capability
to recognize papers that tend to obtain high (and low) judgments. This is
a good property of the Readersourcing model: a reader can be of a high (or
low) quality independently of whether he expressed judgments on papers
that have an average judgments that is either high or low. In other words, if
a reader expresses a high quality judgment on a paper, his score as a reader
will increase no matter what the judgment score is. Ideally, for a model to
be completely fair, this correlation value should be exactly zero.
(c) The higher the mean judgment of a paper, the higher its capability of
recognizing readers that tend to give high judgments. Also in this case, as for
Figure 5(a), the high correlation value is expected and less interesting.
Nevertheless, since the correlation value it exactly one, it can be also interpreted
0.2
0.3
0.4
1e 12
: p=0.11 (p &lt; .01)
: p=0.06 (p &lt; .01)
1e 12
as a bias in how we generate the data: in fact, this plot shows that the
variance of the judgments expressed by readers on each paper is on average
very low, despite the beta distributions we use to generate the data. This is
con rmed also by analyzing the next plot.
(d) The higher the mean judgment by a reader, the higher his ability to
recognize papers that get high scores. As for previous plot, also in this case the
correlation value is exactly one. While the high correlation of the previous
plot is expected, this one is not. On the contrary, if the mean judgment by
a reader is high, then it should not necessarily mean that the papers that
s/he judged should get on average high scores, since the other readers
judging the same papers could give lower judgments. This is an indication of a
possible bias in how we generate the data. We leave for future work to use
more sophisticated statistical methods to generate the data, such as for
example vine copulas, that would allow to consider both the paper and reader
distributions at the same time.
(e) Since the correlation is really low, whatever the goodness of the judgments
received by a paper (i.e., high or low mean goodness), it has the same
capability to recognize readers that tend to give high judgments. This is a good
property of the Readersourcing model: a paper can be either good or bad
(i.e., have a high or low mean goodness) independently from having been
judged by readers biased towards high or low scores. In other words, the
model formalization of the goodness measure of a paper is robust to the
possible reader bias on the judgment scale.
(f) Since the correlation is really low, whatever the mean goodness of the
judgments expressed by a reader (i.e., high or low mean goodness), he has the
same capability to recognize papers that tend to get high scores. As for the
previous plot, also this correlation is an indication of a good model property:
a reader can be either good or bad (i.e., have a high or low mean goodness)
independently from the fact that he has judged papers that get on average
high or low scores. In other words, the model formalization of the goodness
measure of a reader is robust to the possible behavior of other readers that
express judgments on the same paper.
(g) Since the correlation is zero, whatever the steadiness of a paper, it has the
same capability to recognize readers that tend to give high judgments. This
re ects a good property of the model: the formalization of the steadiness
measure of a paper is robust to the fact that the paper will get scores in the
upper or lower part of the judgment scale.
(h) As for the previous plot, since the correlation is zero, whatever the reader
steadiness, he has the same capability to recognize papers that tend to get
high scores. Symmetrically from what derived from the previous plot, this
hints that the steadiness measure of a reader is robust to the judgment
behavior of the other readers that express judgments on the same paper.</p>
        <p>We now turn to discuss Figure 6 which shows the same plots as in Figure 5
but when running the HITS algorithm on the goodness matrix RPG*.
(a) Due to the low correlation, whether a paper has a high or low score, it has
the same capability to recognize readers that tend to express high quality
judgments (judgments with high goodness). This highlights a good property
of the Readersourcing model: the ability of a paper of recognize good readers
is independent from the quality of the paper itself. In an ideal model, the
correlation value of this plot should be zero.
(b) The higher the score of a reader, the higher its capability to recognize papers
that tend to get high quality judgments. This high correlation highlights a
possible bias in how we generate the simulations: in fact, if a high quality
reader judges a paper, all the other readers that judge the same paper will
tend to be of high quality. As for Figure 5(d), we leave for future work the
use of more sophisticated models for the statistical generation of judgments.
(c) This plot is the same as Figure 6(a); this has a double meaning: the paper
score and mean judgment of a paper are almost perfectly correlated (see
the 0.97 value in Table 2) and, as for Figure 6(a), the ability of a paper of
recognize good readers is independent from the mean judgment of the paper.
(d) Due to the very low correlation, whether a reader has a high or low mean
judgment, it has the same capability of recognize papers that tend to get high
quality judgments. As for the previous plot, this highlights a good property
of the model: the ability of a reader to recognize papers with high quality
judgments is independent from the judgment location of the reader (i.e., it
is independent from the judgment scale).
(e) The higher the mean goodness of the judgments received by a paper, the
higher its capability of recognizing readers that tend to express high quality
judgments. In this case the correlation is exactly one, and this is expected
and derived from how the Readersourcing model is de ned.
(f) The higher the mean goodness of the judgments expressed by a reader, the
higher his capability to recognize papers that tend to get judgments having
high quality. This, as the previous plot, is a natural consequence of how the
Readersourcing model is de ned.
(g) Due to the low correlation, whether a paper has a high or low steadiness,
it has the same capability of recognizing readers that tend to express high
quality judgments.
(h) Due to the correlation close to zero, whether a reader has a high or low
steadiness, he has the same capability of recognizing papers that tend to
get high quality judgments. Also in this case this highlights a good property
of the model: the ability of a reader to recognize papers that receives high
quality judgments is independent from its steadiness value.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Future Work</title>
      <p>We have provided a two-fold contribution: (i) we proposed an experimental
validation of the Readersourcing model carried out through a stochastic
simulation, and (ii) we explored model properties using network analysis techniques.</p>
      <p>
        This paper leaves plenty of space for future work like, for example, the
usage of other stochastic models, and the analysis of other models that propose
alternatives to peer review [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
0.2
0.3
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Checco</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roitero</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maddalena</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demartini</surname>
          </string-name>
          , G.:
          <article-title>Let's agree to disagree: Fixing agreement measures for crowdsourcing</article-title>
          .
          <source>In: 5th HCOMP</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>De</given-names>
            <surname>Alfaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Faella</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>TrueReview: A Platform for Post-Publication Peer Review</article-title>
          .
          <source>CoRR</source>
          (
          <year>2016</year>
          ), http://arxiv.org/abs/1608.07878
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Kleinberg</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>Authoritative sources in a hyperlinked environment</article-title>
          .
          <source>J. ACM</source>
          <volume>46</volume>
          (
          <issue>5</issue>
          ),
          <volume>604</volume>
          {632 (Sep
          <year>1999</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/324133.324140
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Quality control in scholarly publishing: A new proposal</article-title>
          .
          <source>JASIST</source>
          <volume>54</volume>
          (
          <issue>11</issue>
          ),
          <volume>989</volume>
          {
          <fpage>1005</fpage>
          (
          <year>2003</year>
          ), https://doi.org/10.1002/asi.22668
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Readersourcing - A Manifesto</surname>
          </string-name>
          .
          <source>JASIST</source>
          <volume>63</volume>
          (
          <issue>8</issue>
          ),
          <volume>1666</volume>
          {
          <fpage>1672</fpage>
          (
          <year>2012</year>
          ), https://onlinelibrary.wiley.com/doi/abs/10.1002/asi.22668
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>HITS Hits TREC: Exploring IR Evaluation Results with Network Analysis</article-title>
          .
          <source>In: Proceedings of 30th ACM SIGIR</source>
          . pp.
          <volume>479</volume>
          {
          <issue>486</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Roitero</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maddalena</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Do easy topics predict e ectiveness better than di cult topics</article-title>
          ? In: ECIR. pp.
          <volume>605</volume>
          {
          <fpage>611</fpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Soprano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Crowdsourcing peer review: As we may do</article-title>
          . In: Digital Libraries: Supporting Open Science. pp.
          <volume>259</volume>
          {
          <fpage>273</fpage>
          . Springer International Publishing (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>