<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Patricia Gutierrez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nardine Osman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carles Sierra</string-name>
          <email>sierrag@iiia.csic.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence Research Institute (IIIA-CSIC)</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>40</fpage>
      <lpage>46</lpage>
      <abstract>
        <p>Consider an evaluator, or an assessor, who needs to assess a large amount of information. For instance, think of a tutor in a massive open online course with thousands of enrolled students, a senior program committee member in a large peer review process who needs to decide what are the final marks of reviewed papers, or a user in an e-commerce scenario where the user needs to build up its opinion about products evaluated by others. When assessing a large number of objects, sometimes it is simply unfeasible to evaluate them all and often one may need to rely on the opinions of others. In this paper we provide a model that uses peer assessments to generate expected assessments and tune them for a particular assessor. Furthermore, we are able to provide a measure of the uncertainty of our computed assessments and a ranking of the objects that should be assessed next in order to decrease the overall uncertainty of the calculated assessments.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Consider an assessor who needs to assess a large amount of
information. For instance, think of a tutor in a massive open
online course with thousands of enrolled students, a senior
program committee member in a large peer review process
who needs to decide what are the final marks of reviewed
papers, or a user in an e-commerce scenario where the user
needs to build up its opinion about products evaluated by
others. When assessing a large number of objects, sometimes it
is simply unfeasible to evaluate them all and often one may
need to rely on the opinions of others. In the process of
building up our opinion, some questions need to be answered, such
as: How much should I trust the opinion of a peer? What
should I believe given a peer’s opinion? What should I
believe when many peers give their different opinions? Which
objects should be assessed next, such that the certainty of my
belief improves?</p>
      <p>This paper addresses these questions through the
Personalised Automated ASsessment model (PAAS). PAAS uses
peer assessment to calculate and predict assessments.
However, what is fundamentally different from many previous
works [Piech et al., 2013; de Alfaro and Shavlovsky, 2013;</p>
      <p>Walsh, 2014; Wu et al., 2015] is that the computed peer-based
assessment is tuned to the perspective of a specific
community member. PAAS aggregates peer assessments giving more
weight to those peers that are trusted by the specific
community member whom the automated assessments are computed
for. How much this specific member trusts a peer is then
based on the similarity or evaluation rate between his (past)
assessments and the peer’s (past) assessments over the same
assignments. To compute such a trust measure, we build a
trust network conformed of direct and indirect trust values
among community members. Direct trust values are derived
from common assessments while indirect trust is based in the
notion of transitivity. We clarify that our target is not
consensus building, but to accurately estimate unknown assessments
from a specific member’s point of view, based on the peers’
assessments and reliability.</p>
      <p>Finally, we are also able to provide a measure of the
uncertainty of our calculated assessments and a ranking of the
objects that should be assessed next in order to decrease the
overall uncertainty of those calculated assessments.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>The PAAS Model</title>
      <sec id="sec-2-1">
        <title>Notation and Problem Definition</title>
        <p>Let represent an assessor who needs to assess a large set of
objects I, and let P be a set of peers that are able to assess
objects in I.</p>
        <p>We understand assessments as probability distributions
over an evaluation space E at a given moment in time. For
example, one can define a set of elements for the
evaluation space for the quality of an English classroom homework
as E = fpoor; good; excellentg. The assessment fpoor 7!
0; good 7! 0; excellent 7! 1g would represent the
highest assessment possible, whereas the assessment fpoor 7!
0; good 7! 1=2; excellent 7! 1=2g would represent that the
quality of the homework is most probably between good and
excellent, and so on.</p>
        <p>We define an assessment ei (also referred to as evaluation
or opinion) as a probability distribution over the evaluation
space E , where 2 I is the object being evaluated and i 2
f [ Pg is the evaluator. We say ei =fx1 7! v1; : : : ; xn 7!
vng, where fx1; : : : ; xng = E and vi 2 [0; 1] represents the
value assigned to each element xi 2 E , with the condition
.
.</p>
        <p>i2jEj</p>
        <p>Finally, we define L as the history of all assessments
performed, and O L as the set of past peer assessments over
the object .</p>
        <p>The ultimate goal of our work is to compute the probability
distribution of ’s evaluation over a certain object , given the
evaluations of several peers over that same object . In other
words, what is the probability that ’s evaluation is x given
the set of peers’ evaluations O ? Such expectation can be
formalized with the conditional probability as follows:
p(X=x j O )</p>
        <p>To calculate the above conditional probability, we take into
account every particular evaluation in O . In other words,
expectations (or probabilities) are calculated for each
individual evaluation in O , before those expectations are
aggregated into p(X=x j O ). The probability that ’s assessment
is x given a particular evaluation e 2 O is formalized as
follows:</p>
        <p>p(X=x j e )</p>
        <p>The more general probability p(X=x j O ) is then defined
as an aggregation of the individual probabilities:</p>
        <p>p(X=x j O )=p(X=x j e )
where the exact definition of the aggregation is presented later
on in Section 2.4.</p>
        <p>We strongly base the intuition behind the computation of
the individual conditional probabilities on the notion of trust
between peers based on previous experiences, where trust is
understood in this context as the expected similarity between
the assessments given by those peers. In other words, our
intuition is that we expect will tend to agree with ’s
assessments if his trust on is high. Otherwise, ’s evaluation will
probably be different. We perform then a sort of analogical
reasoning: if in the past gave opinions that were a certain
degree dissimilar from ’s opinions, then this will probably
happen again now.</p>
        <p>The remainder of this section is divided accordingly. We
first describe in detail how the measure of trust between peers
is calculated (Section 2.2). Then, we illustrate how to
calculate ’s assessment on an object given ’s assessment
over and ’s trust in ’s assessments (Section 2.3). In other
words, we present an approach for calculating the individual
probability p(X=x j e ). We then illustrate how to combine
those probabilities to build the probability distribution of ’s
assessments given the assessments of several peers (Section
2.4). In other words, we present an approach for calculating
the probability p(X=x j O ). Finally, we provide a measure
of the uncertainty of the computed assessments and a ranking
of the objects that should be assessed next by in order to
decrease that uncertainty (Section 2.5).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Step 1. How much should I trust a peer?</title>
        <p>needs to decide how much can he or she trust the assessment
of a peer . We define this trust measure based on the
following two intuitions. Our first intuition states that if and have
both assessed the same object, then the similarity of their
assessments can give a hint of how close their judgments are.
However, cases may arise where there are simply no objects
evaluated by both and . In such a case, one may think of
simply neglecting ’s assessment, as would not know how
much to trust ’s assessment. Our second intuition, however,
proposes an alternative approach for such cases, where we
approximate that unknown trust between and by looking into
a chain of trust between and through other peers. Roughly
speaking, we relay on the transitive notion: “if trusts , and
trusts 0, then will likely trust 0”. In the following, we
define these two intuitions through two different types of trust
relations: direct trust and indirect trust.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Direct Trust</title>
        <p>Direct trust is the trust relation that emerges between
evaluators that have assessed one or more objects in common. One
possible approach is to measure such relation as aggregations
of their evaluations’ similarity over those objects assessed in
common. For instance, let the set Ai;j = f j ei ; ej 2 Lg
be the set of objects that have been assessed by both
evaluators i and j. Then different definitions for the direct trust
between i and j based on the similarity between two
assessments (sim(ej ; ej )) may be adopted, such as as:
The average of the similarities for all commonly
assessed objects:</p>
        <p>TD(i; j)= 2Ai;j</p>
        <p>X sim(ei ; ej )</p>
        <p>jAi;j j
The conjunction of the similarities for all commonly
assessed objects:</p>
        <p>TD(i; j)= ^</p>
        <p>sim(ei ; ej )
2Ai;j
The Pearson coefficient [Upton and Cook, 2008], or
linear correlation between i and j, for all commonly
assessed objects:
TD(i; j)=s</p>
        <p>X sim(ei ; ei) sim(ej ; ej )
2Ai;j
X sim(ei ; ei)2s
2Ai;j</p>
        <p>X
2Ai;j
sim(ej ; ej )2
where ei, ej are the means of the evaluations performed
over the set Ai;j by i and j respectively.</p>
        <p>However when we calculate such aggregations we loose
relevant information. For instance, we are not able to tell if j
usually under rates with respect to i, if it usually over rates,
or neither. We are also not able to tell if the dissimilarities
between i and j’s evaluations are highly variable or not.</p>
        <p>To cope with such loss of information, we define the direct
trust between two peers i and j as a probability distribution
TDi;j : [0; 1] ! [0; 1] built from the historical data of
previous evaluations performed by i and j. This probability
distribution describes, as we will explain shortly, the expected
similarity or the expected evaluation rate between i and j’s
assessments. The support of the distribution is [0; 1] since
both the expected similarity and the expected evaluation rate
are in the range [0; 1], as we will see shortly, and the range
of the distribution is [0; 1] as this is a probability distribution
and the range of any probability is [0; 1]. Note that we do not
consider here any summarizing measure for trust that would
translate that distribution into a single value, although a
number of measures could be used, such as the average similarity
(as the center of gravity of the distribution) or entropy (as a
measure of the uncertainty of the distribution).</p>
        <p>When defining TDi;j we distinguish two cases: (1) a
first case with a non-ordered evaluation space, such as E =
fvisionary; original; soundg; and (2) a second case with an
ordered evaluation space, such as=fbad; good; excellentg. In
the second case, we are interested in maintaining information
about whether a peer under rates or over rates with respect
to another peer, therefore we are interested in the expected
evaluation rate between i and j. In the first case, this is not
an issue as assessments cannot be ordered and therefore the
notion of under/over rating does not exist, therefore we are
rather interested in the expected similarity between i and j’s
assessments. Next we detail the trust probability distributions
TDi;j built for both cases.</p>
        <sec id="sec-2-3-1">
          <title>Non-Ordered Case.</title>
          <p>In the non-ordered case, we are interested in the
similarity between i and j’s assessments. As such, the support
of the distribution representing i’s direct trust on j (i.e.
the x-axis of TDi;j ) consists of the possible degrees of
similarity between i and j’s assessments.</p>
          <p>Trust distribution TDi;j (x) then describes the probability
that peers i and j evaluate an object with a similarity x
(or the probability that the similarity of their evaluations
is x).</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Ordered Case.</title>
          <p>In the ordered case, we are interested in the evaluation
rate ej=ei between evaluations made by peers i and j.
If ej=ei = 1, this means that i and j provide the same
evaluation. If ej=ei &gt; 1, this meas that j over rates with
respect to i. If ej=ei &lt; 1, this means that j under rates
with respect to i.</p>
          <p>We normalize the evaluation rate to values between 0
and 1. To do so, we require a non decreasing function
r : R ! [0; 1] such that limx!1 r(x)=1, and for
convenience we constraint r(1)=0:5. We adopt the following
normalized evaluation rate function that satisfies these
properties:
r(x)=eln 1=2=x
(1)
As such, the support of the distribution representing i’s
direct trust on j (i.e. the x-axis of TDi;j ) consists of the
possible normalized evaluation rates between i and j.
Trust distribution TDi;j (x) then describes the probability
that i and j would assess an object with a normalized
evaluation rate x.</p>
          <p>In what follows, we explain how we build direct trust
distributions computationally, based on previous experiences.</p>
          <p>
            Initially, the direct trust distribution between any two peers
is the uniform distribution F=f1=n; : : : ; 1=ng (describing
ignorance), where n is the size of the distribution’s support.
Every new assessment made would then update the trust
distributions accordingly. Consider a new assessment ei . The
distribution TDi;j 8j s.t. Ai;j 6= ; is updated as follows:
1. We find the element x in TDi;j ’s support whose
probability needs to be adjusted. So we calculate x=sim(ej ; ei )
in the ordered case
            <xref ref-type="bibr" rid="ref2">(where the definition of sim is
domain dependent and outside the scope of this paper,
although we do note that several approaches may be
adopted, such as using semantic similarity measures [Li
et al., 2003])</xref>
            , or x = r(ej =ei ) in the non-ordered case
(Equation 1).
2. We update the probability of the single expectation x in
TDi;j accordingly:
p(X=x) = p(X=x) +
(1
p(X=x))
(2)
The update is based on increasing the latest probability
p(X = x) by a fraction 2 [0; 1] of the total potential
increase (1 p(X =x)). For instance, if the
probability of x is 0.6 and is 0.1, then the new probability of
x becomes 0:6 + 0:1 (1 0:6) = 0:64. We note that
the ideal value of should be closer to 0 than to 1 so
that one single experience does not result in
considerable changes in the distribution. In other words, a single
assessment cannot result in considerable change in the
probability distribution. Considerable changes can only
be the result of information learned from the
accumulation of many assessments.
3. We normalize TDi;j by updating several expectations
following the entropy based approach of [Sierra and
Debenham, 2005]. The entropy-based approach updates
TDi;j such that: (1) the value p(X=x) is maintained and
(2) the resulting distribution has a minimal relative
entropy with respect to the previous one. In other words,
we look for a distribution that contains the updated
probability value p(X =x) and that is at a minimal distance
from the original TDi;j (as the relative entropy is a
measure of the difference between two probability
distributions). Following this approach, we update TDi;j (X) as
follows:
          </p>
          <p>TDi;j (X) = arg min X p(X=x0) log</p>
          <p>P0(X) x0
such that
fp(X=x) = p0(X=x)g
p(X=x0)
p0(X=x0)
(3)
where p(X =x0) is a probability value in TDi;j , p0(X =
x0) is a probability value in P0, and fp(X=x) = p0(X=
x)g specifies the constraint that needs to be satisfied by
the resulting distribution.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Indirect Trust</title>
        <p>Given a direct trust relation between peers i and j and
between peers j and k, the question now is: What can we say
about the indirect trust between peers i and k when i and k
have no objects assessed in common? In other words, given
the direct trust distributions TDi;j and TDj;k, what can we say
about the indirect trust distribution TIi;k?</p>
        <p>As with direct trust distributions, we distinguish two cases:
a first case where assessments cannot be ordered and thus
trust is based on a similarity measure sim; and a second case
where assessments can be ordered and thus trust is based on
a normalized evaluation rate function r(x)=eln 1=2=x.</p>
        <sec id="sec-2-4-1">
          <title>Non-Ordered Case.</title>
          <p>In this case, we want to preserve the fundamental
triangular inequality property of similarity functions that
says that: T-norm(sim(a; b); sim(b; c)) sim(a; c).</p>
          <p>As with TDi;k, the support (or the x-axis) of TIi;k
consists of the possible degrees of similarity between i and
k’s assessments. But since these degrees of similarity
should satisfy the T-norm, the support is defined as the
set:
supp(TIi;k)=fxik=T-norm(xij ; xjk) j xij 2 supp(TDi;j ) i to m and m to k), we obtain multiple indirect trust
distribu^xjk 2 supp(TDj;k)g tions (one from every chain). In those cases, we pick the
re</p>
          <p>sulting distribution which is most optimistic. In other words,
where supp represents the support of a distribution. while our approach to calculate the indirect trust follows the
We then compute the probabilities of the expectations of pessimistic approach (through our choice of the product
operTIi;k as follows: ator in Equations 4 and 5), we now choose the most optimistic
of the pessimistic outcomes. To do that, we choose the
distrifp(X=xik=T-norm(xij ; xjk))=TDi;j (xij ) TDj;k(xjk) j bution that is closest to the equivalence distribution, which is
xij 2 supp(TDi;j ) ^ xjk 2 supp(TDj(;4k))g a distribution that describes that the evaluations of two peers
This could result in more than one probability computed are equivalent. In the non-ordered case, the equivalence
disfor the same expectation xik. As such, we then add up all tribution is PE(1)=1; that is, the similarity between two peers
the probabilities that correspond to the same expectation is maximum. In the non-ordered case, the equivalence
distribution is PE(0:5) = 1; that is, the normalized evaluation
rate between two peers is 0.5, which implies that they always
provide the same evaluation. The distance between an
indirect trust distribution TIi;k and the equivalence distribution
PE can be calculated as:</p>
          <p>The calculations presented above provide an approach for
calculating indirect trust between two peers i and k when
those peers are linked through a direct trust chain passing
through only one intermediate peer j. For direct trust chains
of increasing length between i and k, the previous process
is iterated. For instance, if there is a direct trust chain
linking i to j, j to m, and m to k, then we first compute the
indirect trust distribution TIi;m from the direct trust
distributions TDi;j and TDj;m, and then we compute the indirect trust
distribution TIi;k from the direct/indirect trust distributions
TIi;m and TDm;k, following the same approach as above.</p>
          <p>When multiple chains of direct trust connect two peers (e.g.</p>
          <p>say a chain linking i to j and j to k, and another chain linking
xik.</p>
          <p>We note that we follow a conservative approach by
adopting the product operator (Equation 4), which is a
T-norm that gives the smallest possible values, as we
prefer not to overrate indirect trust values since they are
not inferred directly from historical data. Of course,
other operators could also be used, such as the min
function.</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>Ordered Case.</title>
          <p>In this case, we want to preserve the property: ej=ei
ek=ej=ek=ei with respect to the evaluations performed by
i, j and k. For instance, if the evaluation rate between
ej and ei is 0.5 (j under rates a 50% with respect to i)
and the evaluation rate between ek and ej is 0.5 (k under
rates a 50 % with respect to j) then the evaluation rate
between ek and ei should be 0.25 (then k under rates a
75 % with respect to i).</p>
          <p>As above, the support (or the x-axis) of TIi;k consists
of the possible degrees of similarity between i and k’s
assessments. The support us then defined as the set:
supp(TIi;k) = fxik=xij
xjk j xij 2 supp(TDi;j )
^xjk 2 supp(TDj;k)g
We then compute the probabilities of the expectations of
TIi;k as follows:
fp(X=xik=xij</p>
          <p>xjk) = TDi;j (xij ) TDj;k(xjk) j
xij 2 supp(TDi;j ) ^ xjk 2 supp(TDj;k)g
(5)</p>
          <p>Again, this could result in more than one probability
computed for the same expectation xik. As such, we
then add up all the probabilities that correspond to the
same expectation xik.
emd(TIi;k; PE)
(6)
where emd is the earth mover’s distance which calculates the
distance between two probability distributions [Rubner et al.,
1998].1 We note that the range of emd is [0,1], where 0
represents the minimum distance and 1 represents the maximum
possible distance.</p>
          <p>In the remainder of this paper, when we refer explicitly to
a direct or indirect trust distribution between peers i and j,
we refer to such distribution as TDi;j or TIi;j , respectively.
Whereas when we refer generically to a trust distribution that
could either be the direct or indirect trust distribution, we
refer to such a distribution as Ti;j .</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>Trust Graph</title>
        <p>Direct and indirect trust relations in a community can be
represented by a weighted directed graph. We define a
community’s trust graph as:</p>
        <p>G=hN; E; wi
1If probability distributions are viewed as piles of dirt, then the
earth mover’s distance measures the minimum cost for transforming
one pile into the other. This cost is equivalent to the ‘amount of dirt’
times the distance by which it is moved, or the distance between
elements of the probability distribution’s support.
where the set of nodes N is the set of evaluators in f [ Pg,
E N N are edges between evaluators with direct or
indirect trust relations, and w : E 7! [0; 1]n is the weight of
an edge, described as a trust probability distribution.</p>
        <p>D E is the set of edges that link evaluators with direct
trust relations: D = f(i; j) 2 E j TDi;j 6= ?g. Similarly,
I E is the set of edges that connect evaluators with indirect
trust relations: I = f(i; j) 2 E j TIi;j 6= ?g n D. We note
that the set of edges E is then composed of the union of the
set of direct and indirect edges: E = D [ I. Weights in w
describe direct and indirect trust probability distributions and
are defined as follows:
w(i; j) = TDi;j , if (i; j) 2 D</p>
        <p>TIi;j , if (i; j) 2 I
Our goal is to determine how much a particular evaluator
can trust a peer . So the trust graph is constructed with
respect to ’s point of view only. Therefore, we maintain a
trust graph of the whole community containing all the direct
edges between peers (as they are needed to calculate indirect
trust relations), but we only maintain the indirect edges that
connect with the rest of the peers.</p>
      </sec>
      <sec id="sec-2-6">
        <title>Information Decay</title>
        <p>An important notion in our proposal is the notion of the decay
of information. We say the integrity of information decreases
with time. In other words, the information provided by a trust
probability distribution should lose its value over time and
decay towards a default value. We refer to this default value
as the decay limit distribution D. For instance, D may be the
uniform distribution, which describes that trust information
learned from past experiences tends to ignorance over time.</p>
        <p>To implement such a decay mechanism, we need to:
1. Record all evaluations ei 2 L made at time t with a
timestamp t, noted ei t .
2. Record all direct trust distributions TDi;j with a
timestamp t, noted TDit;j , where t is the timestamp of the last
evaluation that modified the trust distribution. The first
time TDi;j is calculated, t is the timestamp of the latest
evaluation amongst the two evaluations leading to this
calculation. (Recall that it is the similarity between two
evaluations or the evaluation rate that updates the
probability distribution.) Then, every time a new evaluation
with timestamp t0 &gt; t is considered to update TDit;j ,
TDit;j is first decayed from t to t0 before the distribution
is updated.
3. Record all indirect trust distributions TIi;j with a
timestamp t, noted TIit;j , where t is the time the distribution is
calculated. Every time TIi;j is calculated, all probability
distributions involved in this calculation will first need
to be decayed to the time of calculation t. The time of
calculation is usually the latest timestamp amongst the
timestamps of the distributions involved in this
calculation.</p>
        <p>Information in a trust probability distribution Ti;j decays
from t to t0 (where t0 &gt; t) as follows:</p>
        <p>t t0 = (D; Tit;j )
Ti;j
(7)
where
lim
t0!1</p>
        <p>is the decay function satisfying the property:
Tit;j t0 = D. One possible definition for could be:
t t0 =
Ti;j</p>
        <p>t
t;t0 Ti;j + (1
t;t0 )D
(8)
where is the decay rate, and:
t;t0 =</p>
        <p>The definition of t;t0 above serves the purpose of
establishing a minimum grace period, determined by the parameter
!, during which the information does not decay, and that once
reached the information starts decaying. The parameter tmax,
which may be defined in terms of multiples of !, controls the
pace of decay. The main idea behind this is that after the
grace period, the decay happens very slowly; in other words,
t;t0 decreases very slowly.
Given a peer assessment e , the question now is how to
compute the probability distribution of ’s evaluation. In other
words, what is the probability that ’s evaluation of is x
given that evaluated with e . As illustrated earlier, this is
expressed as the conditional probability:</p>
        <p>P(X=x j e )</p>
        <p>To calculate this conditional probability, the intuition is
that would tend to agree with ’s evaluation if his trust on
(that is, the expected similarity between their assessments
or the expected evaluation rate between their assessments) is
high. Otherwise, ’s evaluation would probably be different.
We perform then a sort of analogical reasoning: if in the past
gave assessments that were a certain degree dissimilar from
’s opinions, or with a certain evaluation rate with respect to
, then this will probably happen again now.</p>
        <p>We then calculate the above conditional probability based
on the following desired properties:</p>
        <p>If T ; is a flat distribution (i.e. a distribution
representing ignorance), then P(X j e ) should also be a flat
distribution. That is, the closer ’s trust on is to
ignorance, the less information is giving to with his/her
assessment.</p>
        <p>The degree of belief e = x should increase for those
points x whose similarity (or evaluation rate, in the case
of the ordered case) to e is high (i.e. for higher values
of T ; ).</p>
        <p>The degree of belief e = x should decrease for those
points x whose similarity (or evaluation rate, in the case
of the ordered case) to e is low trust (i.e. for lower
values of T ; ).</p>
        <p>Formally, these properties are achieved by defining the
probabilities accordingly (where the denominator of the
following two equations, Equations 9 and 10, is used for
normalisation to ensure that the resulting distribution is a probability
distribution):
x02E
x02E</p>
        <sec id="sec-2-6-1">
          <title>Non-Ordered Case.</title>
          <p>p(X=x j e ) = X eT ; (sim(e ;x0)) I(T ; )
eT ; (sim(e ;x)) I(T ; )</p>
        </sec>
        <sec id="sec-2-6-2">
          <title>Ordered Case.</title>
          <p>p(X=x j e ) = X eT ; (r(e =x0)) I(T ; )</p>
          <p>eT ; (r(e =x)) I(T ; )
where I(T ; ) is a measure of how informative the probability
distribution T ; is. We calculate I(T ; ) as:</p>
          <p>I(T ; ) = 1</p>
          <p>H(T ; )
where H describes the entropy of a probability distribution.
In other words, the lower the entropy of the distributions then
the more informative it is, and vice versa.</p>
          <p>We finally define the probability distribution of ’s
expected evaluation given ’s opinion accordingly: P(X j e ),
where X varies over the evaluation space E .
2.4</p>
        </sec>
      </sec>
      <sec id="sec-2-7">
        <title>Step 3: What to belief when many give opinions?</title>
        <p>In the previous section we computed P(X j e ). That is,
the probability distribution of ’s evaluation on given the
evaluation of a peer on . But what does do when there is
more than one peer assessing ?</p>
        <p>Given the set of opinions O describing a set of peer
evaluations over the object , we define the probability of ’s
assessment being x as follows:</p>
        <p>Y p(X=x j e )
p(X=x j O )) = X2OY p(X=x0 j e )</p>
        <p>x02E 2O</p>
        <p>In other words, the probability of ’s assessment on being
x given the set of opinions over is an aggregation (a product
in this case) of the probabilities of ’s assessment on being
x given each evaluation e 2 O .</p>
        <p>We then define the probability distribution of ’s expected
evaluation given all opinions in O as P(X j O ), where X
varies over the evaluation space E .</p>
        <p>We note that instead of the product operator Q other
connectives could be used, for instance the min operator might
be used. However, we note that using the minimum operator
does not take into account the number of assessments made.
That is, having assessments of 20 peers could be equivalent to
having the assessment of just one peer. In fact, the proposed
aggregation of Equation 12 ensures that:
(12)
The larger the number of identical opinions, the less
uncertain the final probability distribution is, and
The more trusted the opinions, the less uncertain the
final probability distribution is.
(9)
(10)
(11)</p>
        <p>Finally, to translate the final assessment from a probability
distribution P(X j O ) into a single value, we calculate the
mean (average) of the distribution and select the closest mark
to that mean.
2.5</p>
      </sec>
      <sec id="sec-2-8">
        <title>Step 4: What should be evaluated next?</title>
        <p>The previous three steps have provided a model to
calculate automated assessments of objects that have not been
assessed by , based on peers opinions. The level of
uncertainty of the automated assessments generated by our model
can be calculated as the uncertainty of the probability
distribution of ’s expected evaluation based on those peers
opinions P(X j O ). This level of uncertainty is measured by the
distribution’s entropy:</p>
        <p>H(P(X j O ))</p>
        <p>The question that naturally arises then is what objects can
be assessed next by to decrease such uncertainties? For
example, how many more assignments should a tutor evaluate
so that the uncertainty of the calculated assessments becomes
acceptable. We suggest to evaluate objects with maximum
uncertainty, or maximum entropic value. The ranking of
objects with respect to their entropic value is then defined as
follows:</p>
        <p>Rank( )
= 1 H(P(X j O ))
= 1 + X p(X=x j O ) ln p(x j O )
(13)
x2X
can then continue to evaluate objects one by one until the
uncertainty of the automated assessments becomes less than
some predefined acceptable uncertainty threshold.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and Future Work</title>
      <p>In this paper we have presented the personalised automated
assessments model (PAAS), a trust-based assessment service
that helps compute group assessments from the perspective
of a specific community member. This computation
essentially aggregates peer assessments, giving more weight to
those peers that are trusted by the specific community
member whom the automated assessments are computed for. How
much this specific member trusts a peer is then based on the
similarity or evaluation rate between his (past) assessments
and the peer’s (past) assessments over the same assignments.</p>
      <p>The proposed work is an extension of the work carried out
in [Gutierrez et al., submitted for publication]. In fact, the
COMAS model is a much more simplified model of the
nonordered case. It is much more simplified as it assumes that
the probability of the similarity between two assessors is 1 for
the aggregation of the similarities of past evaluations over the
same objects. PAAS’ use of probability distribution makes
it a richer and more informative model as much more
information is preserved in the calculations. Furthermore, PAAS
computes the uncertainty of the automated assessments,
helping suggesting which objects should be evaluated next in
order to decrease the overall uncertainty of PAAS’ calculations.</p>
      <p>In COMAS, experimental results were conducted on a real
classroom datasets as well as simulated data that considers
different social network topologies (where we say students
assess some assignments of socially connected students).
Results show that the COMAS method 1) is sound, i.e. the error
of the suggested assessments decreases for increasing
numbers of tutor assessments; and 2) scales for large numbers of
students.</p>
      <p>Future work on PAAS should follow a similar approach
for evaluation, where the same real classroom datasets can be
used as the groundtruth of marks, and we can then compare
PAAS’ automated assessments to that groundtruth.</p>
      <p>Additionally, we could also test the ranking of marks
(Section 2.5) by running experiments in a real classroom where
we ask the tutor to evaluate assignments once in a random
order and another time following the suggested ranking. This
could help us check whether the error decreases faster in the
latter case. Also, we expect to find that for a given acceptable
uncertainty threshold, the tutor should evaluate less
assignments in order to reach that threshold than evaluating
randomly.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>
        This work is supported by the CollectiveMind project
        <xref ref-type="bibr" rid="ref1">(funded
by the Spanish Ministry of Economy and Competitiveness,
under grant number TEC2013-49430-EXP)</xref>
        and the PRAISE
project (funded by the European Commission, under grant
number 388770).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[de Alfaro and Shavlovsky</source>
          , 2013] L. de Alfaro and
          <string-name>
            <given-names>M.</given-names>
            <surname>Shavlovsky</surname>
          </string-name>
          .
          <article-title>Crowdgrader: Crowdsourcing the evaluation of homework assignments</article-title>
          .
          <source>Thech. Report 1308.5273</source>
          , arXiv.org,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>[Li</surname>
          </string-name>
          et al.,
          <year>2003</year>
          ]
          <string-name>
            <given-names>Yuhua</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Zuhair A</article-title>
          .
          <string-name>
            <surname>Bandar</surname>
          </string-name>
          , and
          <string-name>
            <surname>David McLean</surname>
          </string-name>
          .
          <article-title>An approach for measuring semantic similarity between words using multiple information sources</article-title>
          .
          <source>IEEE Trans. on Knowl. and Data Eng</source>
          .,
          <volume>15</volume>
          (
          <issue>4</issue>
          ):
          <fpage>871</fpage>
          -
          <lpage>882</lpage>
          ,
          <year>July 2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Piech et al.,
          <year>2013</year>
          ]
          <string-name>
            <given-names>Chris</given-names>
            <surname>Piech</surname>
          </string-name>
          , Jonathan Huang, Zhenghao Chen, Chuong Do,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Ng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Daphne</given-names>
            <surname>Koller</surname>
          </string-name>
          .
          <article-title>Tuned models of peer assessment in moocs</article-title>
          .
          <source>Proc. of the 6th International Conference on Educational Data Mining (EDM</source>
          <year>2013</year>
          ),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Rubner et al.,
          <year>1998</year>
          ]
          <string-name>
            <given-names>Yossi</given-names>
            <surname>Rubner</surname>
          </string-name>
          , Carlo Tomasi, and
          <string-name>
            <given-names>Leonidas J.</given-names>
            <surname>Guibas</surname>
          </string-name>
          .
          <article-title>A metric for distributions with applications to image databases</article-title>
          .
          <source>In Proceedings of the Sixth International Conference on Computer Vision</source>
          (ICCV
          <year>1998</year>
          ), ICCV '
          <volume>98</volume>
          , pages
          <fpage>59</fpage>
          -, Washington, DC, USA,
          <year>1998</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Sierra and Debenham</source>
          , 2005]
          <string-name>
            <given-names>Carles</given-names>
            <surname>Sierra and John Debenham</surname>
          </string-name>
          .
          <article-title>An information-based model for trust</article-title>
          .
          <source>In Proceedings of the Fourth International Joint Conference on Autonomous Agents and Multiagent Systems, AAMAS '05</source>
          , pages
          <fpage>497</fpage>
          -
          <lpage>504</lpage>
          , New York, NY, USA,
          <year>2005</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Upton and Cook</source>
          , 2008]
          <string-name>
            <given-names>G.</given-names>
            <surname>Upton</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <source>A Dictionary of Statistics. Oxford Paperback Reference. OUP Oxford</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Walsh</source>
          , 2014]
          <string-name>
            <given-names>Toby</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <article-title>The peerrank method for peer assessment</article-title>
          . In Torsten Schaub, Gerhard Friedrich, and
          <string-name>
            <surname>Barry</surname>
            <given-names>O</given-names>
          </string-name>
          'Sullivan, editors,
          <source>ECAI 2014 - 21st European Conference on Artificial Intelligence</source>
          ,
          <fpage>18</fpage>
          -22
          <source>August</source>
          <year>2014</year>
          , Prague, Czech Republic - Including
          <source>Prestigious Applications of Intelligent Systems (PAIS</source>
          <year>2014</year>
          ), volume
          <volume>263</volume>
          <source>of Frontiers in Artificial Intelligence and Applications</source>
          , pages
          <fpage>909</fpage>
          -
          <lpage>914</lpage>
          . IOS Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chiclana</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Herrera-Viedma</surname>
          </string-name>
          .
          <article-title>Trust based consensus model for social network in an incompletelinguistic information context</article-title>
          .
          <source>Applied Soft Computing</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>