<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Rating by
Ranking: An Improved Scale for Judgement-Based Labels. In Proceedings of
Joint Workshop on Interfaces and Human Decision Making for Recommender
Systems, Como, Italy, August</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Rating by Ranking: An Improved Scale for Judgement-Based Labels</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jack O'Neill</string-name>
          <email>jack.oneill1@mydit.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarah Jane Delany</string-name>
          <email>sarahjane.delany@dit.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian Mac Namee</string-name>
          <email>brian.macnamee@ucd.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dublin Institute of Technology, School of Computing</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University College Dublin, School of Computer Science</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>27</volume>
      <issue>2017</issue>
      <abstract>
        <p>Labels representing value judgements are commonly elicited using an interval scale of absolute values. Data collected in such a manner is not always reliable. Psychologists have long recognized a number of biases to which many human raters are prone, and which result in disagreement among raters as to the true gold standard rating of any particular object. We hypothesize that the issues arising from rater bias may be mitigated by treating the data received as an ordered set of preferences rather than a collection of absolute values. We experiment on real-world and articially generated data, nding that treating label ratings as ordinal, rather than interval data results in an increased inter-rater reliability. is nding has the potential to improve the eciency of data collection for applications such as Top-N recommender systems; where we are primarily interested in the ranked order of items, rather than the absolute scores which they have been assigned.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Value-judgements, personal preferences and aitudes — oen
supplied in the form of a numerical rating — are an important source of
data for training machine learning models, ranging from emotion
recognition applications [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to recommender systems [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Typically, oracles providing such data are asked to supply a rating on an
N -point scale, representing either the extent to which a particular
aribute is judged to be present (i.e. how common or rare a
particular emotional state is [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), or how strongly an aitude or judgement
is felt (i.e. star-ratings of lms [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). Values provided in this way
(1 - N on an N -point scale) are known as absolute measurements,
as they implicitly depend on an absolute, idealised scale [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] of
measurement. e data we collect from such ratings may be treated
as interval data, as all values are expressed in a single, common,
numerical scale.
      </p>
      <p>
        Psychologists have long known of a number of bias errors to
which these scales are prone [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Researchers found that many
respondents can be categorised as displaying one or more of these
biases. Borman [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], has conducted a comprehensive study of these
biases; including the tendency to rate primarily at the high or low
end of the scale, lenient or severe biases, respectively; the tendency
to avoid making distinctions; and rating primarily around the centre
of the scale, known as range restriction, and its counterpart, the
tendency to rate primarily at either end of the scale, which we refer
to as a bias of extremity . Worryingly, Stockford and Bissel [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
found evidence that the ratings provided can even be inuenced
by the order of questions on the form, which has been termed
proximity bias.
      </p>
      <p>Label rankings represent an alternative method of judgement
elicitation to absolute ratings. When we elicit ranking data from
an oracle, we present a set of objects to be ranked, and the oracle
orders this set of objects in terms of how strongly a particular
aribute is judged to be present. Values provided in this manner
are known as relative measurements, as each value is dependent
on the other items present in the ranking set. e data collected
from such rankings may only be treated as ordinal data; it allows
us to determine a natural ordering of the data, but gives us no
information as to how inherent a particular aribute may be on
any absolute, common scale.</p>
      <p>Inter-rater reliability (IRR) measures the overall level of
agreement between raters as to the values they provide for identical data.
Although some disagreement among raters is inevitable in
situations involving judgements with some degree of subjectivity; we
assume that, given enough ratings for a particular item, the average
rating will eventually converge on its true, gold standard rating (i.e.
the average rating which would be obtained were we to sample the
entire propulation). A high IRR for a label-set indicates that there
is lile disagreement among raters as to the true gold standard
for that item; leaving us with more reliable data. Disagreement
between raters stems from genuine dierences in opinions on the
one hand, and factors such as noise and rater bias on the other. An
increase in IRR is only desirable if the increase is due to reducing
the laer.</p>
      <p>By proposing a general framework for recasting queries
requiring absolutely valued rating labels as a task needing only a series of
relative ratings, we hope to improve the reliability and validity of
scale-based data collection. is proposition relies on the intuition
that corpora of relatively-valued labels will produce higher levels of
inter-rater reliability (IRR) than those consisting of absolute values.
e intuition, in turn, is based on the assumption that although
rater biases aect the numeric label which raters will assign to an
item, as bias is constant for each individual, it should not aect the
order of preference of items for any given rater.</p>
      <p>D</p>
      <p>In our current work, we investigate this intuition empirically,
hypothesizing that ordinal-valued datasets systematically result in
a higher level of inter-rater reliability than their absolute-valued
equivalents. Section 2 reviews related research on absolute and
relative rating methods motivating the current study. Section 3
describes the methodology behind our experiment to test this
hypothesis both on articially generated and real-world datasets. We
relay our ndings in Section 4 and discuss the implications for
future work in Section 5
2</p>
      <p>
        RELATED WORK
e psychologist, Arthur Blumenthal argues that the human mind
is not capable of providing truly absolute ratings; that when we
are asked to provide labels on absolute scale, we compare each
item to similar items we have seen before. He argues that absolute
judgement involves the ”relation between a single stimulus and some
information held in short term memory about some former
comparison stimuli or about some previously experienced measurement scale
using which the observer rates the single stimulus” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. is
suggests that raters would be more comfortable with comparison-type,
or relative measurements, than they would be with the absolute
system of ratings so prevalent in data collection for machine
learning today. A recent study by Moors et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] has shown that
both modes of elicitation — rankings and ratings — produce
similar within-subjects results. is means it should be possible to
use either system without skewing the labels. More importantly,
related work has shown empirically that ratings collected using
comparative methods can, in certain circumstances, be more reliable
between-subjects than data collected using absolute measurement,
      </p>
      <p>
        To take a simple example, consider two raters, rater S, a severe
rater, and rater L, a lenient rater. Both raters are asked to provide
labels for a set of items. Rater S tends to give lower than average
ratings for items in general, whereas rater L tends to give higher
than average ratings. Table 1 shows the interval ratings provided
by both raters; while Table 2 shows these ratings as ordinal (ranked)
data. Although there is very lile agreement between raters on
individual items, there is evident agreement as to the ranking of
items.
for example in multi-criteria decision making [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and collaborative
ltering [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        When gathering absolutely-valued labels, there is no guarantee
that all raters share the same understanding of the absolute,
idealised scale on which the system is based, and this can result in
signicantly dierent behaviours, leading to ultimately unreliable
data. On top of this, it has been shown that results obtained from
rating exercises can be signicantly inuenced by the choice of
scale itself (5-point scale vs 7-point scale, for example) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Marsh
and Ball [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] argue that the eects of these phenomena can be seen
in the contemporary peer-review process, where the mean
singlerater reliability — the relation between two sets of independent
ratings of quality collected for a large number of submissions — of
reviewers for journal articles was very low at 0.27.
      </p>
      <p>
        is possibility is further evidenced by the experience of
Devillers et al. in collecting emotion-annotation data as part of the
HUMAINE project [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Raters used the FEELTRACE annotation
instrument for recording emotions in videos in real time [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
FEELTRACE is a video-annotation tool which allows raters to rate the
intensity of a given emotion in real time. Raters use a slider to
trace the intensity of a target emotion or trait, increasing the value
as the target intensies and decreasing the value as it wanes. e
researchers found strong correlations between raters tracing the
relative changes in emotions from moment-to-moment; however, they
recognised that the absolute values each of the raters chose showed
less uniformity. is suggests that although raters disagreed on the
(absolute) question of the intensity of the target, there was broad
agreement on the (relative) question of whether the intensity right
now is greater or less than the intensity in the moment immediately
preceding.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>EXPERIMENTAL METHODOLOGY</title>
      <p>We hypothesize that label sets collected on an interval scale (i.e.
absolute numbers, for example, on a scale of 1 - 9) will exhibit less
IRR than labels collected on an ordinal scale (i.e., labels provided in
terms of ranks within the dataset, from best to worst). Although
absolute ratings contain more information than relative ratings (a
relative ordering can be deduced from absolute ratings, but not vice
versa) —and as such, may naturally increase the IRR of the resulting
labels —we believe that some of the improvement is not accounted
for by this factor alone. In order to investigate our hypothesis we
examine 4 datasets introduced in previously published literature.
All of the datasets under consideration were rated using interval
labels. We then convert these labels into ordinal data using simple
intra-rater ranking; and use Krippendor’s α , which adjusts for
the inherent diculty of the problem, to compare the inter-rater
reliability of both datasets. is section describes the datasets used
in the experiment and discusses the suitability of Krippendor’s α
as a comparison metric.
3.1
In order to explore the impact of rater bias on IRR we generated
articial datasets of ratings for items provided by raters exhibiting
one of the four rater biases, discussed in Section 1; lenient, severe,
restricted and extreme. Our rst step is to model a set of items to be
rated. ese items could represent, for example, movies, where the
goal is to assign a score to each movie representing how good it
is; or a joke, where the goal is to rate the joke based on how funny
it is. In any case, we are not interested so much in the particular
item being rated so much as the gold standard towards which the
average rating from a large number of raters would converge. Each
item is represented as a randomly selected real number between 1
and 9 representing this gold standard rating. We generated 50 such
items to be rated.</p>
      <p>We next generate a base rating for each rater. is base rating is
determined by adding a modier randomly selected from a normal
distribution with μ 0 and σ 1. is value is rounded to the nearest
whole number and represents the rating this rater would assign in
the absence of rater bias. Each rater’s base rating diers slightly,
representing the inherent subjectivity of labels based on
valuejudgements.</p>
      <p>Next, we generate a bias modier, representing the extent to
which a rater’s inherent bias will aect the label provided. e bias
modier is a strictly positive value drawn from a normal distribution
with μ 0 and σ 1. We ensure the modier is positive by taking the
absolute value of the number drawn and disregarding the sign. We
split the raters into 4 groups of 20; with each group exhibiting one
of the rater biases discussed in 1. e manner in which this bias
modier is applied to each rater’s base rating is dependent on the
type of bias this rater exhibits. ese are as follows:</p>
      <p>Lenient ese raters tend to give ratings above the average
for all items. e bias modier is added to the base rating
for raters falling into this category
Severe Raters falling into this category tend to give
lowerthan-average ratings for all items. e bias modier is
subtracted form the base rating for raters falling into this
category
Restricted ese raters favour ratings around the mid-point
of the scale. If the base rating is below the mid-point, the
bias modier is added to the base rating. If the base rating
is above the mid-point, the bias modier is subtracted from
the base rating.</p>
      <p>Extreme is nal group of raters favours ratings at either
end of the scale. If the base rating is below the mid-point
of the scale, the bias modier is subtracted from the base
rating. If the base rating is above the mid-point, the bias
modier is added to the base rating.</p>
      <p>Figure 1 shows the impact of each of these rating biases on the
labels provided. e histograms in grey show the distribution of
base ratings for a particular item before bias has been applied. e
coloured histograms show the distribution of labels for each of the
4 groups aer bias has been applied.
3.2</p>
      <p>
        Empirical Datasets
e Jester Dataset, rst introduced by Goldberg et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], is a
corpus containing 4.1 million continuous ratings on a scale of 10:00 to
10:00 of 100 jokes from over 70,000 anonymous users. e dataset
in its original format is sparse, with many missing values. In
order to simplify the evaluation, we use a small subset of this data
containing 50 jokes each rated by the same 10 raters. Ratings are
specied to two decimal places.
      </p>
      <p>
        e BoredomVideos dataset is adapted from the corpus
introduced by Soleymani et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. A subset of the dataset has been
chosen to eliminate missing values; resulting in 31 clips, each rated
on a scale of 1-10 by 23 dierent raters. Ratings in this dataset are
provided in integer format.
      </p>
      <p>e MovieLens dataset in its original format consists of 10
million ratings provided by 72,000 users across 10,000 dierent
movies. We extracted a subset of 5,720 ratings, provided by 20 users
across 286 movies. Ratings were provided on a scale of 0:5 - 5, in
steps of 0:5. is subset contained no missing values.</p>
      <p>
        e Vera am Mittag German Audio-Visual Emotional Speech
Database (VAM), created by Grimm et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] contains emotion
recognition ratings gathered from segmented audio clips from a
non-acted German-language talk show. Ratings were provided on
three dimensions, activation (how active or passive the speaker is),
evaluation (how positive or negative the emotion is), and dominance
(how dominant the speaker is). For the purposes of this study, we
selected only the activation ratings. For this experiment, we used
a subset of the data containing 478 speech instances rated by 17
dierent raters. All ratings were provided on a scale of -1.0 to 1.0,
in steps of 0.5.
3.3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Converting Rating Data to Ranking Data</title>
      <p>All of the datasets described in Section 3.2 consist of ratings
provided on an absolute scale. In order to convert these ratings to
rankings we use a simple intra-rater ranking function with average
rank used in the case of ties. Given a set of raters R, and a set of
items to rate, X, with ri j representing the rating provided for the
ith item by the jth rater and Ri representing the set of all ratings
provided by rater i, the interval and ordinal values for these ratings
are described in Equation 1.</p>
      <p>LI nt erval = ri j
LOr dinal = rankRi ¹ri j º
xi 2 X , ri 2 R
(1)
3.4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Metrics</title>
      <p>
        Inter-rater reliability can be computed using a wide variety of
measurements, the choice of which is oen dependent on the properties
of the data being investigated [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. An important factor in
determining the most suitable IRR metric is the scale of the data under
examination. For example, Kendall’s τ assumes that the labels
provided are on an ordinal scale, whereas Pearson’s ρ is stricter in that
it requires labels to be on an interval scale. Further complicating
the comparison is the question of commensurability of results.
Interval data is more ne-grained than ordinal data, providing both
an absolute value measurement and a relative ordering of items on
the scale. For raters to agree on an interval scale, they must agree
on both the ordering and the abolute value of each item rated. On
an ordinal scale, raters are required to agree only on the relative
ordering of each item. is suggests that inter-rater agreement
on ordinal data is fundamentally easier to achieve, and we want
to ensure that any improvement in IRR for ordinal data is not due
simply to the higher standard of agreement required by interval
data.
To overcome these diculties, we use Krippendor’s α [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
to compare the IRR between the ordinal and interval label sets.
Krippendor’s α is a generalized reliability measurement which
expresses the ratio of observed agreement over agreement expected
due to chance. In its most general form, Krippendor’s α is dened
as
where Do is the observed disagreement, and De is the
disagreement that can be expected when chance prevails.
      </p>
      <p>Krippendor’s α can be applied to multiple raters simultaneously,
and allows comparisons to be made between values obtained on
data using diering scales of measurement. e α value ranges
from -1 (perfect disagreement) to 1 (perfect agreement). Having
calculated the overall α for each dataset, we decompose the results
by calculating the α for each pair of raters individually. e α value
obtained for each pair of raters shows the level of agreement for
that pair of raters independent of all others. We plot the individual
α values on a heatmap, demonstrating that the improved accuracy
results from a consistent increase in the pair-wise accuracy across
all raters.
4
4.1</p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS Articial Datasets</title>
      <p>We compared the IRR of the simulated datasets using Krippendor’s
α , treating the data rst as interval and then as ordinal data. Aer
30 repetitions, the α for ordinal rankings was consistently higher
than that of the interval data, and a paired Wilcoxon Signed-Rank
test rejected the alternative hypothesis, that the shi in mean α was
0, with a p value of ¡ 0.001. Figure 2 depicts the Krippendor’s α for
both interval and ordinal data over 30 iterations as a box plot. is
experiment demonstrates that ordinal data is more reliable than
interval data, under the assumptions we made when modelling
our articial labels. However, it remains to be seen whether this
improvement persists when working with real-world datasets.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Empirical Datasets</title>
      <p>Table 3 summarises the α values for each dataset when treated both
as interval and as ordinal data. Although the underlying
agreement in each of the datasets was low —as can be expected with
datasets containing fundamentally subjective labels —we obtained
consistently higher IRR by treating the ratings as ordinal data. e
MovieLens dataset, in particular, showed a considerable
improvement when treated as ordinal data. Although the improvement
in the Jester dataset and the BoredomVideos dataset was smaller,
the improvement relative to the original inter-rater agreement was
quite high. e relative improvement of the VAM dataset was lower
than the others. We noted in Section 4.1 that rater bias is a
particular feature of subjective ratings. e lower performance of the VAM
dataset may be due in part to the fact that identifying emotions in
others is less subjective than making personal value judgements
and so the inherent rater bias in this dataset is less pronounced
than the genuine dierences in opinions between raters.</p>
      <p>Figure 3 visualises the IRR improvement of ordinal treatment
over interval on the BoredomVideo dataset. Each rater is shown
on both the X axis and the Y axis, with each cell representing the
Krippendor’s α of IRR between the raters. Krippendor’s α scores
above 0, indicating higher agreement than that expected by chance
are shown as blue, while Krippendor’s α scores below 0,
indicating lower agreement than that expected by chance are represented
as red. is visualisation shows that the improvement in IRR was
consistent across all pairs of raters reinforcing the ndings in
Section 4.1 that ranking data yields signcantly higher IRR than rating
data.
5</p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSION</title>
      <p>Our experiments suggest that treating label sets as ordinal, rather
than interval in scale tends to suppress the eects of label bias
in generating inter-rater disagreement, and consequently leads to
more reliable datasets. As this increased reliability results from a
reduction of noise (i.e. rater bias), we believe this nding could
be utilised to generate reliable labels with fewer oracle queries,
reducing the overall cost of data collection. is nding, however,
comes with a number of caveats.</p>
      <p>Firstly, ordinal data cannot be directly transposed back to an
interval scale. is approach will work best when we are more
interested in the relative ordering of items rather than their
absolute values, for example, Top-N recommendation. If absolute
label values are required, a further step, (and further information)
will be needed to infer absolute values from our ranked data. We
believe that investigating possible approaches to making such an
inference may expand the applicability of our proposed method of
data collection.</p>
      <p>
        Secondly, our experiments were conducted on real-world data
gathered by researchers using an absolutely-valued measurement
scale. One of the reasons absolute values are more commonly used
in label collection is that it they are easier to elicit from oracles.
Researchers have long been aware of the diculty in geing raters
to rank large collections of items [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. A key challenge in collecting
ranking data lies in building a method to eciently break large
collections of data into smaller subsets which can be easily ranked
by human labellers; and then to recombine these sets of partial
orderings into a reliable super-ordering over all labels in the set.
Developing an algorithm to present items for rating in such a
manner as to maximise the eciency of this process is an essential
next-step in making our proposed approach production-ready
irdly, our experiments do not necessarily reect the potential
results which would be seen on a label set collected using rankings.
However, there is reason to be optimistic that labels collected using
ranking — rather than rating — methods would result in even higher
reliability improvements. Related work, as outlined in Section 2,
suggests that human oracles are more reliable in ranking small sets
of items than they are at providing absolute ratings. Once a method
for eectively collecting ranking labels over a large set of data has
been determined, conducting this experiment on data collected
using ranking would allow us to ascertain the full benets of this
approach to ecient label collection.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Alwin</surname>
            ,
            <given-names>D. F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Krosnick</surname>
            ,
            <given-names>J. A. </given-names>
          </string-name>
          <article-title>e measurement of values in surveys: A comparison of ratings and rankings</article-title>
          .
          <source>Public Opinion arterly 49</source>
          ,
          <issue>4</issue>
          (
          <year>1985</year>
          ),
          <fpage>535</fpage>
          -
          <lpage>552</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Blumenthal</surname>
            ,
            <given-names>A. L. </given-names>
          </string-name>
          <article-title>e process of cognition</article-title>
          .
          <source>Experimental psychology series.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Borman</surname>
            ,
            <given-names>W. C.</given-names>
          </string-name>
          <article-title>Consistency of rating accuracy and rating errors in the judgment of human performance</article-title>
          .
          <source>Organizational Behavior and Human Performance</source>
          <volume>20</volume>
          ,
          <issue>2</issue>
          (
          <year>1977</year>
          ),
          <fpage>238</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Cowie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Apolloni</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Romano</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Fellenz</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>What</surname>
          </string-name>
          <article-title>a neural net needs to know about emotion words</article-title>
          .
          <source>Computational intelligence and applications 404</source>
          (
          <year>1999</year>
          ),
          <fpage>5311</fpage>
          -
          <lpage>5316</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cowie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Cornelius</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          <article-title>Describing the emotional states that are expressed in speech Describing the emotional states that are expressed in speech</article-title>
          .
          <source>Speech Communication</source>
          <volume>40</volume>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          (
          <year>2003</year>
          ),
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Cowie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douglas-Cowie</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savvidou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mcmahon</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sawey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and Schro¨der,
          <string-name>
            <surname>M.</surname>
          </string-name>
          
          <article-title>Feeltrace: An instrument for recording perceived emotion in real time</article-title>
          .
          <source>ISCA Workshop on Speech &amp; Emotion</source>
          (
          <year>2000</year>
          ),
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Dawes</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Do data characteristics change according to the number of scale points used? An experiment using 5-point, 7-point and 10-point scales</article-title>
          .
          <source>International Journal of Market Research</source>
          <volume>50</volume>
          ,
          <issue>1</issue>
          (
          <year>2008</year>
          ),
          <fpage>61</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Devillers</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cowie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>J.-c.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abrilian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mcrorie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Real life emotions in French and English TV video clips : an integrated annotation protocol combining continuous and discrete approaches</article-title>
          .
          <source>In Language Resources and Evaluation</source>
          (
          <year>2006</year>
          ), pp.
          <fpage>1105</fpage>
          -
          <lpage>1110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roeder</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Perkins</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Eigentaste: A constant time collaborative ltering algorithm</article-title>
          .
          <source>Information Retrieval 4</source>
          ,
          <issue>2</issue>
          (
          <year>2001</year>
          ),
          <fpage>133</fpage>
          -
          <lpage>151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Grimm</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kroschel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Narayanan</surname>
            ,
            <given-names>S. </given-names>
          </string-name>
          <article-title>e vera am miag german audio-visual emotional speech database</article-title>
          .
          <source>In Multimedia and Expo</source>
          , 2008 IEEE International Conference on (
          <year>2008</year>
          ), IEEE, pp.
          <fpage>865</fpage>
          -
          <lpage>868</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Krippendorff</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>Answering the call for a standard reliability measure for coding data</article-title>
          .
          <source>Communication methods and measures 1</source>
          ,
          <issue>1</issue>
          (
          <year>2007</year>
          ),
          <fpage>77</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kamishima</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>Nantonac Collaborative Filtering  Recommendation Based on Order Responces</article-title>
          .
          <source>In Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          (
          <year>2003</year>
          ), vol.
          <volume>90</volume>
          , pp.
          <fpage>4</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Krippendorff</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>Content analysis: An introduction to its methodology</article-title>
          .
          <source>Sage</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Marsh</surname>
            ,
            <given-names>H. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ball</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , P. e Peer Review Process Used to Evaluate Manuscripts Submied to Academic Journals : Interjudgmental Reliability.
          <fpage>151</fpage>
          -
          <lpage>169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Moors</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vriens</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelissen</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vermunt</surname>
            ,
            <given-names>J. K.</given-names>
          </string-name>
          <article-title>Two of a kind. similarities between ranking and rating data in measuring values</article-title>
          .
          <source>In Survey Research Methods</source>
          (
          <year>2016</year>
          ), vol.
          <volume>10</volume>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales</article-title>
          .
          <source>Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics</source>
          <volume>3</volume>
          ,
          <issue>1</issue>
          (
          <year>2005</year>
          ),
          <fpage>115</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Saal</surname>
            ,
            <given-names>F. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lahey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>a. Rating the ratings: Assessing the psychometric quality of rating data</article-title>
          .
          <source>Psychological Bulletin</source>
          <volume>88</volume>
          ,
          <issue>2</issue>
          (
          <year>1980</year>
          ),
          <fpage>413</fpage>
          -
          <lpage>428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Saaty</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          <article-title>Rank from comparisons and from ratings in the analytic hierarchy/network processes</article-title>
          .
          <source>European Journal of Operational Research</source>
          <volume>168</volume>
          , 2 SPEC. ISS. (
          <year>2006</year>
          ),
          <fpage>557</fpage>
          -
          <lpage>570</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Soleymani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Crowdsourcing for aective annotation of video: Development of a viewer-reported boredom corpus</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Stockford</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bissel</surname>
            ,
            <given-names>H. W.</given-names>
          </string-name>
          <article-title>Factors involved in establishing a merit-rating scale</article-title>
          .
          <source>Personnel</source>
          (
          <year>1949</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>