<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Expert Disagreement in Sequential Labeling: A Case Study on Adjudication in Medical Time Series Analysis?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mike Schaekermann</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edith Law</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kate Larson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew Lim</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arti cial Intelligence Group, School of Computer Science, University of Waterloo</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Division of Neurology, Sunnybrook Health Sciences Centre, University of Toronto</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>HCI Lab, School of Computer Science, University of Waterloo</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Low inter-rater agreement is typical in various expert domains that rely in part on subjective evaluation criteria. Prior work has predominantly focused on expert disagreement with respect to individual cases in isolation. In this work, we report results from a case study on expert disagreement in sequential labeling tasks where the interpretation of one case can a ect the interpretation of subsequent or previous cases. Three board-certi ed sleep technologists participated in face-to-face adjudication sessions to resolve disagreement in the context of sleep stage classi cation. We collected 1,920 independent scoring decisions from each expert on the same dataset of eight 2-hour long multimodal medical time series recordings. From all disagreement cases (29% of the dataset), a representative subset of 30 cases was selected for adjudication and expert discussions were analyzed for sources of disagreement. We present our ndings from this case study and discuss future application scenarios of expert discussions for the training of non-expert crowdworkers.</p>
      </abstract>
      <kwd-group>
        <kwd>Inter-rater disagreement Adjudication Sequence data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        One of the most common use cases for crowdsourcing is the classi cation of
objects into categories. While crowdsourced classi cation tasks traditionally
focused on problems not requiring domain expertise, recent work suggests that
crowdsourcing can also be e ective for expert-level classi cation. Examples of
such expert tasks from the medical domain include the identi cation of low-level
patterns in sleep-related biosignals [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], the annotation of retinal images [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
and medical relation extraction [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        In many mission-critical expert domains including the interpretation of
medical data, low inter-rater agreement rates are the norm [
        <xref ref-type="bibr" rid="ref11 ref15 ref16 ref5">5, 11, 15, 16</xref>
        ]. Expert
disagreement, however, poses fundamental challenges to quality control
procedures in crowdsourcing, and to the use of data labels in supervised machine
? Supported by NSERC CHRP (CHRP 478468-15) and CIHR CHRP (CPG-140200).
learning, as it is not immediately obvious how cases at the inter-subjective
decision boundary should be disambiguated if multiple equally-quali ed domain
experts exhibit genuine disagreement.
      </p>
      <p>
        Prior work has predominantly paid attention to the nature, sources and
resolvability of expert disagreement on individual classi cation tasks in isolation
[
        <xref ref-type="bibr" rid="ref1 ref13 ref21 ref4">1, 4, 13, 21</xref>
        ]. Many interpretation tasks, however, are sequential in nature, i.e.,
the interpretation of one case a ects the interpretation of subsequent or
previous cases. For example, in text translation, the semantic interpretation of one
phrase or sentence can a ect the translation of subsequent or previous phrases or
sentences. Heidegger called this reciprocity of text and context the hermeneutic
circle. Overall, sequential labeling makes up a large and diverse class of problems
from numerous expert domains.
      </p>
      <p>
        In this work, we present ndings from a case study on expert disagreement in
the context of sleep stage classi cation, the expert task of mapping a sequence of
xed-length pages of continuous multimodal medical time series (polysomnogram,
see Figure 1) to a sequence of discrete sleep stages (hypnogram). Prior work has
established that inter-rater agreement in sleep staging averages around 82.6%
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The objective of this case study is to identify various sources of expert
disagreement in sleep stage classi cation and to investigate if and to what extent
disagreement may be speci c to the sequential nature of the labeling task and
underlying data.
      </p>
      <p>To answer these questions, we collected 1,920 independent sleep scoring
decisions from a committee of three board-certi ed sleep technologists. We then
selected a representative subset of the resulting disagreement cases which were
resolved through in-person adjudication among the members of the expert
committee. The rest of this paper describes the related work, then details our study
for collecting and analyzing the expert deliberation data, and concludes with a
discussion of application scenarios for the training of non-expert crowdworkers.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>Ambiguity and Sources of Inter-rater Disagreement</title>
        <p>
          Ambiguity, the quality of being open to more than one interpretation, and the
phenomenon of expert disagreement are central to the justi cation of knowledge,
and have been extensively discussed in the epistemic literature [
          <xref ref-type="bibr" rid="ref1 ref13 ref21 ref4">1, 4, 13, 21</xref>
          ]. An
early theoretical investigation named three types of expert disagreement [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]:
personality-based disagreement arising from the incompetence, ideology, or
venality of experts, judgment-based disagreement arising from information gaps,
or structural disagreement that arises because experts adopt di erent organizing
principles or problem de nitions. Garbayo [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], on the other hand, distinguished
a form of legitimate disagreement, that arises when experts can access the same
evidence, but still diverge in interpretations, from verbal disagreement, i.e.,
misunderstanding among experts due to discrepancies in terminology.
        </p>
        <p>
          Recent work in the eld of human-computer interaction (HCI) has explored
the issue of disagreement in the context of crowdsourcing tasks. Gurari and
Graumen [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] analyzed visual question answering tasks and found that disagreement
can be attributed to ambiguous and subjective questions, insu cient or
ambiguous visual evidence, di ering levels of annotator expertise, and vocabulary
mismatch. Chang et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] proposed to elicit help from the crowd for re nement
of category de nitions, based on the nding that workers may disagree because
of incomplete or ambiguous classi cation guidelines. Kairam and Heer [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
introduced a technique to identify clusters of workers with diverging, but legitimate
interpretations of the same task. Their work shows that disagreement can arise
from di erences in how liberally or conservatively workers interpret classi caion
guidelines.
        </p>
        <p>
          Our study revolves around the task of biomedical time series classi cation, a
eld with typically low inter-scorer reliability. For example, Rosenberg and van
Hout [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] conducted a large-scale study on inter-scorer reliability in sleep stage
classi cation and found that average expert agreement is as low as 82.6%. In a
comment on this study, Penzel et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] explained that systematic studies on
the inter-rater reliability of sleep automatically bring up the question of truth,
claiming that the \true" state (i.e., sleep stage) is unknown and can only be
approximated through aggregation of expert opinions.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Group Deliberation as a Method for Disambiguation</title>
        <p>Group deliberation is an interactive form of decision making among humans
which typically involves group members with con icting beliefs who try to reach
consensus on a given question by presenting arguments, weighing evidence and
reconsidering individual positions.</p>
        <p>
          Several works explored factors that a ect the process and outcomes of group
deliberation. Solomon [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] appreciates con ict as an important phenomenon of
any fruitful deliberation process. He argues that dissent is both required and
useful|as \dissenting positions are associated with particular data or insights
that would be otherwise lost in consensus formation"|and criticizes procedures
endowed with the a priori aim of reaching consensus. Instead, he advocates for a
structured deliberation procedure that avoids the undesired e ects of groupthink
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] by actively encouraging dissent, organizing individual subgroups to deliberate
on the same question, and ensuring diverse group compositions.
        </p>
        <p>
          Kiesler and Sproull [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] found that time limits imposed on deliberation tend
to polarize discussions and to decrease the number of arguments exchanged. The
same work suggests the use of voting techniques or explicit decision protocols to
structure the deliberation process.
        </p>
        <p>
          Recent work by Schaekermann et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] introduced a real-time deliberation
framework to disambiguate edge cases in crowdsourced classi cation tasks
drawing inspiration from some of these early design considerations. The same work
also introduced a novel public deliberation dataset including all deliberation
dialogues, original and revised classi cation decisions, and evidence regions from
two di erent text classi cation tasks.
        </p>
        <p>
          Navajas et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] studied the e ectiveness of in-person group deliberation
for general-knowledge questions reporting that averaging consensus decisions
yielded better results than averaging individual responses.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Consensus Scoring in Medical Data Analysis</title>
        <p>
          Group deliberation has also been proposed as a technique for disambiguating
edge cases in the interpretation of medical data. Rajpurkar et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] employed
group deliberation among cardiologists to generate a high-quality validation data
set in the context of arrythmia detection from electrocardiograms (ECGs). Their
work revealed that a convolutional neural network trained on independent
labels (i.e., labels collected without deliberation) exceeded the classi cation
performance of individual cardiologists when benchmarked against the consensus
validation set.
        </p>
        <p>
          Krause et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] compared majority vote to in-person deliberation as
techniques for aggregating expert opinions for diagnosing eye diseases from photos
of the eyeground. Compared to majority vote, in-person deliberation yielded
substantially higher recall, suggesting the potential of group deliberation for
mitigating underdiagnosis of diabetic retinopathy and diabetic macular edema.
Krause et al. also showed that performing group deliberation on a small portion
of the entire data set can make tuning of hyperparameters for deep learning
models more e ective. The same consensus data set was later used by Guan
et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to validate the classi cation performance of a novel machine learning
approach involving the training of multiple grader-speci c models. They
demonstrated that training and aggregating separate grader-speci c models can be
more e ective than training a single prediction model on majority labels.
        </p>
        <p>
          In the context of sleep stage classi cation, Penzel et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] refer to the
concept of in-person group deliberation as consensus scoring, concluding that
an \optimal training for [...] sleep scorers is participation in consensus scoring
rounds". In this work, we translate this idea to the non-expert domain suggesting
a method to augment training procedures for crowdworkers through the use of
edge-case examples and the associated expert discussion dialogues in the context
of sleep stage classi cation.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Expert Deliberation Data Set</title>
      <p>An in-person deliberation study was conducted with an expert committee of
three board-certi ed sleep technologists at Sunnybrook Health Sciences Centre
in Toronto to investigate the extent and potential sources of inter-rater
disagreement in sleep stage classi cation, and the e ectiveness of group deliberation as
a method for consensus formation.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Set</title>
        <p>
          We prepared a data set of eight 2-hour-long PSG recording fragments. Each
2hour-long fragment contained a sequence of 240 30-second epochs of biosignal
data, resulting in 1,920 (240 x 8) epochs for the entire data set. Half of the
fragments were from healthy subjects, the other half from patients with
Parkinson's disease. Both parts of the data set (Healthy and Parkinson) contained
examples of di erent transition types. We included examples from four di erent
transition types identi ed by Rosenberg et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] as regions with typically low
inter-rater agreement: the last epoch of stage Wake before sleep onset, the rst
epoch of stage N2 after stage N1, the rst epoch of stage REM after stage N2,
and transitions between stages N2 and N3.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Procedure</title>
        <p>The full data set was rst scored independently by each sleep technologist,
resulting in 5,760 individual scoring decisions, three for each of the 1,920 epochs. We
then identi ed all epochs with disagreement among scorers and selected a subset
of 30 epochs for in-person group deliberation. The selected disagreement epochs
represented both disease states and all four transition types. All 30 epochs were
discussed in person by the three scorers using a graphical scoring interface to
facilitate detailed discussions about patterns present in the time series data. The
experts participants were not explicitly required to reach unanimous consensus,
and could instead choose to declare a case as irresolvable. We did not impose an
explicit voting scheme or limit the amount of time available per discussion, but
instead left the discussion dynamics open until all experts either agreed on one</p>
        <p>Tech A Tech B Tech C Majority # Obs.</p>
        <p>Tech B
Tech C
Majority
sleep stage or declared a case as irresolvable. Unanimous decisions were reached
for all 30 epochs through a process of verbal argumentation and re-interpretation
of the patterns shown in the biosignal data The irresolvable option was never
used. Discussions were recorded (screen capture and audio), transcribed and
qualitatively coded for the di erent sources of disagreement.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Inter-rater Disagreement</title>
        <p>We measured pairwise agreement between all scorers (Tech A, Tech B, Tech C),
as well as the group labels as determined by majority vote (Majority) and the
deliberation process (Deliberation). Agreement was measured by Cohen's kappa.
Table 1 summarizes all agreement results. Pairwise agreement among scorers was
moderate, ranging between 0.68 and 0.71 (N=1920). Agreement between
individual scorers and the majority vote was high, between 0.84 and 0.87 (N=1920).
For the epochs discussed in person, we measured pairwise agreement between the
deliberation decision and individual scorers' decisions. Two of the three scorers
showed weak agreement with deliberation outcomes (Cohen's kappa of 0.63 and
0.50, N=30), while the third scorer showed no systematic agreement with the
deliberation outcomes (Cohen's kappa of 0.02, N=30). Agreement between the
majority vote and deliberation decisions was low (Cohen's kappa of 0.54, N=30).
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Sources of Disagreement</title>
        <p>Initial qualititative coding of the expert discussions for 17 cases from two major
transition types revealed a broad range of reasons why sleep technologists may
disagree on the correct sleep stage label. Figures 2 and 3 compare the relative
frequency of di erent sources of disagreement across two transition types (Last
W before sleep and First N2 after N1 ) and across two disease states (Healthy and
Parkinson) respectively. Overall, we identi ed two sources of disagreement which
occurred with the highest frequency in both transition types and disease states.
These were (a) the presence of multiple stages in one epoch causing disagreement
about which stage was the dominant one, and (b) di erent con gurations of
the graphical scoring interface in terms of amplitude scaling causing divergent
interpretations of visual patterns in the signal.</p>
        <p>Expert s
Crowd Workers
E1 +
E2 +
E3 +
+
-
-
Labeling
(independent)
+ ?
E1: Argument</p>
        <p>E2: Argument
E3: Argument</p>
        <p>While these two sources of disagreement could persist on individual cases
without the sequential context, we identi ed two other sources of disagreement
that explicitly depend on the sequential nature of the labeling task and
underlying data:
{ Number of scoring passes:
for 3 out of 30 adjudicated cases, experts explicitly mentioned that their
scoring decision depended on the number of passes they had taken on a
particular recording. In other words, experts indicated that their interpretation
of biosignals is often updated once certain patient-speci c patterns are
observed towards the end of the recording. A subsequent re-interpretation (i.e.,
second scoring pass) would then allow experts to take into account
observations they have made in the other parts of the data sequence in one of the
earlier scoring passes. Disagreement could therefore arise if one expert had
only performed one initial pass whereas other experts may have performed
two or more passes.
{ Cascade from previous disagreement:
3 out of 30 adjudicated cases could be resolved automatically once the
disagreement on one of the close-by preceding cases had been resolved. This
dynamic was observed since evidence for speci c stages of sleep may
sometimes be observed only at the transition point from one sleep stage to
another. Consequently, disagreement may arise at a \critical" transition point
and persist over multiple steps in the sequence. Once the disagreement at
the transition point is resolved, the resolution can cascade to the subsequent
steps until the next transition point.</p>
        <p>These two sources of disagreement co-occurred once, meaning that 5 out of
30 adjudicated cases (17%) were associated with sources of disagreement that
depend on the sequential nature of the labeling task and underlying data.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>In this work, we provided an initial investigation of expert disagreement in the
context of sequential labeling tasks, studying the e ectiveness of in-person
adjudication for resolving disagreement and for analyzing information about the
original source of disagreement.</p>
      <p>Our results suggest that majority vote is not necessarily a good proxy for
group deliberation decisions in sleep staging. This nding provides some con
dence in the usefulness of expert discussions for the purpose of resolving
disagreement cases. Beyond that, our qualitative analysis of expert discussion dialogues
uncovered a diverse set of di erent reasons why domain experts disagree in the
context of sleep stage classi cation, most of which go beyond the notion of mere
input mistakes.</p>
      <p>Perhaps most importantly, we identi ed two sources of disagreement with
a clear connection to the sequential nature of the labeling task and underlying
data. This observation provides some support for our hypothesis that the
reciprocity of data and context in sequential labeling may lead to unique forms of
expert disagreement that are characteristic for sequential labeling tasks, where
the interpretation of one case a ects the interpretation of subsequent or
previous cases. One exciting avenue for future research is the problem of whether
it is possible to detect the \critical" tasks that might set up a cascade of
disagreement and potentially incorrect labels. Successful detection of such \critical"
tasks would allow for a more cost-e ective use of expert resources by focusing
disambiguation procedures on those cases and saving expert resources on other
cases that may be resolved automatically.</p>
      <p>
        Picking up on Penzel et al.'s comment on the nature of \truth" in sleep
staging [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], some of the inherent di culty may arise because there exists a
certain degree of both temporal and spatial continuity at transitions between
states. In other words, despite the fact that any single neuron or cortical circuit
may be thought of as existing in one state or another at any given moment, it is
possible for local assemblies of neurons to take some time to transition from one
state to another, and also that distant assemblies of neurons in di erent parts
of the brain can exist in di erent states at the same time. These transitions
may take minutes [
        <xref ref-type="bibr" rid="ref18 ref23">18, 23</xref>
        ] which encompasses several 30-second epochs. Thus, we
hypothesize that some of the ambiguity stems from the need to force transitional
states into one sleep stage category or another.
      </p>
      <p>
        We posit that expert disagreement in complex tasks can be used as a signal
to identify ambiguous edge cases, and as a driver for eliciting conclusive expert
discussions to disambiguate such edge cases. For future work, we propose the
idea that example-based training procedures for non-expert crowdworkers may
bene t from the presentation of edge cases and their associated expert
discussions. While expert disagreement may be one signal for the identi cation of edge
cases, other techniques for the automatic selection of edge case examples, e.g.,
based on measures of machine uncertainty, have been proposed in prior work
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We believe that expert disagreement and the associated expert discussions
open up interesting opportunities for optimizing example-based training
procedures for human learners, e.g., to improve disambiguation skills and depth of
understanding.
      </p>
      <p>Figure 4 illustrates a high-level overview of some of these future directions. In
summary, we hope to conduct research on augmenting example-based training
procedures for non-expert crowdworkers using edge-cases and their associated
expert discussions to help human learners develop more accurate classi cation
strategies for expert-level tasks exhibiting a certain amount of ambiguity.</p>
      <p>Another promising avenue for future work will be to explore the minimum
\bandwidth" and e ective protocols of communication between experts needed
to result in successful disambiguation in the context of sequential labeling
settings like the one presented in this work. Comparisons may include di erent
styles of expert communication ranging from online text-based asynchronous
approaches, to in-person verbal real-time communication.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this work, we reported results from a case study on expert disagreement
in sequential labeling tasks where the interpretation of one case can a ect the
interpretation of subsequent or previous cases. Three board-certi ed sleep
technologists scored 1,920 cases in a sequential 5-class labeling task. Out of all
disagreement cases, 30 cases were discussed and resolved through face-to-face
adjudication. We identi ed various sources of disagreement that are speci c to the
sequential nature of the underlying data and labeling procedure. Our work
concluded with a discussion of promising application scenarios of expert discussions
for the training of non-expert crowdworkers that we hope to explore in future
work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Beatty</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Should We Aim for Consensus? Episteme</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <volume>198214</volume>
          (
          <year>2010</year>
          ). https://doi.org/10.3366/E1742360010000948
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <issue>2</issue>
          .
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amershi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamar</surname>
          </string-name>
          , E.:
          <article-title>Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets</article-title>
          .
          <source>In: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems - CHI '17</source>
          . pp.
          <volume>2334</volume>
          {
          <fpage>2346</fpage>
          . ACM, ACM Press, New York, New York, USA (
          <year>2017</year>
          ). https://doi.org/10.1145/3025453.3026044, http://dl.acm.org/citation.cfm?doid=
          <volume>3025453</volume>
          .
          <fpage>3026044</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dumitrache</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aroyo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Crowdsourcing Ground Truth for Medical Relation Extraction</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems</source>
          <volume>8</volume>
          (
          <issue>2</issue>
          ),
          <volume>1</volume>
          {
          <issue>20</issue>
          (7
          <year>2018</year>
          ). https://doi.org/10.1145/3152889, http://dl.acm.org/citation.cfm?doid=
          <volume>3232718</volume>
          .
          <fpage>3152889</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Garbayo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Epistemic Considerations on Expert Disagreement, Normative Justication, and Inconsistency Regarding Multi-criteria Decision Making. Constraint Programming</article-title>
          and
          <source>Decision Making</source>
          <volume>539</volume>
          ,
          <issue>35</issue>
          {
          <fpage>45</fpage>
          (
          <year>2014</year>
          ),
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Guan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gulshan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.:
          <article-title>Who said what: Modeling individual labelers improves classi cation</article-title>
          .
          <source>In: AAAI Conference on Arti cial Intelligence</source>
          (
          <year>2018</year>
          ), https://arxiv.org/pdf/1703.08774.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gurari</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grauman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>CrowdVerge: Predicting If People Will Agree on the Answer to a Visual Question</article-title>
          .
          <source>In: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems - CHI '17</source>
          . pp.
          <volume>3511</volume>
          {
          <fpage>3522</fpage>
          . ACM, ACM Press, New York, New York, USA (
          <year>2017</year>
          ). https://doi.org/10.1145/3025453.3025781, http://dl.acm.org/citation.cfm?doid=
          <volume>3025453</volume>
          .
          <fpage>3025781</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>A.M.:</given-names>
          </string-name>
          <article-title>Victims of Groupthink: A Psychological Study of Foreign Policy Decisions and Fiascoes</article-title>
          .
          <source>The ANNALS of the American Academy of Political and Social Science</source>
          <volume>407</volume>
          (
          <issue>1</issue>
          ),
          <volume>179</volume>
          {
          <issue>180</issue>
          (5
          <year>1973</year>
          ). https://doi.org/10.1177/000271627340700115, http://journals.sagepub.com/doi/10.1177/000271627340700115
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kairam</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heer</surname>
          </string-name>
          , J.: Parting Crowds:
          <article-title>Characterizing Divergent Interpretations in Crowdsourced Annotation Tasks</article-title>
          .
          <source>In: Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work &amp; Social Computing - CSCW '16</source>
          . pp.
          <volume>1635</volume>
          {
          <fpage>1646</fpage>
          . ACM Press, New York, New York, USA (
          <year>2016</year>
          ). https://doi.org/10.1145/2818048.2820016, http://dl.acm.org/citation.cfm?doid=
          <volume>2818048</volume>
          .
          <fpage>2820016</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kiesler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sproull</surname>
          </string-name>
          , L.:
          <article-title>Group decision making and communication technology</article-title>
          .
          <source>Organizational Behavior and Human Decision Processes</source>
          <volume>52</volume>
          (
          <issue>1</issue>
          ),
          <volume>96</volume>
          {
          <issue>123</issue>
          (6
          <year>1992</year>
          ). https://doi.org/10.1016/
          <fpage>0749</fpage>
          -
          <lpage>5978</lpage>
          (
          <issue>92</issue>
          )
          <fpage>90047</fpage>
          -B, http://linkinghub.elsevier.com/retrieve/pii/074959789290047B
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>EcoMeal: A Smart Tray for Promoting Healthy Dietary Habits</article-title>
          .
          <source>In: Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems - CHI EA '16</source>
          . pp.
          <volume>2165</volume>
          {
          <fpage>2170</fpage>
          . ACM Press, New York, New York, USA (
          <year>2016</year>
          ). https://doi.org/10.1145/2851581.2892310, http://dl.acm.org/citation.cfm?doid=
          <volume>2851581</volume>
          .
          <fpage>2892310</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gulshan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahimy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Widner</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Webster</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Grader Variability and the Importance of Reference Standards for Evaluating Machine Learning Models for Diabetic Retinopathy</article-title>
          .
          <source>Ophthalmology (3</source>
          <year>2018</year>
          ). https://doi.org/10.1016/j.ophtha.
          <year>2018</year>
          .
          <volume>01</volume>
          .034, http://arxiv.org/abs/1710.01711 http://linkinghub.elsevier.com/retrieve/pii/S0161642017326982
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mitry</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zutis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dhillon</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peto</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khaw</surname>
          </string-name>
          , K.T.,
          <string-name>
            <surname>Morgan</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moncur</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trucco</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>The Accuracy and Reliability of Crowdsource Annotations of Digital Retinal Images</article-title>
          .
          <source>Translational Vision Science &amp; Technology</source>
          <volume>5</volume>
          (
          <issue>5</issue>
          ),
          <volume>6</volume>
          (
          <year>2016</year>
          ). https://doi.org/10.1167/tvst.5.
          <issue>5</issue>
          .6, http://tvst.arvojournals.org/article.aspx?doi=10.1167/tvst.5.
          <issue>5</issue>
          .
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mumpower</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            ,
            <given-names>T.R.</given-names>
          </string-name>
          :
          <source>Expert Judgement and Expert Disagreement. Thinking &amp; Reasoning</source>
          <volume>2</volume>
          (
          <issue>2-3</issue>
          ),
          <volume>191</volume>
          {
          <issue>212</issue>
          (7
          <year>1996</year>
          ). https://doi.org/10.1080/135467896394500, https://www.tandfonline.com/doi/full/10.1080/135467896394500
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Navajas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niella</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garbulsky</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bahrami</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sigman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Aggregated knowledge from a small number of debates outperforms the wisdom of large crowds</article-title>
          .
          <source>Nature Human Behaviour (1</source>
          <year>2018</year>
          ). https://doi.org/10.1038/s41562-017-0273-4, http://www.nature.com/articles/s41562-017-0273-4
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Penzel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fietze</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Inter-scorer reliability between sleep centers can teach us what to improve in the scoring rules</article-title>
          .
          <source>Journal of Clinical Sleep Medicine</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <volume>81</volume>
          {
          <fpage>87</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rajpurkar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hannun</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haghpanahi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bourn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
          </string-name>
          , A.Y.:
          <article-title>CardiologistLevel Arrhythmia Detection with Convolutional Neural Networks (7</article-title>
          <year>2017</year>
          ), http://arxiv.org/abs/1707.01836
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rosenberg</surname>
            , R.
            <given-names>S.</given-names>
            , van Hout, S.
          </string-name>
          :
          <article-title>The American Academy of Sleep Medicine Inter-scorer Reliability Program: Sleep Stage Scoring</article-title>
          .
          <source>Journal of Clinical Sleep Medicine (1</source>
          <year>2013</year>
          ). https://doi.org/10.5664/jcsm.2350, http://www.aasmnet.org/jcsm/ViewAbstract.aspx?pid=
          <fpage>28772</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Saper</surname>
            ,
            <given-names>C.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuller</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>N.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scammell</surname>
          </string-name>
          , T.E.: Sleep State Switching.
          <source>Neuron</source>
          <volume>68</volume>
          (
          <issue>6</issue>
          ),
          <volume>1023</volume>
          {
          <volume>1042</volume>
          (12
          <year>2010</year>
          ). https://doi.org/10.1016/j.neuron.
          <year>2010</year>
          .
          <volume>11</volume>
          .032, http://linkinghub.elsevier.com/retrieve/pii/S0896627310009748
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Schaekermann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Law</surname>
          </string-name>
          , E.:
          <article-title>Resolvable vs. Irresolvable Disagreement: A Study on Worker Deliberation in Crowd Work</article-title>
          .
          <source>In: Proceedings of the 2018 ACM Conference on Computer Supported Cooperative Work and Social Computing (CSCW'18)</source>
          . New York City, NY (
          <year>2018</year>
          ). https://doi.org/10.1145/3274423
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Solomon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Groupthink versus The Wisdom of Crowds : The Social Epistemology of Deliberation and Dissent</article-title>
          .
          <source>The Southern Journal of Philosophy</source>
          <volume>44</volume>
          (
          <issue>S1</issue>
          ),
          <volume>28</volume>
          {
          <issue>42</issue>
          (3
          <year>2006</year>
          ). https://doi.org/10.1111/j.2041-
          <fpage>6962</fpage>
          .
          <year>2006</year>
          .tb00028.x, http://doi.wiley.
          <source>com/10</source>
          .1111/j.2041-
          <fpage>6962</fpage>
          .
          <year>2006</year>
          .tb00028.x
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Solomon</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The social epistemology of NIH consensus conferences</article-title>
          .
          <source>In: Establishing medical reality</source>
          , pp.
          <volume>167</volume>
          {
          <fpage>177</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Warby</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wendt</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welinder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munk</surname>
            ,
            <given-names>E.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorensen</surname>
            ,
            <given-names>H.B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jennum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peppard</surname>
            ,
            <given-names>P.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perona</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mignot</surname>
          </string-name>
          , E.:
          <article-title>Sleep-spindle detection: crowdsourcing and evaluating performance of experts, non-experts and automated methods</article-title>
          .
          <source>Nature Methods</source>
          <volume>11</volume>
          (
          <issue>4</issue>
          ),
          <volume>385</volume>
          {
          <issue>392</issue>
          (2
          <year>2014</year>
          ). https://doi.org/10.1038/nmeth.2855, http://www.nature.com/doi nder/10.1038/nmeth.2855
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Wright</surname>
            Jr,
            <given-names>K.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badia</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wauquier</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Topographical and temporal patterns of brain activity during the transition from wakefulness to sleep</article-title>
          .
          <source>Sleep</source>
          <volume>18</volume>
          (
          <issue>10</issue>
          ),
          <volume>880</volume>
          {
          <fpage>889</fpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>