<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Human Computation Must Be Reproducible</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Praveen Paritosh Google 345 Spear St</institution>
          ,
          <addr-line>San Francisco, CA 94105</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Human computation is the technique of performing a computational process by outsourcing some of the difficult-toautomate steps to humans. In the social and behavioral sciences, when using humans as measuring instruments, reproducibility guides the design and evaluation of experiments. We argue that human computation has similar properties, and that the results of human computation must be reproducible, in the least, in order to be informative. We might additionally require the results of human computation to have high validity or high utility, but the results must be reproducible in order to measure the validity or utility to a degree better than chance. Additionally, a focus on reproducibility has implications for design of task and instructions, as well as for the communication of the results. It is humbling how often the initial understanding of the task and guidelines turns out to lack reproducibility. We suggest ensuring, measuring and communicating reproducibility of human computation tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Some examples of tasks using human computation are:
labeling images [Nowak and Ruger, 2010], conducting user
studies [Kittur, Chi, and Suh, 2008], annotating natural
language corpora [Snow, O’Connor, Jurafsky and Ng, 2008],
annotating images for computer vision research [Sorokin and
Forsyth, 2008], search engine evaluation [Alonso, Rose and
Stewart, 2008; Alonso, Kazai and Mizzaro, 2011], content
moderation [Ipeirotis, Provost and Wang, 2010], entity
reconciliation [Kochhar, Mazzocchi and Paritosh, 2010],
conducting behavioral studies [Suri and Mason, 2010; Horton,
Rand and Zeckhauser, 2010].</p>
      <p>These tasks involve presenting a question, e.g.,“Is this
image offensive?,” to one or more humans, whose answers are
aggregated to produce a resolution, a suggested answer for
the original question. The humans might be paid
contributors. Examples of paid workforces include Amazon
Mechanical Turk [www.mturk.com] and oDesk [www.odesk.com]. An
example of a community of volunteer contributors is Foldit
[www.fold.it], where human computation augments machine
computation for predicting protein structure [Cooper et al.,
2010]. Another example involving volunteer contributors is
games with a purpose [von Ahn, 2006], and the upcoming
Duolingo [www.duolingo.com], where the contributors are
translate previously untranslated web corpora while
learning a new language.</p>
      <p>The results of human computation can be characterized
by accuracy [Oleson et al., 2011], information theoretic
meaCopyright c 2012 for the individual papers by the papers’ authors.
Copying permitted for private and academic purposes. This volume is published
and copyrighted by its editors.</p>
      <p>CrowdSearch 2012 workshop at WWW 2012, Lyon, France
sures of quality [Ipeirotis, Provost and Wang, 2010], utility
[Dai, Mausam and Weld, 2010], among others. In order for
us to have confidence in any such criteria, the results must
be reproducible, i.e., not a result of chance agreement or
irreproducible human idiosyncrasies, but a reflection of the
underlying properties of the questions and task instructions, on
which others could agree as well. Reproducibility is the
degree to which a process can be replicated by different human
contributors working under varying conditions, at different
locations, or using different but functionally equivalent
measuring instruments. A total lack of reproducibility implies
that the given results could have been obtained merely by
chance agreement. If the results are not differentiable from
chance, there is little information content in them. Using
human computation in such a scenario is wasteful of an
expensive resource, as chance is cheap to simulate.</p>
      <p>A much stronger claim than reproducibility is validity. For
a measurement instrument, e.g., a vernier caliper, a
standardized test, or, a human coder, the reproducibility is the
extent to which a measurement gives consistent results, and
the validity is the extent to which the tool measures what
it claims to measure. In contrast to reproducibility,
validity concerns truths. Validity requires comparing the results
of the study to evidence obtained independently of that
effort. Reproducibility provides assurances that particular
research results can be duplicated, that no (or only a
negligible amount of) extraneous noise has entered the process and
polluted the data or perturbed the research results, validity
provides assurances that claims emerging from the research
are borne out in fact.</p>
      <p>We might want the results of human computation to have
high validity, high utility, low cost, among other desirable
characteristics. However, the results must be reproducible
in order for us to measure the validity or utility to a degree
better than chance.</p>
      <p>More than a statistic, a focus on reproducibility offers
valuable insights regarding the design of the task and the
guidelines for the human contributors, as well as the
communication of the results. The output of human computation is
thus akin to the result of a scientific experiment, and it can
only be considered meaningful if it is reproducible — that is,
the same results could be replicated in an independent
exercise. This requires clearly communicating the task
instructions, and the criterion of selecting the human contributors,
ensuring that they work independently, and reporting an
appropriate measure of reproducibility. Much of this is well
established in the methodology of content analysis in the
social and behavioral sciences [Armstrong, Gosling, Weinman
and Marteau, 1997; Hayes and Krippendorff, 2007], being
required of any publishable result involving human
contributors. In Section 2 and 3, we argue that human computation
resembles the coding tasks of behavioral sciences. However,
in the human computation and crowdsourcing research
community, reproducibility is not commonly reported.</p>
      <p>We have collected millions of human judgments
regarding entities and facts in Freebase [Kochhar, Mazzocchi and
Paritosh, 2010; Paritosh and Taylor, 2012]. We have found
reproducibility to be a useful guide for task and guideline
design. It is humbling how often the initial understanding
of the task and guidelines turns out to lack reproducibility.
In section 4, we describe some of the widely used measures of
reproducibility. We suggest ensuring, measuring and
communicating reproducibility of human computation tasks.</p>
      <p>In the next section, we describe the sources of
variability human computation, which highlight the role of
reproducibility.</p>
    </sec>
    <sec id="sec-2">
      <title>SOURCES OF VARIABILITY IN HUMAN</title>
    </sec>
    <sec id="sec-3">
      <title>COMPUTATION</title>
      <p>There are many sources of variability in human
computation that are not present in machine computation. Given
that human computation is used to solve problems that are
beyond the reach of machine computation, by definition,
these problems are incompletely specified. Variability arises
due to incomplete specification of the task. This is convolved
with the fact that the guidelines are subject to differing
interpretations by different human contributors. Some
characteristics of human computation tasks are:
• Task guidelines are incomplete: A task can span a wide
set of domains, not all of which are anticipated at the
beginning of the task. This leads to incompleteness
in guidelines, as well as varying levels of performance
depending upon the contributor’s expertise in that
domain. Consider the task of establishing relevance of
an arbitrary search query [Alonso, Kazai and Mizzaro,
2011].
• Task guidelines are not precise: Consider, for
example, the task of declaring if an image is unsuitable for
a social network website. Not only is it hard to write
down all the factors that go into making an image
offensive, it is hard to communicate those factors to
human contributors with vastly different predispositions.
The guidelines usually rely upon shared common sense
knowledge and cultural knowledge.
• Validity data is expensive or unavailable: An oracle
that provide the true answer for any given question is
usually unavailable. Sometimes for a small subset of
gold questions, we have answers from another
independent source. This can be useful in making estimates of
validity, subject to the degree that the gold questions
are representative of the set of questions. These gold
questions could be very useful for training and
feedback to the human contributors, however, we have to
be ensure their representatitiveness in order to make
warranted claims regarding validity.</p>
      <p>Each of the above might be true to a different degree for
different human computation tasks. These factors are
similar to the concerns of behavioral and social scientists in using
humans as measuring instruments.
3.</p>
    </sec>
    <sec id="sec-4">
      <title>CONTENT ANALYSIS AND CODING IN</title>
    </sec>
    <sec id="sec-5">
      <title>THE BEHAVIORAL SCIENCES</title>
      <p>In the social sciences, content analysis is a methodology
for studying the content of communication [Berelson, 1952;
Krippendorff, 2004]. Coding of subjective information is a
significant source of empirical data in social and behavioral
sciences, as they allow techniques of quantitative research
to be applied to complex phenomena. These data are
typically generated by trained human observers who record or
transcribe textual, pictorial or audible matter in terms
suitable for analysis. This task is called coding, which involves
assigning categorical, ordinal or quantitative responses to
units of communication.</p>
      <p>An early example of a content analysis based study is “Do
newspapers now give the news?” [Speed, 1893], which tried
to show that the coverage of religious, scientific and literary
matters was dropped in favor of gossip, sports and scandals
between 1881 and 1893, by New York newspapers.
Conclusions from such data can only be trusted if the reading of
the textual data as well as of the research results are
replicable elsewhere, that the coders demonstrably agree on what
they are talking about. Hence, the coders need to
demonstrate the trustworthiness of their data by measuring their
reproducibility. To perform reproducibility tests, additional
data are needed: by duplicating the research under various
conditions. Reproducibility is established by independent
agreement between different but functionally equal
measuring devices, for example, by using several coders with
diverse personalities. The reproducibility of coding has been
used for comparing consistency of medical diagnosis [e.g.,
Koran, 1975], for drawing conclusions from meta-analysis of
research findings [e.g., Morley et al., 1999], for testing
industrial reliability [Meeker and Escobar, 1998], for establishing
the usefulness of a clinical scale [Hughes et al., 1982].
3.1</p>
    </sec>
    <sec id="sec-6">
      <title>Relationship between Reproducibility and Validity</title>
      <p>• Lack of reproducibility limits the chance of validity: If
the coding results are a product of chance, it may well
include a valid account of what was observed, but
researchers would not be able to identify that account
to a degree better than chance. Thus, the more
unreliable a procedure, the less likely it is to result in data
that lead to valid conclusions.
• Reproducibility does not guarantee validity: Two
observers of the same event who hold the same
conceptual system, prejudice, or, interest may well agree on
what they see but still be objectively wrong, based on
some external criterion. Thus a reliable process may
or may not lead to valid outcomes.</p>
      <p>In some cases, validity data might be so hard to obtain
that one has to contend with reproducibility. In tasks such
as interpretation and transcription of complex textual
matter, suitable accuracy standards are not easy to find.
Because interpretations can only be compared to
interpretations, attempts to measure validity presuppose the
privileging of some interpretations over others, and this puts any
claims regarding validity on epistemologically shaky grounds.
In some tasks like psychiatric diagnosis, even
reproducibility is hard to attain for some questions. Aboraya et al.
[2006] review the reproducibility of psychiatric diagnosis.
Lack of reproducibility has been reported for judgments of
schizophrenia and affective disorder [Goodman et al., 1984],
calling such diagnosis into question.
3.2</p>
    </sec>
    <sec id="sec-7">
      <title>Relationship with Chance Agreement</title>
      <p>In this section, we look at some properties of chance
agreement, and its relationship to reproducibility. Given two
coding schemes for the same phenomenon, the one with fewer
categories will have higher chance agreement. For example,
in reCAPTCHA [von Ahn et al., 2008], two independent
humans are shown an image containing text and asked to
transcribe it. Assuming that there is no collusion, chance
agreement, i.e., two different humans typing in the same
word/phrase by chance is very small. However, in a task
in which there are only two possible answers, e.g., true and
false, the probability of chance agreement between two
answers is 0.5.</p>
      <p>If a disproportionate amount of data falls under one
category, then the expected chance agreement is very high, so in
order to demonstrate high reproducibility, even higher
observed agreement is required [Feinstein and Cicchetti 1990;
Di Eugenio and Glass 2004].</p>
      <p>Consider a task of rating a proposition as true or false.
Let p be the probability of the proposition being true. An
implementation of chance agreement is the following: toss a
biased coin with the same odds, i.e., p is the probability that
it turns heads, and declare the proposition to be true when
the coin lands heads. Now, we can simulate n judgments
by tossing the coin n times. Let us look at the properties
of unanimous agreement between two judgments. The
likelihood of a chance agreement on true is p2. Independent
of this agreement, the probability of the proposition being
true is p, therefore, the accuracy of this chance agreement
on true is p3. By design, these judgments do not contain any
information other than the a priori distribution across
answers. Such data has close to zero reproducibility, however,
sometimes it can show up in surprising ways when looked
through the lens of accuracy.</p>
      <p>For example, consider the task of an airport agent
declaring a bag as safe or unsafe for boarding on the plane. A
bag can be unsafe if it contains toxic or explosive materials
that could threaten the safety of the flight. Most bags are
safe. Let us say that one in a thousand bags is potentially
unsafe. Random coding would allow two agents to jointly
assign “safe” 99.8% of the time, and since 99.9% of the bags
are safe, this agreement would be accurate 99.7% of the time!
This leads to the surprising result that when data are highly
skewed, the coders may agree on a high proportion of items
while producing annotations that are accurate, but of low
reproducibility. When one category is very common, high
accuracy and high agreement can also result from
indiscriminate coding. The test for reproducibility in such cases is
the ability to agree on the rare categories. In the airport
bag classification problem, while chance accuracy on safe
bags is high, chance accuracy on unsafe bags is extremely
low, 10−7%. In practice, the cost of errors vary: mistakenly
classifying a safe bag as unsafe causes far less damage than
classifying an unsafe bag as safe.</p>
      <p>In this case, it is dangerous to consider an averaged
accuracy score, as different errors do not count equal: a chance
process that does not add any information has an average
accuracy which is higher than 99.7%, most of which is
reflecting the original bias in the distribution of safe and unsafe
bags. A misguided interpretation of accuracy or a poor
estimate of accuracy can be less informative than
reproducibility.
3.3</p>
    </sec>
    <sec id="sec-8">
      <title>Reproducibility and Experiment Design</title>
      <p>The focus on reproducibility has implications on design of
task and instruction materials. Krippendorff [2004] argues
that any study using observed agreement as a measure of
reproducibility must satisfy the following requirements:
• It must employ an exhaustively formulated, clear, and
usable guidelines;
• It must use clearly specified criteria concerning the
choice of contributors,so as others may use such
criteria to reproduce the data;
• It must ensure that the contributors that generate the
data used to measure reproducibility work
independently of each other. Only if such independence is
assured can covert consensus be ruled out the observed
agreement be explained in terms of the given guidelines
and the task.</p>
      <p>The last point cannot be stressed enough. There are
potential benefits from multiple contributors collaborating, but
data generated in this manner neither ensure
reproducibility nor reveal its extent. In groups like these, humans are
known to negotiate and to yield to each other in quid pro
quo exchanges, with prestigious group members dominating
the outcome [see for example, Esser, 1998]. This makes the
results of collaborative human computation a reflection of
the social structure of the group, which is nearly impossible
to communicate to other researchers and replicate. The data
generated by collaborative work are akin to data generated
by a single observer, while reproducibility requires at least
two independent observers. To substantiate the contention
that collaborative coding is superior to coding by separate
individuals, a researcher would have to compare the data
generated by at least two such groups and two individuals,
each working independently.</p>
      <p>A model in which the coders work independently, but
consult each other when unanticipated problems arise, is also
problematic. A key source of these unanticipated problems
is the fact that the writers of the coding instructions did
not anticipate all the possible ways of expressing the
relevant matter. Ideally, these instructions should include every
applicable rule on which agreement is being measured.
However, discussing emerging problems could create re-interpretation
of the existing instructions in ways that are a function of the
group and not communicable to others. In addition, as the
instructions become reinterpreted, the process loses its
stability: data generated early in the process use instructions
that differ from those later.</p>
      <p>In addition to the above, Craggs and McGee Wood [2005]
discourage researchers from testing their coding instructions
on data from more than one domain. Given that the
reproducibility of the coding instructions depends to a great
extent on how complications are dealt with, and that every
domain displays different complications, the sample should
contain sufficient examples from all domains which have to
be annotated according to the instructions.</p>
      <p>Even the best coding instructions might not specify all
possible complications. Besides the set of desired answers,
the coders should also be allowed to skip a question. If the
coders cannot prove any of the other answers is correct, they
skip that question. For any other answer, the instructions
define an a priori model of agreement on that answer, while
skip represents the unanticipated properties of questions and
coders. For instance, some questions might be too difficult
for certain coders. Providing the human contributors with
an option to skip is a nod to the openness of the task, and
can be used to explore the poorly defined parts of the task
that were not anticipated at the outset. Additionally, the
skip votes can be removed from the analysis for computing
reproducibility, as we do not have expectation of agreement
on them [Krippendorff, 2012, personal communication].</p>
    </sec>
    <sec id="sec-9">
      <title>MEASURING REPRODUCIBILITY</title>
      <p>In measurement theory, reliability is the more general
guarantee that the data obtained are independent of the
measuring event, instrument or person. There are three
different kinds of reliability:</p>
      <p>Stability: measures the degree to which a process is
unchanging over time. It is measured by agreement between
multiple trials of the same measuring or coding process. This
is also called test-retest condition, in which one observer
does a task, and after some time, repeats the task again.
This measures intra-observer reliability. A similar notion is
internal consistency [Cronbach, 1951], which is the degree
to which the answers on the same task are consistent.
Surveys are designed so that the subsets of similar questions
are known a priori, and measures for internal consistency
metrics are based on correlation between these answers.</p>
      <p>Reproducibility: measures the degree to which a process
can be replicated by different analysts working under varying
conditions, at different locations, or, using different but
functionally equivalent measuring instruments. Reproducible
data, by definition, are data that remain constant
throughout variations in the measuring process [Kaplan and
Goldsen, 1965].</p>
      <p>Accuracy: measures the degree to which the process
produces valid results. To measure accuracy, we have to
compare the performance of contributors with the performance
of a procedure that is known to be correct. In order to
generate estimates of accuracy, we need accuracy data, i.e.,
valid answers to a representative sample of the questions.
Estimating accuracy gets harder in cases where the
heterogeneity of the task is poorly understood.</p>
      <p>The next section focuses on reproducibility.
4.1</p>
    </sec>
    <sec id="sec-10">
      <title>Reproducibility</title>
      <p>There are two different aspects of reproducibility:
interrater reliability and inter-method reliability. Inter-rater
reliability focuses on the reproducibility by agreement between
independent raters, and inter-method reliability focuses on
the reliability of different measuring devices. For example,
in survey and test design, parallel forms reliability is used
to create multiple equivalent tests, of which more than one
are administered to the same human. We focus on
interrater reliability as the measure of reproducibility typically
applicable to human computation tasks, where we generate
judgments from multiple humans per question. The simplest
form of inter-rater reliability is percent agreement, however
it is not suitable as a measure of reproducibility as it does
not correct for chance agreement.</p>
      <p>For extensive survey of measures of reproducibility, refer
to Popping [1988], Artstein and Poesio [2007]. The different
coefficients of reproducibility differ in the assumptions they
make about the properties of coders, judgments and units.
Scott’s π [1955] is applicable to two raters and assumes that
the raters have the same distribution of responses, where
Cohen’s κ [1960; 1968] allows for a a separate distribution
of chance behavior per coder. Fleiss’ κ [1971] is a
generalization of Scott’s π for an arbitrary number of raters.
All of these coefficients of reproducibility correct for chance
agreement similarly. First, they find how much agreement
is expected by chance: let us call this vallue Ae. The data
from the coding is a measure of the observed agreement,
Ao. Various inter-rater reliabilities measure the proportion
of the possible agreement beyond chance that was actually
observed.</p>
      <p>S, π, κ =</p>
      <p>Ao − Ae
1 − Ae</p>
      <p>Krippendorff’s α [1970; 2004] is a generalization of many
of these coefficients. It is a generalization of the above
metrics for an arbitrary number of raters, not all of whom have
to answer every question. Krippendorff’s α has the following
desirable characteristics:
• It is applicable to an arbitrary number of contributors
and invariant to the permutation and selective
participation of contributors. It corrects itself for varying
amounts of reproducibility data.
• It constitutes a numerical scale between at least two
points with sensible reproducibility interpretations, 0
representing absence of agreement, and 1 indicating
perfect agreement.
• It is applicable to several scales of measurement:
ordinal, nominal, interval, ratio, and more.</p>
      <sec id="sec-10-1">
        <title>Alpha’s general form is:</title>
      </sec>
      <sec id="sec-10-2">
        <title>Where Do is the observed disagreement: α = 1 −</title>
        <p>Do</p>
        <p>De
Do = 1 X X ock.δc2k</p>
        <p>n c k
and De is the disagreement one would expect when the
answers are attributable to chance rather than to the
properties of the questions:</p>
        <p>De =
1</p>
        <p>2</p>
        <p>X X nc.nk.δck
n(n − 1) c k</p>
        <p>The δc2k term is the distance metric for the scale of the
answerspace. For a nominal scale,
2
δck =
(
0 if c = k
1 if c 6= k
4.2</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Statistical Significance</title>
      <p>The goal of measuring reproducibility is to ensure that
the data does not deviate too much from perfect agreement,
not that the data is different from chance agreement. In the
definition of α, chance agreement is one of the two anchors
for the agreement scale, the other, more important, reference
point being that of perfect agreement. As the distribution of
α is unknown, confidence intervals on α are obtained from
empirical distribution generated by bootstrapping — that
is, by drawing a large number of subsamples from the
reproducibility data, computing α for each. This gives us a
probability distribution of hypothetical α values that could
occur within the constraints of the observed data. This can
be used to calculate the probability of failing to reach the
smallest acceptable reproducibility αmin, q|α &lt; αmin, or a
two tailed confidence interval for chosen level of significance.
4.3</p>
    </sec>
    <sec id="sec-12">
      <title>Sampling Considerations</title>
      <p>To generate an estimate of the reproducibility of a
population, we need to generate a representative sample of the
population. The sampling needs to ensure that we have
enough units from the rare categories of questions in the
data. Assuming that α is normally distributed, Bloch and
Kraemer [1989] provide a suggestion for minimum number of
questions from each category to be included in the sample,
Nc, by,</p>
      <p>Nc = z
2 (1 + αmin)(3 − αmin)
4pc(1 − pc)(1 − αmin)
− αmin
Where,
• pc is the smallest estimated proportion values of the
category c in the population,
• alphamin is the smallest acceptable reproducibility
below which data will have to be rejected as unreliable,
and
• z is the desired level of statistical significance,
represented by the corresponding z value for one-tailed tests
This is a simplification, as it assumes α is normally
distributed, and binary data, and does not account for the
number of raters. A general description of sampling requirement
is an open problem.
4.4</p>
    </sec>
    <sec id="sec-13">
      <title>Acceptable Levels of Reproducibility</title>
      <p>Fleiss [1981] and Krippendorff [2004] present guidelines
for what should acceptable values of reproducibility based
on surveying the empirical research using these measures.
Krippendorff suggests,
• Rely on variables with reproducibility above α = 0.800.</p>
      <p>Additionally don’t accept data if the confidence
interval reaches below the smallest acceptable
reproducibility, αmin = 0.667, or, ensure that the probability, q, of
the failure to have less than smallest acceptable
reproducibility alphamin is reasonably small, e.g., q &lt; 0.05.
• Consider variables with reproducibility between α =
0.667 and α = 0.800 only for drawing tentative
conclusions.</p>
      <p>These are suggestions, and the choice of thresholds of
acceptability depend upon the validity requirements imposed
on the research results. It is perilous to “game” α by
violating the requirements of reproducibility: for example, by
removing a subset of data post-facto to increase α.
Partitioning data by agreement measured after the experiment
will not lead to valid conclusions.
4.5</p>
    </sec>
    <sec id="sec-14">
      <title>Other Measures of Quality</title>
      <p>Ipeiritos, Provost and Wang [2010] present an information
theoretic quality score, which measures the quality of a
contributor in terms of comparing their score to a spammer who
is trying to advantage of chance accuracy. In that regard, it
is a similar metric to Krippendorff’s alpha, and additionally
models a notion of cost.</p>
      <p>QualityScore = 1 −</p>
      <p>ExpCost(Contributor)</p>
      <p>ExpCost(Spammer)</p>
      <p>Turkontrol [Dai, Mausam and Weld, 2010], uses both a
model of utility and quality using decision-theoretic control
to make trade-offs between quality and utility for workflow
control.</p>
      <p>Le, Edmonds, Hester and Biewald [2010] develop a gold
standard based quality assurance framework that provides
direct feedback to the workers and targets specific worker
errors. This approach requires extensive manually generated
collection of gold data. Oleson et al. [2011], further develop
this approach to include pyrite, which are programmatically
generated gold questions on which contributors are likely to
make an error, for example by mutating data so that it is
no longer valid. These are very useful metrics for training,
feedback and protection against spammers, but these do not
reveal the accuracy of the results. The gold questions, by
design, are not representative of the original set of questions.
These lead to wide error bars on the accuracy estimates, and
it might be valuable to measure reproducibility of results.
5.</p>
    </sec>
    <sec id="sec-15">
      <title>CONCLUSIONS</title>
      <p>We describe reproducibility as a necessary but not
sufficient requirement for results of human computation. We
might additionally require the results to have high validity
or high utility, but our ability to measure validity or utility
with confidence is limited if the data are not reproducible.
Additionally, a focus on reproducibility has implications for
design of task and instructions, as well as for the
communication of the results. It is humbling how often the initial
understanding of the task and guidelines turns out to lack
reproducibility. We suggest ensuring, measuring and
communicating reproducibility of human computation tasks.
6.</p>
    </sec>
    <sec id="sec-16">
      <title>ACKNOWLEDGMENTS</title>
      <p>The author would like to thank Jamie Taylor, Reilly Hayes,
Stefano Mazzocchi, Mike Shwe, Ben Hutchinson, Al Marks,
Tamsyn Waterhouse, Peng Dai, Ozymandias Haynes, Ed
Chi and Panos Ipeirotis for insightful discussions on this
topic.
7.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Aboraya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rankin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , France,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>El-Missiry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>John</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>The Reliability of Psychiatric Diagnosis Revisited</article-title>
          ,
          <source>Psychiatry (Edgmont)</source>
          .
          <source>2006 January; 3</source>
          (
          <issue>1</issue>
          ):
          <fpage>41</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Alonso</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rose</surname>
            ,
            <given-names>D. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            ,
            <given-names>B..</given-names>
          </string-name>
          <article-title>Crowdsourcing for relevance evaluation</article-title>
          .
          <source>SIGIR Forum</source>
          ,
          <volume>42</volume>
          (
          <issue>2</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Alonso</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazai</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2011</year>
          ,
          <article-title>Crowdsourcing for search engine evaluation</article-title>
          ,
          <year>Springer 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Armstrong</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gosling</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Marteau</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <year>1997</year>
          ,
          <article-title>The place of inter-rater reliability in qualitative research: An empirical study</article-title>
          ,
          <source>Sociology</source>
          , vol
          <volume>31</volume>
          , no.
          <issue>3</issue>
          ,
          <fpage>597</fpage>
          -
          <lpage>606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Artstein</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Inter-coder agreement for computational linguistics</article-title>
          .
          <source>Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Bennett</surname>
            ,
            <given-names>E. M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alpert</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Goldstein</surname>
          </string-name>
          .
          <year>1954</year>
          .
          <article-title>Communications through limited questioning</article-title>
          .
          <source>Public Opinion Quarterly</source>
          ,
          <volume>18</volume>
          (
          <issue>3</issue>
          ):
          <fpage>303</fpage>
          -
          <lpage>308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Berelson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>1952</year>
          .
          <article-title>Content analysis in communication research</article-title>
          , Free Press, New York.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] Bloch, Daniel A. and Helena Chmura Kraemer.
          <year>1989</year>
          .
          <article-title>2 x 2 kappa coefficients: Measures of agreement or association</article-title>
          .
          <source>Biometrics</source>
          ,
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>269</fpage>
          -
          <lpage>287</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.,
          <year>2008</year>
          , Freebase:
          <string-name>
            <given-names>A Collaboratively</given-names>
            <surname>Created</surname>
          </string-name>
          <article-title>Graph Database for Structuring Human Knowledge</article-title>
          .
          <source>In the Proceedings of the 28th ACM SIGMOD International Conference on Management of Data</source>
          , Vancouver.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          , Jacob.
          <year>1960</year>
          .
          <article-title>A coefficient of agreement for nominal scales</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          , Jacob.
          <year>1968</year>
          .
          <article-title>Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit</article-title>
          .
          <source>Psychological Bulletin</source>
          ,
          <volume>70</volume>
          (
          <issue>4</issue>
          ):
          <fpage>213</fpage>
          -
          <lpage>220</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Cooper</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khatib</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Treuille</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barbero</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beenen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leaver-Fay</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Popovic</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <article-title>Predicting protein structures with a multiplayer online game</article-title>
          .
          <source>Nature</source>
          ,
          <volume>466</volume>
          :
          <fpage>7307</fpage>
          ,
          <fpage>756</fpage>
          -
          <lpage>760</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Cronbach</surname>
            ,
            <given-names>L. J.</given-names>
          </string-name>
          ,
          <year>1951</year>
          ,
          <article-title>Coefficient alpha and the internal structure of tests</article-title>
          .
          <source>Psychometrica</source>
          ,
          <volume>16</volume>
          ,
          <fpage>297</fpage>
          -
          <lpage>334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Craggs</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>McGee</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2004</year>
          .
          <article-title>A two-dimensional annotation scheme for emotion in dialogue</article-title>
          .
          <source>In Proc.of AAAI Spring Symmposium on Exploring Attitude and Affect in Text</source>
          , Stanford.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mausam</surname>
            , Weld,
            <given-names>D. S.</given-names>
          </string-name>
          <article-title>Decision-Theoretic Control of Crowd-Sourced Workflows</article-title>
          .
          <source>AAAI</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Di</surname>
            <given-names>Eugenio</given-names>
          </string-name>
          , Barbara and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Glass</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The kappa statistic: A second look</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>30</volume>
          (
          <issue>1</issue>
          ):
          <fpage>95</fpage>
          -
          <lpage>101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Esser</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          ,
          <year>1998</year>
          ,
          <article-title>Alive and Well after 25 Years: A Review of Groupthink Research, Organizational Behavior and Human Decision Processes</article-title>
          , Volume
          <volume>73</volume>
          ,
          <string-name>
            <surname>Issues</surname>
          </string-name>
          2-
          <issue>3</issue>
          ,
          <fpage>116</fpage>
          -
          <lpage>141</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Feinstein</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alvan R</surname>
          </string-name>
          . and
          <string-name>
            <surname>Domenic</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Cicchetti</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>High agreement but low kappa: I. The problems of two paradoxes</article-title>
          .
          <source>Journal of Clinical Epidemiology</source>
          ,
          <volume>43</volume>
          (
          <issue>6</issue>
          ):
          <fpage>543</fpage>
          -
          <lpage>549</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Fleiss</surname>
            ,
            <given-names>Joseph L.</given-names>
          </string-name>
          <year>1971</year>
          .
          <article-title>Measuring nominal scale agreement among many raters</article-title>
          .
          <source>Psychological Bulletin</source>
          ,
          <volume>76</volume>
          (
          <issue>5</issue>
          ):
          <fpage>378</fpage>
          -
          <lpage>382</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Goodman</surname>
            <given-names>AB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahav</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popper</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ginath</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pearl</surname>
            <given-names>E.</given-names>
          </string-name>
          <article-title>The reliability of psychiatric diagnosis in Israel's Psychiatric Case Register</article-title>
          .
          <source>Acta Psychiatr Scand</source>
          . 1984 May;
          <volume>69</volume>
          (
          <issue>5</issue>
          ):
          <fpage>391</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Krippendorff</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Answering the call for a standard reliability measure for coding data</article-title>
          .
          <source>Communication Methods and Measures</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>77</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Hughes</surname>
            <given-names>CD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danziger</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coben</surname>
            <given-names>LA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            <given-names>RL</given-names>
          </string-name>
          .
          <article-title>A new clinical scale for the staging of dementia</article-title>
          .
          <source>Then British Journal of Psychiatry</source>
          <year>1982</year>
          ;
          <volume>140</volume>
          :
          <fpage>56</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Ipeirotis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Provost</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Quality management on amazon mechanical turk</article-title>
          .
          <source>In KDD-HCOMP '10.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Krippendorff</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>1970</year>
          .
          <article-title>Bivariate agreement coefficients for reliability of data</article-title>
          .
          <source>Sociological Methodology</source>
          ,
          <volume>2</volume>
          :
          <fpage>139</fpage>
          -
          <lpage>150</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Krippendorff</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Content Analysis: An Introduction to Its Methodology, Second edition</article-title>
          . Sage, Thousand Oaks.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Kochhar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mazzocchi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2010</year>
          ,
          <article-title>The Anatomy of a Large-Scale Human Computation Engine</article-title>
          ,
          <source>In Proceedings of Human Computation Workshop at the 16th ACM SIKDD Conference on Knowledge Discovery and Data Mining, KDD 2010</source>
          , Washington D.C.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Koran</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          ,
          <year>1975</year>
          ,
          <article-title>The reliability of clinical methods, data and judgments (parts 1 and 2</article-title>
          ),
          <source>New England Journal of Medicine</source>
          ,
          <volume>293</volume>
          (
          <issue>13</issue>
          /14).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Kittur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chi</surname>
            ,
            <given-names>E. H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Suh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Crowdsourcing user studies with Mechanical Turk. In Proceedings of the Proceeding of the twenty-sixth annual SIGCHI conference on Human factors in computing systems</article-title>
          , Florence.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Mason</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Siddharth</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Conducting Behavioral Research on Amazon's Mechanical Turk</article-title>
          , Working Paper, Social Science Research Network.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Meeker</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Escobar</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <year>1998</year>
          ,
          <article-title>Statistical Methods for Reliability Data</article-title>
          , Wiley.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Morley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Eccleston</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>1999</year>
          ,
          <article-title>Systematic review and meta-analysis of randomized controlled trials of cognitive behavior therapy and behavior therapy for chronic pain in adults, excluding headache</article-title>
          ,
          <source>Pain</source>
          ,
          <volume>80</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ruger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation</article-title>
          .
          <source>In Multimedia Information Retrieval</source>
          , pages
          <fpage>557</fpage>
          -
          <lpage>566</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Oleson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hester</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorokin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laughlin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Biewald</surname>
          </string-name>
          , L. Programmatic Gold:
          <article-title>Targeted and Scalable Quality Assurance in Crowdsourcing</article-title>
          .
          <source>In HCOMP '11: Proceedings of the Third AAAI Human Computation Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.,
          <year>2012</year>
          , Freebase: Curating Knowledge at Scale, To appear
          <source>in the 24th Conference on Innovative Applications of Artificial Intelligence</source>
          ,
          <string-name>
            <surname>IAAI</surname>
          </string-name>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Popping</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <year>1988</year>
          ,
          <article-title>On agreement indices for nominal data</article-title>
          ,
          <source>Sociometric Research: Data Collection and Scaling</source>
          ,
          <volume>90</volume>
          -
          <fpage>105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>W. A.</given-names>
          </string-name>
          <year>1955</year>
          .
          <article-title>Reliability of content analysis: The case of nominal scale coding</article-title>
          .
          <source>Public Opinion Quarterly</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>321</fpage>
          -
          <lpage>325</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Snow</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Jurafsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            , and
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Y.</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>Cheap and fast, but is it good?: evaluating nonexpert annotations for natural language tasks</article-title>
          .
          <source>In EMNLP '08: Proceedings of the Conference on Empirical Methods in Natural Language Processing.</source>
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <surname>Sorokin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Forsyth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Utility data annotation with amazon mechanical turk</article-title>
          .
          <source>In First International Workshop on Internet Vision</source>
          , CVPR
          <volume>08</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Speed</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <year>1893</year>
          ,
          <article-title>Do newspapers now give the news</article-title>
          , Forum,
          <year>August</year>
          ,
          <year>1893</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>von Ahn</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Games With A Purpose</article-title>
          . IEEE Computer Magazine,
          <year>June 2006</year>
          . Pages 96-
          <fpage>98</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>von Ahn</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dabbish</surname>
            ,
            <given-names>L. Labeling</given-names>
          </string-name>
          <article-title>Images with a Computer Game</article-title>
          .
          <source>ACM Conference on Human Factors in Computing Systems, CHI 2004. Pages</source>
          <volume>319</volume>
          -
          <fpage>326</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>von Ahn</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maurer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McMillen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abraham</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Blum</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>reCAPTCHA: Human-Based Character Recognition via Web Security Measures</article-title>
          .
          <source>Science, September</source>
          <volume>12</volume>
          ,
          <year>2008</year>
          . pp
          <fpage>1465</fpage>
          -
          <lpage>1468</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>