<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manfred Klenner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne G o¨hring</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Amsler</string-name>
          <email>mamslerg@cl.uzh.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computational Linguistics University of Zurich</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we argue that harmonization is not the preferred way to produce a gold standard in all cases. Neither does a majority vote based harmonization produce an appropriate gold standard centroid, nor would a mere centroid be a good basis for training a system that reproduces prototypical user reactions given some understanding task. We discuss these claims in the context of sentiment inference.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>It is common practice to harmonize annotated data
produced by a couple of human raters in order to
create a gold standard. The quality of the
annotated data not only depends on the quality of the
annotation guidelines, but also on various personal
traits of the raters (cognitive capacity, reliability,
motivation etc.). We cannot fully control these
parameters, we are even expecting raters to fail and
produce wrong annotations. Hence the need for
harmonization where the right annotation decision
is fixed either on the basis of a discussion among
raters or by (simple) majority vote. This process
of harmonization is suited for all those annotation
tasks that have clear decision boundaries. The
situation changes when the annotation is more based
on subjective understandings and evaluations of
the material (e.g. text). This might even touch
upon personal standards, mental dispositions and
ethic obligations not to mention political stance
or religious premises. To give quite a harmless
example: is the occasional (or even unique) use
of a work phone for private purposes negative?
Clearly, we could force raters to penalize any
misuse, any (white) lies, and misdemeanor.</p>
      <p>Would a gold standard created that way
represent any real opinion or would it be an artifact of
the guidelines.</p>
      <p>In this paper, we argue that for our application,
sentiment inference, we should leave room for
individual perspectives and should avoid
harmonization and accept that the annotation process leaves
us with a distribution of annotations representing a
diversity of opinions. We first introduce the notion
of sentiment inference, then we sketch our
annotation guidelines and discuss lessons learned from
the initial annotations efforts.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Sentiment Inference</title>
      <p>
        Sentiment inference1 is a variant of stance
detection where both the source and the target of an
attitude need to be found and where the pro and con
relations sometimes are only transitively given. In
stance detection, the source of an opinion usually
is the writer of a text and the target is some
controversial topic (e.g. death penalty). However,
any text might discuss sources and targets of
attitudes, the text author just tells us about it and
has or has not an opinion directed towards these
proponents and opponents. For instance, a
statement like Obama criticizes the spread of fake news
on Facebook gives rise to the inference that the
opinion source Obama is against (con) the
target fake news. This is also an example of an
inference, since literally, Obama just is against the
spread, but this implies he is against fake news
as well. The kind of reasoning that gives rise to
con(Obama, fake news) could be captured by the
following inference rule: If A is against B and B is
good for C, then A is against C. Of course, one has
1We have implemented a rule-based
system for German, a demo is available
pub.cl.uzh.ch/demo/stancer/
baseline
under
to accept the idea that it is good for the fake news
that it was spread. Then, we no longer only are
talking about attitude, but also about positive and
negative effects on entities. Effects are defined as
the perceived consequences of an event that
happened (this is similar to the good-for/bad-for
distinction of
        <xref ref-type="bibr" rid="ref1">(Deng et al., 2013)</xref>
        ). Sometimes, there
are only effects, but no attitudes (she wins) and
sometimes attitudes and effects are even somewhat
contrary (He criticizes that she was honored).
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Sentiment Inference Annotation</title>
      <p>Initially, our goal was to produce a traditional gold
standard for verb-based sentiment inference. We
have a lexicon of about 1000 verbs that express
polar relations and for a subset of 100 verbs, we
extracted 500 sentences from a newspaper corpus.
We let 4 raters (experts) annotate them. The
annotators were trained beforehand on 100 sentences
from another newspaper corpus. On that basis, the
initial annotation guidelines were refined as a
result of our discussions of difficult examples and
borderline cases.</p>
      <p>vote(s) # annot.</p>
      <p>4 340
3 325
2 358
1 782</p>
      <p>The distribution of annotations produced is
shown in table 1. It also gives the absolute
frequencies for effect (eff), relations (rel) and actors
(act) (we won’t discuss this last dimension here,
because only a few cases were found).</p>
      <p>In 43.32% of the cases only a single annotator
(last row, 1 vote) produced a particular annotation
in contrast to the other ones. In 19.83% we have 2
annotators that agree, in 18.01% three agree and in
18.84% of the cases all agreed. This is quite a
diverse picture and we used Fleiss’ Kappa to further
quantify it. For the reached interannotator
agreement see table 2.</p>
      <p>The result was an agreement of 11.98%
(overall value), which is considered slight (0.0-0.20 is
slight, 0.21-0.40 is fair). We also measured the
pairwise agreements which are -0.16%, 6.94%,
8.60%, 8.62%,14.85% and 27.02%. Only one
pairing reached a fair agreement (27.02%).
overall
relation
- pro relation
- contra relation
effect
- positive effect
- negative effect
actor:</p>
      <p>The agreement in general is low. Searching
for the reasons, we again discussed our guidelines
(see the next section), but in the end, after we
finished our attempt to harmonize, we opted against
more rigid guidelines and in favor of a more liberal
notion of what is called a gold standard.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Sketch of the Annotation Guidelines</title>
      <p>We distinguish two polar relations (pro, con),
positive and negative effects, and positive as well as
negative actors. According to our guidelines, a pro
relation is directed from a source toward a target,
if there is
1. a positive attitude of an actor toward an
actor/object/situation
2. an action of an actor which yields something
positive for an actor/object/situation
3. an object which yields something positive for
an actor/object/situation
The con relation is defined accordingly.</p>
      <p>We allow pro/con relations to hold between
non-animate discourse referents, thus these
relations are not meant to be interpreted as strict
attitudes where the source must be an (intentional)
opinion bearer. The reason is, among others,
that in order to be able to infer pro/con
relations transitively, non-animate arguments are
useful as a bridge. In The CEO is against the
contract, since the contract is bad for the
company we can only infer that pro(CEO,company)
if con(CEO,contract) and con(contract,company).
But contract is not a well-defined attitude bearer.
This is why we have widened the definition of
pro/con. If a given application needs a stricter
perspective, an additional animacy classifier could be
used to separate cases where the opinion source
is animate - even metonomically given like in
Moscow criticizes Washington - from cases with
non-animate sources like in The snow blocks the
entrance to the hospital.</p>
      <p>The guidelines also state that the pro and con
relations are situation specific, they do not hold in
general. If A criticizes B, then this is true only for
the situation at hand.</p>
      <p>Effects are consequences of the truth of a
particular situation. They hold if the event denoted by a
clause is factual (cf. a positive effect on she given
She won the competition, but not in She might
win). Like attitudes, effects are verb-specific. We
distinguish on a conceptual level (but we do not
annotate it) moral (to accuse), social (to honor),
emotional (to insult) and physical (to hurt) effects.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Disagreement Example</title>
      <p>In order to give an example of the problems we
encountered, we discuss the following (translated
version of a German) sentence:</p>
      <p>Jim Crace, whose books depress many readers,
seemed like the most cheerful man on earth.</p>
      <p>All annotators agreed that there is a negative
effect on ‘readers’ (1). Three see a con relation
between ‘book’ and ‘readers’ (2), one believes in
a con relation between ‘readers’ and ‘Jim Crace’
(3), one opts for a con between ‘readers’ and
‘book’ (4) and one annotator postulates a
negative effect on ‘Jim Crace’ (5). In the
harmonization process, the annotation 1 and 2 survived; the
annotator of 3 sticks with her opinion; it was
unproblematic to cancel 5; there were longer
discussions on relation 4 that was finally given up (by its
proponent).</p>
      <p>Clearly as a reader one immediately has a strong
opinion here, but there are so many aspects to
consider if you start a discussion on such cases – it is
really amazing.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Disagreement Analysis</title>
      <p>For 50 sentences, we performed a disagreement
analysis. This means, we harmonized our
annotations to get a gold standard, thereby checking
where and why our annotations diverged.</p>
      <p>In total, the four annotators produced 432
individual annotation decisions for the 50 sentences,
which amounts to 204 distinct annotations. After
harmonization, we ended up with 113 gold labels.
Sometimes all 4 annotators agreed upon a
particular decision, then 4 of the 432 individual
annotations yield a single gold standard annotation (one
of 113). Counting on individual annotations, there
was a 100% agreement (all four agreed) for 37
annotations (18% of all annotations). See table 3
for a detailed description. A comparison with the
whole sample (see table 1) reveals that our 50
sentences sample (table 3) well reflects the underlying
distribution of classes and thus might be
considered as representative.</p>
      <p>votes
annotations
%
The main question was: could we harmonize
the data without (much) dispute and discord? This
clearly would be the case if individual mistakes are
the reason for disagreement.</p>
      <p>It turned out that just 6% of the annotations
(27 out of 432) are based on plain mistakes (e.g.
wrong head of a phrase) and thus could be
resolved immediately. For another 34% of the
annotations, the annotators agreed during the
discussion that they missed out on an annotation (87
cases) or that their annotation was wrong (60 cases
out of 432). So 40% of the disagreement was away
and here there was hardly the need for strict
argumentation to convince each other to accept an
additional annotation or to drop one.</p>
      <p>But there are also cases where no agreement
was reached, i.e. there was at least one
annotation in a sentence which not all four annotators
could agree on. There are two cases: 26 out of
432 (6%) annotations on which not all four agreed
and 6 cases (1.4%), where one annotator did not
agree on an annotation which the others proposed.</p>
      <p>The rest of the cases (53.4%) are cases where
only after some discussion a harmonization step
was carried out. While 40% of the cases are valid
harmonizations, the nature of the harmonization of
the remaining 53.4% cases is unclear. In the next
two sections, we argue that strict harmonization is
harming (for our task).
7</p>
    </sec>
    <sec id="sec-7">
      <title>Majority Harmonization Means Harm</title>
      <p>One might argue that for the remaining 53.4%
cases a majority resolution was appropriate. Just
get rid of all singletons and adopt those
annotations where the majority of the voters agrees (3 or
4 voters). Only the 2 voter cases would have to be
dealt with on the basis of further discussions.</p>
      <p>This presupposes that the majority perspective
is the most valid one. We found out that,
statistically, this is not the case. We reached that
conclusion afterwards, i.e. after we have carried out
harmonization on the basis of discussion (which
was meant to clarify the reasons for disagreement
in the first place).</p>
      <p>Table 4 shows the frequency of adaptation
decisions. For instance, the first row shows that the
first annotator switched his opinion 22 times in the
case when the three other annotators voted in a
different way, i.e. voter 1 adopted his annotation
decision. There are two variants of this: voter 1
canceled his annotation since the others have not
approved it or voter 1 has not seen an annotation
step the others have and now he adopts it.
annotator id
1
2
3
4</p>
      <p>If majority vote proved to be superior over
singleton votes, then, in general, the inclination to
modify a decision should be dependent (increase)
on the number of voters that stand in opposition to
it. This means the more counter voters, the higher
the probability of a modification toward that
majority perspective. In order to test this, we
specified as a null hypothesis that the modification
decisions are independent (sic!) from the number
of voters that stand in opposition to it: three
voters change their mind quite as often as a singleton
voter (i.e. with the same probability). Then we
had independence.</p>
      <p>If we can reject the null hypothesis, if the
majority vote more often prevails, than the
harmonization strategy majority voting has proved valid
(and useful). If not, we have evidence that
singleton opinions are quite as valid as majority
decisions. We then have to discuss the status and
consequences of such a finding, namely whether
we should harmonize at all (see below).</p>
      <p>We applied Fisher’s exact test (in R) to the
table 4 and get as a p-value 0.35 which obviously is
not significant at any level (e.g. p &lt; 0.01). The
null hypothesis (independence, i.e. P(annotator’s
inclination of revisionjnumber of counter raters)
= P(annotator’s inclination of revision)) cannot be
rejected thus. This strengthens our claim that
harmonization should not just be realized as majority
voting.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Any Harmonization is Harm</title>
      <p>We found that a single opinion might turn out to be
as valid/strong as three opposing ones. But what
does this mean for the harmonization idea? The
majority harmonization strategy produces a gold
standard that only seemingly represents the best
choice - as we have seen. The discussion-based
harmonization strategy on the other hand produces
a gold standard where non-representative opinions
are as frequent as representative ones (those of the
majority of raters). Such a gold standard no longer
represents the prototypical reader - an entity we
would like to model.</p>
      <p>As a consequence: we neither should harmonize
by majority vote nor by discussion-based
agreements. We should not harmonize at all (beyond
the elimination of mistakes, of course).</p>
      <p>No harmonization means that our systems could
make use of all annotations in order to learn a
distribution of opinions. We then could interpret the
probabilities such a system would assign to a
particular decision as an indicator of its
prototypicality or prominence (visibility). It also would
produce singletons which might represent interesting
perspectives as well. This, of course, requires
further investigation and further proof.
9</p>
    </sec>
    <sec id="sec-9">
      <title>Related Work</title>
      <p>
        There exists a couple of papers dealing with
sentiment inference, see e.g.
        <xref ref-type="bibr" rid="ref2 ref3">(Deng and Wiebe,
2015a)</xref>
        , Rashkin et al. (2016), Klenner and Amsler
(2016),Klenner et al. (2017). There are also some
annotated resources, e.g. the MPQA corpus
        <xref ref-type="bibr" rid="ref2 ref3">(Deng
and Wiebe, 2015b)</xref>
        , but all approaches that we
know rely on a harmonization step. There are also
a number of papers dealing with (mostly
crowdsourcing related) annotation quality (
        <xref ref-type="bibr" rid="ref9">(Plank et al.,
2014)</xref>
        ,
        <xref ref-type="bibr" rid="ref5">(Hovy et al., 2013)</xref>
        ,
        <xref ref-type="bibr" rid="ref4">(Geva et al., 2019)</xref>
        ,
        <xref ref-type="bibr" rid="ref11">(Sheng et al., 2008)</xref>
        ). But none of these
approaches argues in favor of a gold standard in the
form of an decision distribution, as we do.
      </p>
      <p>As a notable exception, Kenyon-Dean et al.
(2018) present an interesting discussion on the
disagreement of annotators for the sentiment
labelling task. They demonstrate that bare
averaging or simple heuristics, such as the majority vote,
should be avoided. In contrast to our work, they
consider a task that includes only one label per
utterance, whereas we focus on effects and
relations, including multiple, possibly independent
instances per sentence. Additionally, in their work,
they investigate crowd-sourced annotations rather
than observing the outcomes of the harmonization
stage in the form of a discussion among the
annotators as we do. However, we see many
similarities in this paper to study sentiment annotation
disagreement, but viewed from a different
perspective. Also, we fully support the authors on their
postulate to release corpora with annotations from
all annotators, and withstanding from discarding
data samples for which the agreement is low.
10</p>
    </sec>
    <sec id="sec-10">
      <title>Conclusion</title>
      <p>In this paper, we argued in favor of an annotation
strategy that harmonizes as much as reasonable
(in order to get rid of errors and annotation
omissions), but otherwise leaves the distribution of
annotator decisions intact. We provided first
statistical evidence for such a strategy in the area of
sentiment inference, where annotation decisions are
far more subjective than, say, in PoS tagging. We
are interested in models of a prototypical reader,
we strive to model his/her understanding and we
believe that training a system on the basis of a
distribution of opinions better serves our purposes as
if an artificially created gold standard was used.
Acknowledgements We would like to thank
Noe¨mi Aepli, Sophia Conrad and Andreas
Sa¨uberli for their valuable contributions. This
work is supported by the Swiss National Science
Foundation under the project ID 105215 179302.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Lingjia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Yoonjung Choi, and
          <string-name>
            <given-names>Janyce</given-names>
            <surname>Wiebe</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Benefactive/malefactive event and writer attitude annotation</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>120</fpage>
          -
          <lpage>125</lpage>
          , Sofia, Bulgaria. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Lingjia</given-names>
            <surname>Deng</surname>
          </string-name>
          and
          <string-name>
            <given-names>Janyce</given-names>
            <surname>Wiebe</surname>
          </string-name>
          . 2015a.
          <article-title>Joint prediction for entity/event-level sentiment analysis using probabilistic soft logic models</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2015</year>
          , Lisbon, Portugal,
          <source>September 17-21</source>
          ,
          <year>2015</year>
          , pages
          <fpage>179</fpage>
          -
          <lpage>189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Lingjia</given-names>
            <surname>Deng</surname>
          </string-name>
          and
          <string-name>
            <given-names>Janyce</given-names>
            <surname>Wiebe</surname>
          </string-name>
          .
          <source>2015b. MPQA 3</source>
          .
          <article-title>0: An entity/event-level sentiment corpus</article-title>
          .
          <source>In Human Language Technologies</source>
          :
          <article-title>The 2015 Annual Conference of the North American Chapter of the ACL</article-title>
          , Denver, Colorado,
          <fpage>May31</fpage>
          -
          <lpage>June5</lpage>
          ,
          <year>2015</year>
          , pages
          <fpage>1323</fpage>
          -
          <lpage>1328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Mor</given-names>
            <surname>Geva</surname>
          </string-name>
          , Yoav Goldberg, and
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Berant</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Are we modeling the task or the annotator? An investigation of annotator bias in natural language understanding datasets</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>1161</fpage>
          -
          <lpage>1166</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Dirk</given-names>
            <surname>Hovy</surname>
          </string-name>
          , Taylor Berg-Kirkpatrick,
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Eduard</given-names>
            <surname>Hovy</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Learning whom to trust with MACE</article-title>
          .
          <source>In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>1120</fpage>
          -
          <lpage>1130</lpage>
          , Atlanta, Georgia. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Kian</given-names>
            <surname>Kenyon-Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Eisha</given-names>
            <surname>Ahmed</surname>
          </string-name>
          , Scott Fujimoto, Jeremy Georges-Filteau, Christopher Glasz, Barleen Kaur, Auguste Lalande, Shruti Bhanderi, Robert Belfer, Nirmal Kanagasabai, Roman Sarrazingendron, Rohit Verma, and
          <string-name>
            <given-names>Derek</given-names>
            <surname>Ruths</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sentiment analysis: It's complicated! In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , Volume
          <volume>1</volume>
          (
          <issue>Long Papers)</issue>
          , pages
          <fpage>1886</fpage>
          -
          <lpage>1895</lpage>
          , New Orleans, Louisiana. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Klenner</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Amsler</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Sentiframes: a resource for verb-centered German sentiment inference</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), pages
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          , Paris, France.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Klenner</surname>
          </string-name>
          , Don Tuggener, and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Clematide</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Stance detection in Facebook posts of a German right-wing party</article-title>
          .
          <source>In LSDSem</source>
          <year>2017</year>
          /
          <article-title>LSD-Sem Linking Models of Lexical, Sentential and Discourse-level Semantics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          , Dirk Hovy, and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning part-of-speech taggers with inter-annotator agreement loss</article-title>
          .
          <source>In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics</source>
          , pages
          <fpage>742</fpage>
          -
          <lpage>751</lpage>
          , Gothenburg, Sweden. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Hannah</given-names>
            <surname>Rashkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Connotation frames: A data-driven investigation</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL)</source>
          , Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Victor S. Sheng</surname>
            ,
            <given-names>Foster J.</given-names>
          </string-name>
          <string-name>
            <surname>Provost</surname>
          </string-name>
          , and
          <string-name>
            <surname>Panagiotis</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Ipeirotis</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Get another label? improving data quality and data mining using multiple, noisy labelers</article-title>
          .
          <source>In KDD</source>
          , pages
          <fpage>614</fpage>
          -
          <lpage>622</lpage>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>