<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multidimensional News Quality: A Comparison of Crowdsourcing and Nichesourcing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eddy Maddalena</string-name>
          <email>E.Maddalena@soton.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Ceolin</string-name>
          <email>davide.ceolin@cwi.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Mizzaro</string-name>
          <email>mizzaro@uniud.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centrum Wiskunde &amp; Informatica</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Southampton</institution>
          ,
          <addr-line>Southampton</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Udine</institution>
          ,
          <addr-line>Udine</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the age of fake news and of lter bubbles, assessing the quality of information is a compelling issue: it is important for users to understand the quality of the information they consume online. We report on our experiment aimed at understanding if workers from the crowd can be a suitable alternative to experts for information quality assessment. Results show that the data collected by crowdsourcing seem reliable. The agreement with the experts is not full, but in a task that is so complex and related to the assessor's background, this is expected and, to some extent, positive.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Online information is used by a variety of
stakeholders as a basis for decision making, knowledge
discovery, studies, and many more activities. However, as
a consequence of the democratic nature of the Web,
such information shows an extremely diverse level of
quality. Making explicit this level of quality for each
information item is crucial to allow the stakeholders an
overall adequate information perusal. Given their
pervasiveness and in uence on the public opinion, online
news are a kind of information whose quality
assessment becomes a particularly critical task to contrast
the spread of misinformation and disinformation.</p>
      <p>Assessing the quality of online news and
information in general is a challenging task, because of its
Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
intrinsic complexity. Information quality can be
assessed by considering diverse points of views; how they
can be assessed, and how the assessment results should
be combined, depends on the assessors and on their
requirements. This calls for a combined approach, where
automated computation is required to handle the huge
amount of information available on the Web, while
human computation is required to understand how the
quality dimensions are assessed and combined. An
important aspect of human computation in this context is
its regularity: when human assessments are consistent
enough, automated computation can leverage them to
scale the computation up.</p>
      <p>In a previous work by Ceolin, Noordegraaf, and
Aroyo [CNA16], two user studies are performed to
collect quality assessments regarding Web documents on
the vaccination debate. Assessments were collected
by means of a Web application, in a scenario
similar to crowdsourcing with the only di erence that the
assessments were expressed by a few experts (media
scholars and journalism students) rather than a large
crowd of anonymous workers. This approach has been
named nichesourcing [Boe+12]. Ceolin, Noordegraaf,
and Aroyo noted that, when the task at hand is
constrained, experts who show a similar background tend
to signi cantly agree with each other. However, they
also noted that the task of deeply assessing online
information is rather demanding, and expert availability
is limited. Crowdsourcing could be a solution to the
limited availability of human assessors.</p>
      <p>In this paper, we repeat that study [CNA16] though
crowdsourcing to analyse similarities and di erences
among the two ways of collecting human assessments.
Our ultimate goal is to determine if and how
crowdsourcing is a suitable alternative to nichesourcing for
information quality assessment. Section 2 brie y
surveys related work, Section 3 describes the experimental
setup we adopted, Section 4 presents the results, and
Section 5 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>In the age of fake news [Laz+18; VRA18] and of
the lter bubble [Par11], assessing the quality of
information is a compelling issue: it is important
for users to understand the quality of the
information they consume online. Two important
initiatives that are worth being mentioned in this eld are
the W3C Credible Web Community Group (https://
credweb.org/) and the Credibility Coalition (http:
//credibilitycoalition.org). While the rst is
meant to establish standards to model and share data
about the credibility of information online, the second
aims at identifying markers and strategies for
establishing the credibility of the same information. To this
extent, the work we present in this paper is
complementary to these initiatives, as it aims at providing
gold standards to reason on the credibility (and, more
broadly, quality) of online information.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setup</title>
      <sec id="sec-3-1">
        <title>Dataset Description</title>
        <p>We ran our experiment on a sample from the
vaccination debate dataset provided by the QuPiD project
(http://qupid-project.net) and used by Ceolin,
Noordegraaf, and Aroyo [CNA16]. In 2015, a measles
outbreak took place at Disneyland, California. Such
outbreak triggered a erce debate that eshed out the
already hot discussions regarding vaccinations, where
pro and anti vaccination individuals blamed each other
for the responsibility of the event. The vaccination
debate dataset collects a number of documents
regarding that speci c debate. While the dataset is limited
in size (about 50 documents), it is rather diverse in
terms of types of documents represented (newspaper
articles, activist blog posts, etc.) and stances (pro,
anti, neutral).
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>The Crowdsourcing Task</title>
        <p>The crowdsourcing task we ran aimed at collecting
laymen judgments concerning the quality of a subset
of 20 articles assessed by the experts (media scholars
and journalism students). We asked each worker to
assess one document along eight di erent quality
dimensions derived from Ceolin, Noordegraaf, and Aroyo
[CNA16] (we slightly reformulated some of them to
have a shorter description, more adequate for crowd
workers):
1. Accuracy - How accurate is the information in this
article?
2. Neutrality - Is the document neutral with respect
to the topic addressed, or does it clear stance (e.g.,
pro, against)?
3. Readability - Does the document read well?
4. Precision - How precise is the information in this
document (as opposed to vague)?
5. Completeness - How complete is the information in
this document?
6. Trustworthiness - How trustworthy is the source? Is
the source trustworthy or does it exhibit malicious
intentions?
7. Relevance - How relevant is the article to the task?
8. Overall quality - Which is your general opinion
about the quality of the article?
We also asked two further questions requiring workers
personal opinion, to understand how personal belief
a ects quality judgment:
9. Your personal opinion - Do you agree with the
document content?
10. Your con dence - How knowledgeable/expert are
you about the topic?
All the 10 assessments were collected on a 5-stars
Likert scale, as in the original experiment [CNA16]. For
each quality dimension, we also asked the users to
motivate their judgment by some free text.</p>
        <p>The task ran on the Figure Eight (https://www.
figure-eight.com/) crowdsourcing platform by
selecting level-three workers who are highest accuracy
contributors. Each worker was paid 0.2 USD and could
not judge more than three articles. Besides
redundancy (each article was judged by 10 workers), we also
adopted some standard quality checks: each worker
was shown a pair of articles of clearly low and high
quality, and the work was rejected if the collected
values were ranked in the wrong way; there was also a
time threshold (the worker needed to spend at least
120 seconds on the task), and some syntactic checks
on the free text motivations.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Research Questions</title>
        <p>This experiment allows us to address three research
questions:
Q1. Relationships between quality dimensions: what
are the correlations between the quality
dimensions? Do some of the quality dimensions
correlate in a way that makes one derivable from
another? What is the di erence between experts
and workers?
Q2. Internal agreement (between individual workers):
can di erent workers agree to a reasonable extent
when assessing quality dimensions? Are there
differences among the dimensions?
Q3. External agreement (between individual workers
and experts): what is the individual external
agreement, i.e., the agreement between the
individual workers and the experts, on all
dimensions? What is the aggregate external agreement,
cccyuaA152345
r
tryaue43
lililitttyaaNRQliliitttrrrssscssvscoepneenuhoeeenaenoeTPCwRm21353442154512312512321543543
y
li4
b
d
a
e
u
rvae2
O1
=.57, p=4.8e-04 =.73, p=9.4e-07 =.83, p=1.4e-09 =.76, p=1.5e-07 =.88, p=7.0e-12 =.64, p=5.1e-05 =.9, p=2.2e-13
*** *** *** *** *** *** ***
r=.54, p=9.8e-04 r=.71, p=2.0e-06 r=.82, p=2.9e-09 r=.76, p=2.4e-07 r=.89, p=1.0e-12 r=.58, p=3.0e-04 r=.93, p=1.1e-15
=.49, p=8.0*e*-0*4 =.6, p=1.8e-05 ***
*** =.74, p=1.1*e*-0*7 =.67, p=2.5*e*-0*6 =.81, p=7.4*e*-0*9 =.51, p=3.2*e*-0*4 =.88, p=3.5e-10
*** =.42, p=1.4*e*-0*2 =.5, p=2.5e*-0*3* =.7, p=3.7e*-0*6* =.62, p=1.0*e*-0*4 =.33, p=5.8*e*-0*2 =.45, p=8.2*e*-0*3
r=.41, p=1.7e-0*2r=.48, p=3.9e-*0*3r=.7, p=4.1e*-0*6*r=.61, p=1.4*e-*0*4r=.25, p=1.5e-01 r=.46, p=6.4e-*0*3
=.33, p=2.1e-0*2 =.44, p=1.9e*-0*3 =.62, p=2.8*e*-0*5 =.57, p=8.0*e*-0*5 =.23, p=1.1e-01 =.42, p=3.6e-03
**
* =.67, p=1.3e*-0*5 =.62, p=1.1*e*-0*4 =.76, p=2.2*e*-0*7 =.65, p=3.4e-05 =.69, p=5.7e*-0*6</p>
        <p>*** *** *** *** ***
r=.64, p=4.0e-05 r=.61, p=1.2e-04 r=.75, p=3.3e-07 r=.59, p=2.5e-04 r=.68, p=1.0e-05
=.53, p=1.3*e*-0*4 =.55, p=1.1*e*-0*4 =.64, p=4.2*e*-0*6 =.52, p=2.0*e*-0*4 =.57, p=3.2e-05
***
*** =.78, p=4.4*e*-0*8 =.8, p=1.3e*-0*8* =.73, p=1.2*e*-0*6 =.79, p=2.1*e*-0*8</p>
        <p>*** *** *** ***
r=.77, p=1.2e-07 r=.81, p=5.2e-09 r=.65, p=2.7e-05 r=.79, p=2.7e-08
=.66, p=2.5*e*-0*6 =.74, p=1.1*e*-0*7 =.55, p=7.8*e*-0*5 =.69, p=5.1e-07</p>
        <p>***
*** =.72, p=1.5*e*-0*6 =.67, p=1.2*e*-0*5 =.7, p=3.5e*-0*6*</p>
        <p>*** *** ***
r=.73, p=1.0e-06 r=.61, p=1.2e-04 r=.72, p=2.0e-06
=.64, p=6.6*e*-0*6 =.52, p=2.6*e*-0*4 =.63, p=7.9e-06</p>
        <p>***
*** =.6, p=1.8e*-0*4* =.82, p=3.7*e*-0*9</p>
        <p>*** ***
r=.55, p=7.5e-04 r=.83, p=1.8e-09
=.47, p=8.3*e*-0*4 =.76, p=5.4e-08</p>
        <p>***
*** =.67, p=1.4*e*-0*5</p>
        <p>***
r=.64, p=4.7e-05</p>
        <p>***
=.54, p=1.0e-04
***
1 A2ccu3rac4y 5 1 N2eut3rali4ty 5 1Re2ad3abi4lity5 1 P2rec3isio4n 5 C1om2ple3ten4es5s T1rus2two3rthi4nes5s 1 Re2lev3an4ce 5 O1ve2ral3Qu4alit5y
1 A2ccu3rac4y 5 1 N2eut3rali4ty 5 1Re2ad3abi4lity5 1 P2rec3isio4n 5 C1om2ple3ten4es5s T1rus2two3rthi4nes5s 1 Re2lev3an4ce 5 O1ve2ral3Qu4alit5y
i.e., the agreement between the aggregated
assessments by the workers and the experts, on all
dimensions?
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The main results are grouped on the basis of the
research questions.
4.1</p>
      <sec id="sec-4-1">
        <title>Q1: Quality Dimensions Relationships</title>
        <p>A rst result is presented in Figure 1, that shows a
scatterplot matrix. For each pair of dimensions
(indicated on the diagonal), a scatterplot is shown (in
the bottom triangular matrix, with some random
jitter to avoid some overlap). Each dot in a scatterplot
represents one individual worker/article pair, and its
coordinates are the values expressed by the worker on
the corresponding two dimensions. In the upper
triangular part, the correlation values are shown with their
p-values to measure statistical signi cance.</p>
        <p>Figure 2 allows to compare the data to experts.
Comparing correlation values, it is clear that experts
are more consistent across dimensions; p-values are
roughly similar in the two cases.</p>
        <p>As it is common practice in crowdsourcing, in place
of using raw values by individual workers, we
compute aggregated values. We select a simple (if not the
simplest) aggregation function: the arithmetic mean.
Figure 3 shows the correlations obtained when
aggregating with the mean the 10 values expressed by 10
workers on the same article. When comparing to
Figure 1, one can see that correlations increase, although
they are less statistically signi cant. When comparing
to Figure 2 one can see that usually the correlation
between dimensions are higher for the experts than for
the aggregate workers, but values are de nitely more
comparable than the individual raw values, and indeed
the aggregate workers have higher correlations than
the experts in three cases (the correlations between
Accuracy and Relevance those between Overall
Quality and both Neutrality and Precision). We also tried
aggregating with the median, obtaining worse results.</p>
        <p>Another remark that can be made by observing the
histograms on the diagonals of Figures 1 and 2 is that
the values provided by the experts tend to follow a
more Bimodal distributions (they use more the
extremes of the scale) than the workers. This is even
clearer when looking at the aggregated values since the
mean of the values will pull them even more towards
the middle of the scale, as it can be seen in Figure 3.
The distributions also show that the workers tend to
express higher values than the experts.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Q2: Internal Agreement among Workers</title>
        <p>Table 1 shows the agreement among the workers,
overall and on each quality dimension, measured by both
Krippendor 's [Kri07] and [Che+17]. Both
measures assume values in [ 1; +1] (with 1
corresponding to complete disagreement, 0 to random agreement,
and +1 to complete agreement). For the table also
shows, besides the most likely value, the Highest
Posterior Density (HPD) interval, i.e., the interval that
contains the actual value with a 95% probability:
these are quite small intervals, so we can be con
cccyuaA215345
r
tryaue34
illilitttyaaRNQliliitttrrrscvsssscsouneaeehenoepneeeonTPwRCm23154312451234152343215512543
y
li4
b
d
a
e
u
rvae2
O1
=.48, p=4.3e-02 =.37, p=1.3e-01 =.75, p=3.4e-04 =.57, p=1.3e-02 =.79, p=9.8e-05 =.77, p=1.8e-04 =.79, p=1.0e-04
r=.49, p=3.8e-0*2r=.39, p=1.1e-01 r=.7, p=1.3e*-0*3*r=.6, p=8.3e-03*r=.77, p=1.8*e-*0*4r=.76, p=2.6*e-*0*4r=.73, p=5.8*e-*0*4
=.33, p=6.5e-0*2 =.31, p=8.8e-02 =.57, p=1.6e*-0*3 =.47, p=9.8e*-0*3 =.66, p=2.3*e*-0*4 =.65, p=4.0*e*-0*4 =.61, p=7.2e-04
***
** ** *** *** ***
=.39, p=1.1e-01 =.33, p=1.8e-01 =.48, p=4.6e-02 =.54, p=2.2e-02 =.23, p=3.6e-01 =.57, p=1.4e-02
r=.31, p=2.1e-01 r=.28, p=2.6e-01 r=.43, p=7.7e-0*2r=.58, p=1.2e-0*2r=.055, p=8.3e-01r=.57, p=1.4e-0*2
=.25, p=1.6e-01 =.2, p=2.7e-01 =.34, p=6.0e-02 =.4, p=2.4e-02* =.035, p=8.5e-01 =.42, p=1.9e-02
*
=.42, p=8.6e-02 =.5, p=3.5e-02 =.39, p=1.1e-0*1 =.13, p=6.0e-01 =.42, p=8.1e-0*2
r=.24, p=3.5e-01 r=.36, p=1.5e-0*1r=.25, p=3.1e-01 r=.18, p=4.8e-01 r=.23, p=3.6e-01
=.19, p=3.1e-01 =.27, p=1.4e-01 =.19, p=3.0e-01 =.14, p=4.6e-01 =.17, p=3.5e-01
=.74, p=4.0e-04 =.71, p=1.1e-03 =.67, p=2.5e-03 =.82, p=3.0e-05
r=.75, p=3.4*e-*0*4r=.6, p=8.6e-0*3*r=.58, p=1.1e-*0*2r=.86, p=4.6*e-*0*6
=.6, p=9.1e-04 ***
*** =.47, p=9.1e*-0*3 =.48, p=9.1e-0*3 =.71, p=8.8e-05
*** =.63, p=5.0e*-0*3 =.4, p=1.0e-0*1* =.82, p=3.2*e*-0*5</p>
        <p>** ***
r=.62, p=6.1e-03 r=.43, p=7.8e-02 r=.86, p=5.0e-06
=.47, p=9.0e*-0*3 =.35, p=6.1e-02 =.72, p=7.4e-05</p>
        <p>***
** =.57, p=1.4e-02 =.75, p=3.5*e*-0*4
r=.57, p=1.4e-0*2r=.65, p=3.5*e-*0*3
=.43, p=2.0e-0*2 =.54, p=2.5e-03</p>
        <p>**
* =.49, p=4.0e*-0*2
r=.43, p=7.8e-0*2
=.36, p=5.2e-02
1 A2ccu3rac4y 5 1 N2eut3rali4ty 5 1Re2ad3abi4lity5 1 P2rec3isio4n 5 C1om2ple3ten4es5s T1rus2two3rthi4nes5s 1 Re2lev3an4ce 5 O1ve2ral3Qu4alit5y
dent that the most likely value is correct.
values are quite low, but ones are much higher. Most
likely, as we have discussed above, assessment values
have a quite low variability. In such a case, exhibits
a pathological behavior, which is of the issues with
that is solved by as discussed by Checco et al.
[Che+17]. The much higher values, together with
the narrow HPD intervals, show that the agreement
among the workers is consistent even if not complete.</p>
        <p>The results presented so far hint that the data
collected by our crowdsourcing experiment are reliable. It
is also important to remark that although the workers
in some cases fail to exactly replicate the assessments
by the experts (as we discuss shortly), the task is quite
complex and assessor background might have a critical
role. In this respect, a full agreement might even be
a problem rather than a feature. If this is the case,
it might be necessary to treat in a di erent way
different worker groups, and/or decrease the granularity
and ask to evaluate passages of an article instead of a
full article. In this light, we observe a low correlation
(between 0 and 0.20) between the workers con dence,
i.e., question number 10, and all the quality
dimensions and a moderate correlation (about 0.6) between
the workers agreement, i.e., question 9, with the article
assessed and Precision, Accuracy, and Overall Quality
scores. While this correlation is not complete, it still
hints at the possibility that a subgroup of the workers
shows a con rmation bias, meaning that these tend to
judge positively the articles they agree with, and
viceversa. In this short paper we do not have the space
to discuss these issues in full, and we leave them for
future work.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Q3: External Agreement with the Experts</title>
        <p>Turning to the agreement between workers and
experts, the scatterplots and correlations values in
Figure 4 (top row) show that the agreement of the
individual workers with the experts is rather low, as
correlation values are positive but quite small, and
often not signi cant. Figure 4 (center row) shows the
agreement with the experts that is obtained when
aggregating the worker values with the mean.
Correlation values are systematically higher than individual
workers, although almost never greater than 0:5 and
often not statistically signi cant. As previously
observed, the aggregation reduces the range of the values:
whereas the experts usually use the full spectrum, the
aggregated workers score is more limited. In all these
plots, the eight dimensions show quite similar
correlation values with the exception of Neutrality: workers
particularly disagree with the experts about it.</p>
        <p>Figure 4 (bottom row) demonstrates the previous
claim that in general the median is a worse
aggregation function: lower correlation values are obtained for
Completeness, Trustworthiness, Relevance, and,
especially, Overall Quality (which has not correlation with
the experts when using the median). However,
Readability and Precision are similar, and Neutrality and,
especially, Accuracy are higher. This suggests that
di erent and more sophisticate aggregation functions
might lead to a higher agreement with the experts, an
issue that for space limits we leave for future work.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>In this paper we present an experiment that aims at
comparing crowd and nichesourcing as methods for
assessing the quality of online information from a
multidimensional standpoint. We collect 10 assessments
about 20 articles from a dataset on the vaccination
debate, and we analyze them internally and in
comparison to previously published expert assessments. We</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Checco</surname>
          </string-name>
          , Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, and Gianluca Demartini. \
          <article-title>Let's Agree to Disagree: Fixing Agreement Measures for Crowdsourcing"</article-title>
          .
          <source>In: The 5th AAAI Conference on Human Computation and Crowdsourcing (HCOMP</source>
          <year>2017</year>
          ).
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Ceolin</surname>
          </string-name>
          , Julia Noordegraaf, and Lora Aroyo.
          <article-title>\Capturing the Ine able: Collecting, Analysing, and Automating Web Document Quality Assessments"</article-title>
          .
          <source>In: Knowledge Engineering and Knowledge Management</source>
          . Springer International Publishing,
          <year>2016</year>
          , pp.
          <volume>83</volume>
          {
          <fpage>97</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Krippendor</surname>
          </string-name>
          . \
          <article-title>Computing Krippendor 's alpha reliability"</article-title>
          .
          <source>In: Departmental papers (ASC)</source>
          (
          <year>2007</year>
          ), p.
          <fpage>43</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Metzger</surname>
            , Brendan Nyhan, Gordon Pennycook, David Rothschild,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Schudson</surname>
          </string-name>
          , Steven A.
          <string-name>
            <surname>Sloman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Cass R. Sunstein</surname>
          </string-name>
          , Emily A.
          <string-name>
            <surname>Thorson</surname>
            ,
            <given-names>Duncan J.</given-names>
          </string-name>
          <string-name>
            <surname>Watts</surname>
          </string-name>
          , and Jonathan L. Zittrain. \
          <article-title>The science of fake news"</article-title>
          .
          <source>In: Science 359.6380</source>
          (
          <year>2018</year>
          ), pp.
          <volume>1094</volume>
          {
          <fpage>1096</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Pariser</surname>
          </string-name>
          .
          <article-title>The Filter Bubble: What the Internet Is Hiding from You</article-title>
          . The Penguin Group,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Soroush</given-names>
            <surname>Vosoughi</surname>
          </string-name>
          , Deb Roy, and Sinan Aral. \
          <article-title>The spread of true and false news online"</article-title>
          .
          <source>In: Science 359.6380</source>
          (
          <year>2018</year>
          ), pp.
          <volume>1146</volume>
          {
          <fpage>1151</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Victor de Boer</surname>
          </string-name>
          , Michiel Hildebrand, Lora Aroyo, Pieter De Leenheer, Chris Dijkshoorn, Binyam Tesfa, and Guus Schreiber. \
          <article-title>Nichesourcing: Harnessing the Power of Crowds of Experts"</article-title>
          .
          <source>In: Knowledge Engineering and Knowledge Management</source>
          . Springer Berlin Heidelberg,
          <year>2012</year>
          , pp.
          <volume>16</volume>
          {
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>[VRA18] [Par11]</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>