<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>user study on people's perception to the credibility of online health information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcos Fernández-Pichel</string-name>
          <email>marcosfernandez.pichel@usc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markus Bink</string-name>
          <email>markus.bink@ur.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David E. Losada</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Elsweiler</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Compostela</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spain</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela</institution>
          ,
          <addr-line>Santiago de</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Chair of Information Science, Universität Regensburg</institution>
          ,
          <addr-line>Regensburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Health-related content</institution>
          ,
          <addr-line>Credibility, User study</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Judging the credibility of information is a subjective process and prone to biases. This issue can be especially concerning in health information seeking. Some eforts have been made to define robust credibility assessment guidelines that support the development of reliable test collections. This is of the utmost importance since the applicability of retrieval algorithms to real use case scenarios relies on the quality of the labelled data. Yet, the question persists as to whether the labels created by these guidelines can efectively serve as a surrogate for the genuine judgements of credibility as perceived by end-users. Motivated by this, we conducted a user study with 1,000 participants. We demonstrate that there is a correlation between participants' judgements and the reference values produced following existing guidelines. Further analyses of the data reveal worrying insights into people's ability to judge the credibility of online medical content, leading to potential personal harm.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        The Internet has become the dominant platform for accessing health information, ofering
convenient access to a wealth of medical knowledge [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. Nonetheless, the abundance
of information poses a challenge for users in discerning trustworthy sources from unreliable
ones, potentially resulting in ill-informed choices regarding their health [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ]. In extreme
cases, this situation can have severe consequences and even poses a risk to personal
wellbeing [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Credibility has been defined as the extent to which information from a webpage or
other online source can be believed [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. It is a highly subjective concept that is susceptible to
individual diferences, such as user’s reading skills [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The subjective nature of credibility
represents a barrier in creating reliable and robust test collections. In the context of shared-task
evaluation campaigns, some researchers have critically analysed the quality of the credibility
assessments and proposed a set of robust and traceable guidelines to improve the robustness of
nEvelop-O
annotations [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This is important since the applicability of retrieval and machine learning
algorithm relies on the quality of the annotation process.
      </p>
      <p>
        Nevertheless, there is still a need for a rigorous examination of the relationship between this
type of guidelines and credibility as real end-users perceive it. While we know that judgements
vary across users, annotations need to be both consistent and reflective of average users’
perceptions. Previous research has studied the main elements influencing individual credibility
perceptions [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. However, no user-oriented study has attempted to understand annotation
practices for shared-tasks. Such a study could also provide valuable cues on how people evaluate
the credibility of websites posting medical information.
      </p>
      <p>In this work, we perform a study to understand how end-users perceive the credibility of
online health information1. The ultimate goal being to determine whether the labels created
from guidelines can serve as surrogate of credibility perceived by end-users. We attempt to
answer the following research questions:
• RQ1. Can current credibility annotation guidelines act as a proxy of the real perception
of credibility of end-users?
• RQ2. To what extent users are able to recreate the judgements of experts?
• RQ3. How do user variables, such as familiarity with the search topic, educational
background, and other human factors, afect the user’s perception of credibility?</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related work</title>
      <p>
        The credibility of online information and the spread of misinformation have been extensively
studied [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16, 17, 18</xref>
        ]. Viviani and Pasi reviewed the main automatic methods to estimate
credibility in social media, focusing mainly on health content [19]. A further body of work has
sought to understand how end-users assess the credibility of online content and why people
make certain assessments. For instance, Fogg defined the prominence-interpretation theory,
which helps to determine which website elements influence end-users’ credibility [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This
theory was later tested through a user study involving 2,500 participants, where authors found
that 46% of the users mentioned design as a critical aspect influencing credibility [ 20].
      </p>
      <p>
        Easting et al. [21] demonstrated that both the source and the prior knowledge about the
content have influence on users’ perception of online health information. Other studies have
also demonstrated that, apart from the characteristics of the web elements, the receiver’s
characteristics also influence the perception of the information [ 22]. Other researchers analysed
in-depth the factors that influence end-users’ perceived credibility [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Previous studies have
also evaluated the correlation between diferent users’ judgements to test their feasibility as
ground truth values [23, 24].
      </p>
      <p>In this paper, we present a systematic user study that shows how well expert annotations
reflect the subjective judgements of a broad population of users. We also evaluate a number of
personal factors that may influence the credibility estimations.
1https://github.com/MarcosFP97/perceived-credibility-study</p>
      <sec id="sec-3-1">
        <title>Topic id T1 T5 T8</title>
        <p>T10</p>
      </sec>
      <sec id="sec-3-2">
        <title>Question</title>
        <p>Do antioxidants help female subfertility?
Do sealants prevent dental decay in permanent teeth?
Does melatonin help treat and prevent jetlag?
Does traction help low back pain?</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Experimental setup</title>
      <p>We hypothesise that reference values (created by expert annotators using formal guidelines)
will correlate with users’ judgements. To test this, we conducted a crowd-sourced user study
whereby participants provided credibility judgements for webpages.</p>
      <sec id="sec-4-1">
        <title>3.1. Dataset</title>
        <p>
          We utilised a pre-existing dataset from the medical domain that had been originally compiled
by Pogacar et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and later extended by Zimmerman et al. [25]. We extracted 162 screenshots
of webpages from it, such that the selected webpages provide answers spanning four distinct
health-related topics, as detailed in Table 12. Each participant was presented with a full-scale
screenshot of a randomly selected webpage and was asked to assess the webpage’s credibility
on a 7-point Likert-scale. Each webpage was evaluated at least 5 times by diferent participants
to minimise personal bias [
          <xref ref-type="bibr" rid="ref10">18, 10</xref>
          ].
        </p>
        <p>
          For these webpages, we also produced annotations generated by human assessors according
to the guidelines from [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], detailed in Table 2. We recruited four diferent assessors 3. For a
given topic, the webpages were annotated by the same pair of assessors. Next, the two assessors
responsible for each topic convened a meeting to discuss and consolidate their annotations and
generate a final set of labels. These final annotations were used as reference values in our user
study.
3.2. Variables
        </p>
        <sec id="sec-4-1-1">
          <title>3.2.1. Independent variables</title>
          <p>• Reference values: variable indicating the credibility-level perceived by the human
annotators according to the guidelines. There are three possible levels: 0 (non-credible),
1 (credible), and 2 (highly credible).</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>3.2.2. Dependent variables</title>
          <p>• User credibility score: the credibility score assigned by the crowdsourcers in a 7-point</p>
          <p>Likert scale.
2We did not use the raw HTML, since we consider visual elements as key to the perception of credibility.
3Three PhD students with background in British Studies, Computer Science and Information Science, respectively,
and a Master’s Degree student in Information Science</p>
          <p>G1
G2
G3
G4
G5
G6</p>
          <p>Label
2
1
1
0
0
0</p>
          <p>Guideline
Source is a scientific paper, or a Medical publisher or hospital/clinic or government
website or university.</p>
          <p>Document is citing the information they provide in their articles. They provide links
or specific references to their sources. They cite sources with credibility 2 (i.e. medical
publications and/or lab studies).</p>
          <p>Document is written by an expert in the field/someone qualified to write this
document (irrespective of publishing venue).</p>
          <p>The document is actually for advertising or marketing purposes. If so, the website
might be biased or a scam designed to trick people into fake treatments or into
buying medical products that do not live up to their claim.</p>
          <p>The information posted by a non-expert person providing a medical product review
or providing medical advice without proper citations (links/list of references).</p>
          <p>The website provides or states claims that go against well-known medical consensus
(e.g. smoking cigarettes does not cause cancer).</p>
          <p>NOTE: It is generally allowed to look up authors to check whether they have the required knowledge to
be regarded as an expert and look up websites to find out if they are legitimate.
• Time of completion (in minutes): variable representing the total amount of time it
took for a crowdsourcer to complete the assessment (measured from the moment the
screenshot of the website was shown until successful completion).</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>3.2.3. Descriptive and exploratory variables</title>
          <p>We also studied some variables that can influence or have some connection with the credibility
scores gathered in the study (this relation is further explored in Section 4):
• Topic familiarity: in the pre-task questionnaire, see Figure 1 (left side), participants
were asked about their prior knowledge on the topic.</p>
          <p>• Personal data: in a post-study questionnaire, see Figure 1 (right side), we gathered
additional information about the participants in the study (educational background,
gender, and age).
• Justifications : we also provided the participants with the possibility of justifying their
rating in their own words (free text field).</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.3. Procedure</title>
        <p>Once participants had been presented with the goals of the study, its methodology, and the
implications of their involvement, they provided their permission by signing a consent form.
Next, they proceeded with 3 steps to satisfactorily complete the study:
1. Each participant was randomly assigned one webpage from the collection. Before seeing
the webpage’s screenshot, they needed to fulfil a pre-task questionnaire about their
expertise on the topic of the website, see Figure 1 (left side).
2. Subjects were shown a screenshot of the entire webpage and they needed to assess its
credibility in a 7-point Likert scale (from not credible at all to very credible), see Figure
2. They had no time limit to provide this estimate, and they could scroll through the
entire screenshot and provide a free-text justification about their judgement (this step
was not mandatory, however a high number of participants provided this feedback, see
Section 4.6).
3. Before ending the study, participants were shown a post-study questionnaire, see Figure 1
(right side). Our main goal was to gather additional data, such as educational background,
age, and gender.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.4. Participants</title>
        <p>We recruited a total of 1,000 users to guarantee at least 5 judgements per webpage. We used
the Prolific 4 platform and each participant received £0.32 (equivalent to £9.60 per hour). Our
participants were fluent English speakers, resident in the United States or the United Kingdom.
The annotators belonged to an age range between 18 and 85 years old. 53% identified as female,
45% as male, and the remaining participants identified either as diverse or other. In terms of
educational background, 40% had a bachelor’s degree, 37% completed secondary education, and
only 2.6% of the participants reported a level of education below high school.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <sec id="sec-5-1">
        <title>4.1. Reference credibility values vs users’ judgements</title>
        <p>
          RQ1 seeks to determine whether the current annotation guidelines, whose goal is to produce
robust assessments of credibility for medical websites, serve as a proxy for the human’s perception
of credibility [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. We analysed the distribution of the crowdsourced judgments according to the
three levels of reference values. As can be seen in Figure 3, it seems that there is a relationship
between both variables. The participants’ judgements tend to be higher when presented with
webpages of increasing credibility (according to experts). This is confirmed by a Spearman’s
rank correlation ( = 0.26 and a  −   &lt; 0.01 ), indicating weak agreement according to [26].
        </p>
        <p>Despite this correlation, there are some signs of concern regarding RQ2. Webpages annotated
as non-credible according to the guidelines were often perceived as reliable by the participants.
This can be observed by the fact that non-credible documents have very high perception
scores and their median score is 5. This confirms previous research findings that people
tend to overestimate credibility and have problems identifying low-quality sites [27, 28]. In
general, webpages labelled as credible or highly credible by the reference judgements were
also considered as of high quality by the crowdworkers. However, people struggled to detect
contents that are regarded as low quality by the reference annotations.</p>
        <p>Summing up, regarding RQ1, we can conclude that the participants’ judgements and the
reference values derived from the guidelines are correlated. As for RQ2, we found out that
crowdworkers are less prone to errors when evaluating high quality pages, but they struggled
for pages with lower levels of credibility.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Topic familiarity</title>
        <p>Prior to completing the study and to partially answer RQ3, we asked participants about their
level of knowledge or familiarity with their assigned topic in a 5-point Likert scale (ranging
from Not at all familiar to Extremely familiar, see Figure 1 (left side)).</p>
        <p>Figure 4 (left side) shows the relation between levels of familiarity and the participants’
judgements. It seems that the higher the familiarity, the higher the credibility judgements
provided by participants. Spearman’s correlation yielded a  = 0.10 and  −   &lt; 0.01 ,
demonstrating that there is a very weak correlation between the two variables [26].</p>
        <p>To further explore the user study data, we also computed the deviation per level of familiarity
between the reference values and the user study’s judgements. First, we applied a Min-Max
normalisation to both sets of scores. Then, the diference between the crowdsourcing judgements
and the reference values was computed –0 represents a perfect match, while a positive (negative)
value means that people overestimated (underestimated) credibility– see Figure 4 (right side).
From the figure, we might conclude that there is not a strong relation between familiarity and
how efective are users at rating webpages. However, Spearman’s correlation yielded a  = 0.13
and  −   &lt; 0.01 .</p>
        <p>As a complementary analysis, we also computed the mean familiarity per topic (and its
standard deviation): T1 has a mean familiarity of 1.55 (0.89), T5 of 1.79 (1.07), T8 of 2.44 (1.25),
and T10 of 2.14 (1.15).</p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Topic analysis</title>
        <p>Figure 5 (left side) reports a topic-level analysis. As can be expected, there are individual
diferences among the topics. Spearman’s test also revealed statistically significant correlations
between reference credibility scores and crowdsourced credibility scores for all topics. However,
the correlation for T1 was lower. These results fit with the familiarity scores described above,
where T1 was shown to be the topic that users had less knowledge about.</p>
        <p>Again, we also computed the deviation per topic between the reference values and the user
study’s judgements, see Figure 5 (right side). Spearman’s test revealed an statistically significant
correlation between both variables ( = 0.20 and  −   &lt; 0.01 ). An interesting finding is
that for the topics users are more familiar with, T8 and T10, they tend to overestimate their
perceived credibility.</p>
      </sec>
      <sec id="sec-5-4">
        <title>4.4. Other user variables</title>
        <p>To fully answer RQ3, the relation between additional user variables (gender, age, and educational
background) and their perception of credibility was explored. For the first two variables, no
significant correlations or revealing trends were found. However, for the educational background
(Figure 6), we found an interesting conclusion: all groups were equally good at estimating
credibility, except for the less educated (less than high school group) and the most educated
(doctorate group). The Spearman’s test reported a  = 0.05 and  −   = 0.12 rejecting
the hypothesis that there is a correlation between the educational level and the quality of the
assessments, measured by the deviation between the crowdsourced and reference values.</p>
      </sec>
      <sec id="sec-5-5">
        <title>4.5. Time of completion</title>
        <p>We also analysed the time (in minutes) users needed to complete the assessment. Figure 7 shows
that users who spent less time analysing the web (between 0-6 minutes) tended to deviate less
from the reference values (deviation close to 0). We speculate that “overthinking” might be
counterproductive for this task. Alternatively, the lower quality of the estimates at the right
end of the graph could be due to other factors such as distractions. Related to this, previous
studies showed that people who take more time on this type of tasks tend to be more influenced
by visual elements and their prior knowledge [29]. In any case, the Spearman’s correlation test
revealed no statistical significance (with a  = 0.009 and a  −   = 0.78 ) between completion
time and deviation between crowdsourced and reference judgements.</p>
      </sec>
      <sec id="sec-5-6">
        <title>4.6. Analysis of justifications</title>
        <p>We ofered users the possibility of justifying their judgements. This was actively used by
participants, with 93% providing a textual explanation. This gave us valuable evidence to analyse
the reasons behind credibility judgements. Yet, manual inspection was infeasible because we
had thousands of datapoints. We therefore opted for exploiting the summarisation capabilities
of current Large Language Models (LLMs). To that end, the justifications were grouped by the
diferent levels of perceived credibility and GPT-4 was provided with these textual extracts
and asked to generate a summary for each level. The template used for prompting the LLM
was as follows: “We are a group of scientists that have conducted an online survey on webpage
credibility. For each webpage, we asked a human assessor to provide a score credibility from 1
to 7 (very low credible to very credible). Assessors could also provide a justification on why they
assigned a given credibility score. Given a series of justifications between &lt;&gt;, I want you to generate
an understandable summary. &lt;justifications&gt;. The summary is: ”. The resulting summaries are
presented in Table 3.</p>
        <p>
          Some interesting patterns were observed in the provided explanations. Low credibility was
usually associated with a poor visual appearance of the webpage. Moreover, several users
mentioned the lack of reputable references as a critical aspect to mistrust information. This
suggests the existence of a direct relationship between the real perception of end-users and
the existing guidelines [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], which incorporate criteria such as the presence of authoritative
references. Regarding cases of high perception of credibility, users tend to trust professional
appearance and well-structured content. They also mentioned the importance of the
qualifications of the authors or the inclusion of credible sources. These two criteria were also considered
in the original assessment guidelines (citing reputable sources or the expertise of the author).
Crowdsourcers mentioned governmental sites, e.g. NHS or CDC, as the highest trustful sources.
This also matches with the criteria stated in the guidelines. This qualitative analysis provides
tangible evidence to support current guidelines as proxies of users’ perceptions of credibility.
        </p>
        <sec id="sec-5-6-1">
          <title>Many participants cited unprofessional design, poor layout, and</title>
          <p>lack of reliable sources as reasons for low credibility scores.
Additionally, some expressed distrust towards Wikipedia and belief
that the information could be biased or incorrect.</p>
          <p>The justifications mainly indicate that the webpages appear
amateur, unprofessional, and lacking credible sources or evidence.
Additionally, many assessors express skepticism due to outdated
information, poor design, and potential bias from
advertisements or afiliate links.</p>
          <p>The webpages were considered less credible due to factors such as
unprofessional design, outdated information, informal
language, lack of evidence or citations, and clickbait-like content.
The credibility of some pages was questioned due to unfamiliar
brands, personal opinions, and the absence of proper
qualifications or accreditations.</p>
          <p>Assessors found some webpages to be credible due to professional
appearance, links to studies or association with reputable
organizations, while others were seen as less credible due to informal
language, lack of citations or references, and potential for errors.
The credibility of some pages was dificult to judge without further
investigation or knowledge of the subject matter.</p>
          <p>The justifications highlight the presence of credible sources,
professional appearance, and author qualifications as positive
factors for credibility. However, some concerns are raised due to
missing citations, outdated information, and potential biases.
Survey participants found the webpages credible due to their
professional appearance, use of medical facts and references,
reputable sources, well-structured content, and qualified authors.
The credibility was also often influenced by personal experiences
or previous knowledge about the subjects discussed.</p>
          <p>The majority of the justifications indicate that the webpages are
credible due to their professional appearance, reputable sources, and
being associated with trusted organizations such as the NHS, CDC,
and various academic journals. Additionally, assessors mentioned
the presence of scientific research, citations, author credentials,
and detailed information as contributors to the credibility of the
webpages.</p>
        </sec>
        <sec id="sec-5-6-2">
          <title>Summary of the Justifications</title>
          <p>1 (low credibility)</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>
        In this study, we showed that there is correlation between the annotations produced from
existing credibility guidelines and the end-user’s perception of credibility. This highlights the
value of existing guidelines [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] as proxies of credibility, thus endorsing these guidelines as a
roadmap in the complex and subjective task of credibility tagging.
      </p>
      <p>We also demonstrated that this relation is topic-dependent, as users tend to deviate more from
the reference ground truth when they are less familiar with the topic. Results also confirmed
previous research that suggests that people tend to overestimate credibility and, often, struggle
to identify sites labelled as low-quality. We also studied the influence of other variables, such
as the time of completion or the educational background, and some interesting conclusions
arose. Regardless of the educational background, people have dificulties judging the credibility.
This even happens with individuals that have a strong educational background (for example,
graduated students who have often been trained in skills such as critical thinking). It also
appears that “overthinking” and spending too much time to emit a judgement does not lead to
better estimates of credibility. Text-free justifications were also inspected, confirming a direct
relationship between certain elements from the credibility guidelines and user’s perceptions.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusions</title>
      <p>In this paper, we have conducted a user study on people’s perception to the credibility of online
health information. First, we used a previous study in the field to produce reference values
based on a series of guidelines. We found out a correlation between these values and the
judgements collected in the study. However, some worrying facts were also found: people
tend to overestimate the credibility of the sites (this can be specially damaging when health
information seeking) and it seems that the educational background has not a direct efect in
their perceptions. As future work, we want to diferentiate between closely related concepts
such as credibility (more subjective) or correctness (more factual) and study how they afect
users judgements.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>The authors thank: i) the financial support supplied by the Xunta de Galicia - Consellería de
Cultura, Educación, Formación Profesional e Universidades (Centro de investigación de Galicia
accreditation 2019-2022 ED431G-2019/04 and Reference Competitive Group accreditation
20212024, ED431C 2022/19) and the European Union (European Regional Development Fund - ERDF)
and ii) the financial support supplied by project PID2022-137061OB-C22 (Ministerio de Ciencia
e Innovación, Agencia Estatal de Investigación, Proyectos de Generación de Conocimiento;
supported by the European Regional Development Fund).</p>
      <p>The third author thanks the financial support obtained from project SUBV23/00002 (Ministerio
de Consumo, Subdirección General de Regulación del Juego).</p>
      <p>The authors also thank the funding of project PLEC2021-007662
(MCIN/AEI/10.13039/501100011033, Ministerio de Ciencia e Innovación, Agencia
Estatal de Investigación, Plan de Recuperación, Transformación y Resiliencia, Unión Europea-Next
Generation EU).
and mitigating online misinformation, IEEE Transactions on Computational Social Systems
(2023).
[17] A. L. Ginsca, A. Popescu, M. Lupu, et al., Credibility in information retrieval, Foundations
and Trends in Information Retrieval 9 (2015) 355–475.
[18] D. H. McKnight, C. J. Kacmar, Factors and efects of information credibility, in: Proceedings
of the ninth international conference on Electronic commerce, 2007, pp. 423–432.
[19] M. Viviani, G. Pasi, Credibility in social media: opinions, news, and health information—a
survey, Wiley interdisciplinary reviews: Data mining and knowledge discovery 7 (2017)
e1209.
[20] B. J. Fogg, C. Soohoo, D. R. Danielson, L. Marable, J. Stanford, E. R. Tauber, How do users
evaluate the credibility of web sites? a study with over 2,500 participants, in: Proceedings
of the 2003 conference on Designing for user experiences, 2003, pp. 1–15.
[21] M. S. Eastin, Credibility assessments of online health information: The efects of source
expertise and knowledge of content, Journal of Computer-Mediated Communication 6
(2001) JCMC643.
[22] C. N. Wathen, J. Burkell, Believe it or not: Factors influencing credibility on the web,</p>
      <p>Journal of the American society for information science and technology 53 (2002) 134–144.
[23] S. Sikdar, B. Kang, J. ODonovan, T. Höllerer, S. Adah, Understanding information credibility
on twitter, in: 2013 International Conference on Social Computing, IEEE, 2013, pp. 19–24.
[24] S. K. Sikdar, B. Kang, J. O’Donovan, T. Hollerer, S. Adal, Cutting through the noise:</p>
      <p>Defining ground truth in information credibility on twitter, Human 2 (2013) 151–167.
[25] S. Zimmerman, A. Thorpe, C. Fox, U. Kruschwitz, Privacy nudging in search:
Investigating potential impacts, in: Proceedings of the 2019 Conference on Human Information
Interaction and Retrieval, 2019, pp. 283–287.
[26] N. S. Chok, Pearson’s versus Spearman’s and Kendall’s correlation coeficients for
continuous data, Ph.D. thesis, University of Pittsburgh, 2010.
[27] J. Schwarz, M. Morris, Augmenting web pages and search results to support credibility
assessment, in: Proceedings of the SIGCHI conference on human factors in computing
systems, 2011, pp. 1245–1254.
[28] E. R. Carlson, Evaluating the credibility of sources: A missing link in the teaching of
critical thinking, Teaching of Psychology 22 (1995) 39–41.
[29] M. Kattenbeck, D. Elsweiler, Understanding credibility judgements for web search snippets,
Aslib Journal of Information Management (2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shepperd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Charnock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gann</surname>
          </string-name>
          ,
          <article-title>Helping patients access high quality health information</article-title>
          ,
          <source>Bmj</source>
          <volume>319</volume>
          (
          <year>1999</year>
          )
          <fpage>764</fpage>
          -
          <lpage>766</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Cline</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Haynes</surname>
          </string-name>
          ,
          <article-title>Consumer health information seeking on the internet: the state of the art</article-title>
          ,
          <source>Health education research</source>
          <volume>16</volume>
          (
          <year>2001</year>
          )
          <fpage>671</fpage>
          -
          <lpage>692</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fox</surname>
          </string-name>
          , Health topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          ,
          <source>Pew Internet &amp; American Life Project</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Pogacar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghenai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Smucker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <article-title>The positive and negative influence of search results on people's decisions about the eficacy of medical treatments</article-title>
          ,
          <source>in: Proceedings of the ACM SIGIR Int. Conf. on Theory of Information Retrieval</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>209</fpage>
          -
          <lpage>216</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Eysenbach</surname>
          </string-name>
          ,
          <article-title>Infodemiology: The epidemiology of (mis) information</article-title>
          ,
          <source>The American Journal of Medicine</source>
          <volume>113</volume>
          (
          <year>2002</year>
          )
          <fpage>763</fpage>
          -
          <lpage>765</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Eysenbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Powell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kuss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.-R.</given-names>
            <surname>Sa</surname>
          </string-name>
          ,
          <article-title>Empirical studies assessing the quality of health information for consumers on the world wide web: a systematic review</article-title>
          ,
          <source>Jama</source>
          <volume>287</volume>
          (
          <year>2002</year>
          )
          <fpage>2691</fpage>
          -
          <lpage>2700</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E. V.</given-names>
            <surname>Bernstam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Shelton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Walji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Meric-Bernstam</surname>
          </string-name>
          ,
          <article-title>Instruments to assess the quality of health information on the world wide web: what can our patients actually use?</article-title>
          ,
          <source>International journal of medical informatics 74</source>
          (
          <year>2005</year>
          )
          <fpage>13</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Vigdor</surname>
          </string-name>
          ,
          <article-title>Man fatally poisons himself while self-medicating for coronavirus, doctor says</article-title>
          ,
          <year>2020</year>
          . URL: https://www.nytimes.com/
          <year>2020</year>
          /03/24/us/chloroquine-poisoning-coronavirus. html,
          <source>[accessed June 9</source>
          ,
          <year>2022</year>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Fogg</surname>
          </string-name>
          , Persuasive technologie301398,
          <source>Communications of the ACM</source>
          <volume>42</volume>
          (
          <year>1999</year>
          )
          <fpage>26</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hahnel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Goldhammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Kröhne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>The role of reading skills in the evaluation of online information gathered from search engine environments</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>78</volume>
          (
          <year>2018</year>
          )
          <fpage>223</fpage>
          -
          <lpage>234</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kąkol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jankowski-Lorek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Abramczuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wierzbicki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Catasta</surname>
          </string-name>
          ,
          <article-title>On the subjectivity and bias of web content credibility evaluations</article-title>
          ,
          <source>in: Proceedings of the 22nd international conference on world wide web</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>1131</fpage>
          -
          <lpage>1136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fernández-Pichel</surname>
          </string-name>
          , S. Meyer,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Frummet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Elsweiler</surname>
          </string-name>
          ,
          <article-title>Improving the reliability of health information credibility assessments</article-title>
          ,
          <source>in: Proceedings of the 3rd Workshop on Reducing Online Misinformation through Credible Information Retrieval</source>
          <year>2023</year>
          co
          <article-title>-located with The 45th European Conference on Information Retrieval (ECIR</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>50</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3406</volume>
          /paper4_jot.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Fogg</surname>
          </string-name>
          ,
          <article-title>Prominence-interpretation theory: Explaining how people assess credibility online, in: CHI'03 extended abstracts on human factors in computing systems</article-title>
          ,
          <year>2003</year>
          , pp.
          <fpage>722</fpage>
          -
          <lpage>723</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Unkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Haas</surname>
          </string-name>
          ,
          <article-title>The efects of credibility cues on the selection of search engine results</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>68</volume>
          (
          <year>2017</year>
          )
          <fpage>1850</fpage>
          -
          <lpage>1862</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Sharif</surname>
          </string-name>
          ,
          <article-title>A review on credibility perception of online information</article-title>
          ,
          <source>in: 2020 14th International Conference on Ubiquitous Information Management and Communication (IMCOM)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bodaghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Watine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <article-title>A literature review on detecting, verifying,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>