<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Bozen-Bolzano, IT
$ davide.ceolin@cwi.nl (D. Ceolin)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>CrowdIQ: An Ontology for Crowdsourced Information Quality Assessments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Ceolin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dafne van Kuppevelt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ji Qi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centrum Wiskunde &amp; Informatica</institution>
          ,
          <addr-line>Science Park 123, 1098XG, Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Netherlands eScience Center</institution>
          ,
          <addr-line>Science Park 402, 1098 XH Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Fact-checking is a common journalistic practice adopted to verify the truthfulness of claims and information items. Because of the demanding nature of fact-checking, a significant amount of research has been devoted to the use of crowdsourcing to scale up this practice. The idea to use laypeople to fact-check information allows accessing a vast amount of human computation resources, but introduces an issue of reliability: when these tasks are performed by laypeople instead of experts, their quality might be questioned. In this paper, we introduce an ontology for modeling crowdsourced datasets of information quality assessments. We emphasize that we allow modeling information about the items evaluated as well as important metadata such as the authors of such assessments. The goal of this model is to favor interoperability among diferent datasets of the same kind, as well as to support internal analyses of the dataset themselves in terms of bias and reliability of the collected assessments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Information Quality</kwd>
        <kwd>Crowdsourcing</kwd>
        <kwd>Data Modeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Fact-checking is a common journalistic practice employed in order to test the veracity of claims
to be reported in newspapers. In order to fact-check information, journalists look up
information sources, appraise their reliability, extract information out of them, and then reason on
the resulting information set. This sequence of operations is rather time-consuming, while
the amount of claims that require fact-checking online is constantly growing. On the hand,
this means that the workload of specialists is high and, on the other hand, their possibility to
intervene real-time is limited. Crowdsourcing has revealed a useful tool to allow scaling up
fact-checking. Crowd workers can, in fact, be instructed so to produce expert-like assessments,
and by collecting multiple assessments about the same item (wisdom of the crowds), quality
can be assured. In order to maximize the usefulness of the fact-checking (and, more in general,
information quality assessment) datasets obtained from crowdsourcing, it is important to
annotate them with specific metadata. This paper describes a lightweight ontology to describe and
annotate crowdsourced information quality assessments. The ontology allows characterizing
information about the items being assessed, the author of the assessment, and information
quality details. We leverage existing ontologies like the Data Quality Vocabulary (DQV) and
Schema.org (Schema) and we select and specialize their elements to serve this specific niche
of information. The paper is structured as follows. Section 2 presents related work. Section 3
introduces the ontology, while Section 4 discusses example applications. Section 5 concludes.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Relevant to this line of work is the Data Quality Vocabulary (DQV) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], that allows defining
quality dimensions and metrics for annotating quality of data. Our model extends and specializes
DQV in order to model specifically crowdsourced information quality assessments. In turn, this
links to the extensive line of research on data quality modeling and measuring [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While we
can observe similarities with the measurement and modeling of data quality, the information
quality assessments we are interested in aim at assessing the information content of items,
rather than their data serialization. For an overview on information quality and its philosophy,
we refer the reader to the book edited by Floridi and Illari [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The Data Cube Vocabulary [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is
also a predecessor of our model, since it allows modeling multidimensional metadata. We see
our model as a specialization of this vocabulary as well.
      </p>
      <p>
        The work presented in this paper is also relevant to the field of FAIR data principles, as it
aims at favoring findability (especially on principle F2) and interoperability (I1) of data. We
refer the reader to the work of Poveda et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] who provide an extensive analysis on this topic.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The CrowdIQ Ontology</title>
      <p>The ontology we present here aims at achieving the highest interoperability while allowing
modeling the necessary information resulting from the crowdsourcing tasks of interest. In the
ontology, we identify three main “superclasses” that represent the core of the information we
intend to model. We describe them as follows, while Figure 1 provides an overview.
Target Item The target item represents the item to be evaluated. This ontology means to
model quality assessments on information items online, therefore this class specializes the
DigitalDocument class of Schema.org.</p>
      <p>
        Worker The worker class is meant to characterize the author of a quality assessment, and is
thus modeled from the Person class of Schema.org. The instances of this class could be more or
less populated depending on the level of anonymity granted to the crowd workers. Information
to populate instances of the worker class can be provided both by the crowdsourcing platform
or by the worker who responds to demographic questions in the crowdsourcing task.
IQ Assessment This class models the information quality assessment that the worker provides
about the target item. This class requires the specification of a quality dimension (e.g., precision,
accuracy, truthfulness; see [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for an overview of information quality dimensions) and of a
metric. E.g., the precision of an online document might be expressed using a 5-level Likert scale.
We also record provenance information (time, platform of creation).
rdf:subClassOf
:target_item
rdf:subClassOf
:worker
:platform
dqv:hasQualityMeasurement
prov:wasAttributedTo
dqv:inDimension
:iq_assessment
rdf:subClassOf prov:atLocation
      </p>
      <p>prov:startedAtTime
dqv:isMeasurementOf "^Y^xYsYdY:d-MatMeT-DimDeTHH:MM:SSZ"
dqv:dimension</p>
      <p>dqv:metric</p>
    </sec>
    <sec id="sec-4">
      <title>4. Applications</title>
      <sec id="sec-4-1">
        <title>4.1. Data Conversion Requirements</title>
        <p>
          The proposed ontology is meant to produce and annotate Linked Data representation of
crowdsourced information quality assessments. Given that these data are often produced in CSV
format, we refer the reader to the CSV on the Web standard [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] to convert CSV files in RDF or
JSON-LD. We also provide an example metadata file online 1 for converting CSV files using our
ontology.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Analysis workflow in Knime</title>
        <p>
          In the real-world research project “The Eye of the Beholder” 2, we aim to building or leveraging
a platform for scholars to train machine learning pipelines to automatically assess quality of
online information. To train these pipelines, we make use of a dataset collected from a crowd
sourcing platform called CrowdFrame [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The training data contains information from three
parts, namely worker information, document information, and assessment information. Details
about the corresponding crowdsourcing experiment is available at [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          Through investigation and comparison, we chose the KNIME [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], a free and open-source data
analysis platform. KNIME integrates various components for machine learning, data mining,
and reporting. Users can visually create workflows with these components and execute part or
all of them, and subsequently use interactive views to examine results, models, etc.
1Available at https://github.com/EyeofBeholder-NLeSC/assessments-ontology/blob/main/metadata.json
2https://www.esciencecenter.nl/projects/the-eye-of-the-beholder-transparent-pipelines-for-assessing-onlineinformation-quality/
        </p>
        <p>The desired tool is composed of several KNIME workflows working in tandem with each other.
These workflows automate the process of data exploration, model training, result interpretation,
and pipeline comparison. To achieve this goal, we designed the CrowdIQ ontology to design
the data interfaces for these workflows, and we also need to validate the training data files
corresponding to this ontology before feeding the to those workflows.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. User stories</title>
        <p>The requirements for integrating the ontology in KNIME are specified as follows.
Upload data from the crowdsourcing platform Once the crowd sourcing task has been
completed by enough workers, our user wants to load the resulting assessments into KNIME for
further analysis. The platform provides our user with a URL of the metadata of the file (the url
of the data itself can be inferred from the metadata). The KNIME component that we develop
takes this URL as an input. The component downloads the data and metadata, and interprets
the data according to the CrowdIQ ontology. It executes some sanity checks: it returns an error
when, for example, no documents or no users are present. It outputs three tables: for workers,
documents and assessments respectively. It also outputs a visualization of the realization of
some of the attributes in the data: e.g., which dimensions and worker attributes are present.
Specify metadata When the user wants to provide data from a platform that does not give
machine readable metadata compatible with our ontology, the user needs to specify which fields
relate to which attributes in the data. Instead of providing an URL to the metadata as above, the
user should be able to specify in a graphical user interface, for each column of their CSV data,
which attribute in the ontology it represents. Depending on the format of the data that the user
provides, the CSV data can either consist of one table with all information combined, or three
diferent tables for workers, assessments, and documents respectively.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Proposed solutions</title>
        <p>To account for the user stories as described in the previous section, we propose three types of
validation to integrate in KNIME workflows. First, if the training data is stored in CSV format,
a CSV-on-the-Web validation is essential. In this step, metadata annotations are added to the
input CSV files (manually or by referring to an existing metadata to interpret the data files
(encoding, data types of fields, etc.). At the same time, these files will be converted into a linked
data format. Next, the obtained linked data should also be tested for consistency with the
ontology. This is done through an ontology-based validation that maps the data fields to the
expected classes/properties defined in the ontology and further cleans the data. Finally, the
resulting data should also be validated against the constraints defined in the ontology. This can
include various checks such as missing/duplicate properties, missing or improper type arcs,
and inconsistent value ranges. In addition, we propose to extend the CrowdFrame platform to
output a metadata file describing the output CSV according to the CSV on the web standard. In
the interface for creating a new crowd sourcing task, the user should be able to specify which
ifelds correspond to which attributes in the ontology, for those fields for which it cannot be
inferred automatically. A prototype of a Knime pipeline that reads CSV and validates against
the ontology, can be found on GitHub3. It wraps the csvw library4 in a Python scripting node.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we introduce CrowdIQ, an ontology for modeling crowdsourced information
quality assessments. The ontology aims at achieving minimal commitement, while specializing
existing models to represent common information about information quality crowdsourced
tasks. The ontology is described along with a possible utilization scenario. The goals of this
model are twofold: allow interoperability among existing crowdsourced datasets, as well as to
allow inspecting each dataset separately in order to assess its reliability and bias. Future work
will aim at extending the model further and at providing pipeline automation components that
leverage the model’s expressivity.</p>
      <sec id="sec-5-1">
        <title>Acknowledgements</title>
        <p>3https://github.com/EyeofBeholder-NLeSC/assessments-ontology
4https://github.com/cldf/csvw</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Albertoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Isaac</surname>
          </string-name>
          ,
          <article-title>Introducing the data quality vocabulary (dqv</article-title>
          ),
          <source>Semant. Web</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>81</fpage>
          -
          <lpage>97</lpage>
          . URL: https://doi.org/10.3233/SW-200382. doi:
          <volume>10</volume>
          .3233/SW-200382.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V. C.</given-names>
            <surname>Storey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Modeling quality requirements in conceptual database design</article-title>
          ., in: I. N.
          <string-name>
            <surname>Chengalur-Smith</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          Pipino (Eds.), IQ, MIT,
          <year>1998</year>
          , pp.
          <fpage>64</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          , P. Illari (Eds.),
          <source>The Philosophy of Information Quality</source>
          , Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Richard</surname>
          </string-name>
          <string-name>
            <surname>Cyganiak</surname>
          </string-name>
          ,
          <source>The RDF Data Cube Vocabulary</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Poveda-Villalón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Espinoza-Arias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garijo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Corcho</surname>
          </string-name>
          ,
          <article-title>Coming to terms with fair ontologies, in: Knowledge Engineering and Knowledge Management: 22nd International Conference</article-title>
          ,
          <source>EKAW 2020</source>
          , Springer-Verlag, Berlin, Heidelberg,
          <year>2020</year>
          , p.
          <fpage>255</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Tennison</surname>
          </string-name>
          ,
          <article-title>Csv on the web: A primer</article-title>
          , https://www.w3.org/TR/tabular-data-primer/,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soprano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Roitero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bombassei De Bona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          ,
          <article-title>Crowd_frame: A simple and complete framework to deploy complex crowdsourcing tasks of-the-shelf</article-title>
          ,
          <source>in: Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, WSDM '22</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2022</year>
          , p.
          <fpage>1605</fpage>
          -
          <lpage>1608</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soprano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Roitero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. La</given-names>
            <surname>Barbera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ceolin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Demartini,</surname>
          </string-name>
          <article-title>The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale</article-title>
          ,
          <source>Information Processing Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <fpage>102710</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Berthold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cebron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Gabriel</surname>
          </string-name>
          , T. Kötter,
          <string-name>
            <given-names>T.</given-names>
            <surname>Meinl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ohl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Thiel</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Wiswedel,</surname>
          </string-name>
          <article-title>KNIME: The Konstanz Information Miner, in: Studies in Classification, Data Analysis, and Knowledge Organization (GfKL</article-title>
          <year>2007</year>
          ), Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>