<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multimodal Online Manipulation: Empirical Analysis of Fact-Checking Reports</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Uryupina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering and Computer Science, University of Trento</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an in-depth exploratory quantitative study of the interaction between multimedia and textual components in online manipulative content. We discuss relations between content layers (such as proof or support) as well as unscrupulous techniques compromising visual content. The study is based on fakes reported and analyzed by PolitiFact and comprises documents from Facebook, Twitter and Instagram. We identify several pervasive phenomena currently, afecting the impact of manipulative content on the reader and the possible strategies for efective de-bunking actions, and discuss possible research directions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;fact checking</kwd>
        <kwd>multi modal</kwd>
        <kwd>annotation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>understanding of the way the authors integrate
multimedia into their content: most research so far has focused
Manipulative online content (fake news, propaganda, on a specific component and not on their interplay. Our
among others) is growing at an alarming rate, hinder- study aims at identifying the role of multimedia part of
ing our access to truthful and unbiased information and manipulative messages.
thus threatening principles of the democratic society. Figure 1 shows some examples from potential fakes
The problem has been addressed by professional jour- analyzed by PolitiFact. We observe diferent relations
nalists, who – with the help of crowd-workers – fight a between the text and the image. In particular, in (1a),
never-ending battle to prevent information contamina- the video is supposed to prove the claim by providing
tion. To enable a large-scale response to the misinforma- direct evidence, whereas in (1b), the image provides a
tion threat, the AI community has invested a considerable support (appeal to authority). In (1c), the image is a
viefort into building competitive models for identifying sual paraphrase of the claim, enhancing its appeal but
non-transparent content, such as false claims or altered not providing extra proof, support or informational
mavideos (deep fakes). However, we still lack a thorough terial. Finally, in (1d), the photo is an illustration that,
understanding of the manipulative content and multi- while depicting the discussed person, does not aim at
ple aspects afecting its perception and impact on the being relevant to the claim’s veracity or impact. While
reader. This paper aims at an in-depth analysis of one of understanding the relation between the image and the
such aspects, namely, the interaction between diferent text is interesting from the scientific perspective, it is
(multimedia) layers of the manipulative message. More also a crucial prerequisite for eficient and meaningful
specifically, we study the semantics underlying the re- fact-checking response. For example, if a supposed proof
lation between multimedia and textual parts of the fake is a compromised photo, the response should highlight
news. Our study is based on around 800 fakes from Jan- this fact (e.g., the video in (1a) has been cropped
misuary till September 2022, as identified and analysed by representing the quote, which should be highlighted in
PolitiFact.1 the fact-checking report). On the contrary, if a
compro</p>
      <p>Multimedia content, such as videos, reels, photos, mised photo is used as a mere illustration, the efective
screenshots or images is becoming increasingly popu- fact-checking report should focus on the textual claim
lar in social media: it is an appealing and powerful way per se.
of expressing and/or enhancing one’s message. Never- Another important angle is the issue with the
multitheless, as a scientific community, we still have little media part. In our example, the video in (1a) is cropped.
On the contrary, (1b) represents an authentic screenshot,
yet, it has been miscaptioned by the claim: an older
content, irrelevant for the current events/topics, has been
CLiC-it 2024: Tenth Italian Conference on Computational Linguistics,
Dec 04 — 06, 2024, Pisa, Italy
$ uryupina@gmail.com (O. Uryupina)</p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License repurposed.
1PolitiFactAt(trhibtuttipons:4/.0/Iwntewrnawtio.npaol(lCiCtiBfYa4c.0t)..com/) is an independent journal- The current paper focuses on these two aspects to
anistic agency and one of the most experienced fact-checking orga- alyze empirically the interplay between multimedia and
nizations, providing detailed analytics for non-transparent online textual components in fake news, as identified by
Politicontent since 2007.</p>
      <p>
        (a) Biden to teachers: “They’re not somebody else’s
children. They’re yours when you’re in the classroom."
(VIDEO)
(b) Now you know why there’s suddenly "a formula
shortage". The new age robber barrons have conveniently
invested in some unholy breast milk made from
human organs.
(c) In honor of #TaxDay, I remind you that Governor Evers
wanted to increase your taxes by $1 billion just for
heating your homes. Instead, Republicans cut your
taxes by more than $2 billion.
(d) Italian football agent Mino Raiola has died after
suffering from an illness. RIP
Fact. To this end, we reannotate the PolyFake dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] extra content to the textual message. Cheema et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
with fine-grained labels reflecting multimedia aspects. propose a dataset of multimodal tweets, annotated for
visual relevancy and checkworthiness. Finally, Biamby
et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] propose a larger-scale dataset of multimodal
2. Related Work tweets, where "falsified" claims have been added
synthetically to address the image repurposing problem.
      </p>
      <p>
        While fact checking has been receiving an increasing These studies have paved the way for evaluation
camamount of attention recently both from NLP and Vision paigns and benchmarking resources, for example, [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
communities, only very few studies focus on the interac- Yet, these studies rely on rather straightforward
annotation between diferent modalities. tion guidelines to reduce the per-claim cost. Moreover,
      </p>
      <p>
        A breakthrough approach by Vempala and Preoţiuc- the annotators are not professional fact-checkers: while
Pietro [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] focuses on two dimensions of the relationship they can assess some aspects of the compromised content,
between text and image on Twitter: whether the text is they still can get deceived by more challenging cases –
represented in the image and whether the image adds after all, the manipulative content has been created on
Layer
none
video
photo
screenshot
link
image
thread
total
purpose to influence and bias the reader. is based on the first nine months of PolyFake (818
en
      </p>
      <p>
        In a recent survey, Mubashara et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] highlight tries). Each entry has been re-assessed by two annotators,
the importance of an interdisciplinary approach to fact- with further adjudication by the supervisor. The
origichecking, proposing a framework to model diferent axes nal PolyFake labels are binary and encode more generic
of online manipulation, most importantly, fusing the tex- properties of fake news (e.g. whether the reasoning is
tual and visual fact-checking and survey benchmarks and fallacious or whether the document triggers emotions).
models developed by respective communities. Our study For the present study, we have designed and iteratively
is built upon the same motivation – and our main goal refined annotation guidelines for labelling multimedia
is to study empirically the interplay between diferent aspects of manipulative content.
modalities, based on real-world (i.e., not simulated or The annotation process is based on consulting jointly
synthesized) fakes data. not only the original content, but the PolitiFact report as
      </p>
      <p>Our study aims at an in-depth exploratory analysis well. This way we make use of the wealth of analytics
of the multimodal online content. To this end, we focus provided by experienced professional fact-checkers by
on more specific labels to describe the relationship be- encoding it in more structured annotation labels.
tween diferent layers/modalities. We extend the scope PolyFake covers fakes from diferent social media
of our study to cover all the three major platforms (Face- (Twitter, Facebook, Instagram, TikTok, Threads and
book, Instagram and Twitter). Moreover, our input is YouTube). Note that manipulative content often gets
not only the claim per se, but the professionally created propagated across platforms through re-posts, sharing,
fact-checking report from PolitiFact. In our experience, linking or just copying. For example, a large
proporPolitiFact reports contain a wealth of information about tion of Facebook videos originates from TikTok (in this
online manipulation: as opposed to 2-3 binary labels of case, PolitiFact typically analyzes the Facebook message,
common NLP fact-checking benchmarks, PolitiFact char- hence a low number of TikTok entities in the table). In
acterizes each claim with 1-3 pages of analytics. This the following study, we omit TikTok, YouTube and
Teleanalytics, however, comes in a free textual form. While gram as largely underrepresented categories with rather
it might be still impossible for the NLP community to en- straightforward patterns.
code these reports for building high-quality fact-checking
systems, we believe that we should at least learn from 3.2. Multimedia and Layered Content
them to get better insights, stop trivializing the task and
highlight understudied, yet impactful, subtasks.</p>
    </sec>
    <sec id="sec-2">
      <title>3. Analyzing Multimedia Content</title>
      <p>
        3.1. PolyFake
Our study is based on the PolyFake dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] covering
fake news from 2022, as analyzed by professional
factcheckers from the PolitiFact agency.2 The current study
      </p>
      <sec id="sec-2-1">
        <title>2PolyFake annotation guidelines cover a wide range of phenomena</title>
        <p>related to online manipulation: from fallacious/propaganda
reasoning to emotive appeals, factual veracity etc. Current study aims at
an in-depth analysis of a specific angle. The Appendix discusses
the distribution of veracity labels across PolyFake documents.</p>
        <p>Layer Types. Table 1 shows the distribution of diferent
media types for each platform. We have identified several
types of layered content: parts of the message rendered
together with the initial post. The most common ones
are videos (including reels), photos and screenshots
(typically, complex visual objects combining textual content
with photos/images and referring the reader to a
diferent source). We have also observed images (infographics,
maps or drawings), links (this content typically is
rendered with a photo/stillshot, yet it explicitly points to a
diferent online location, for example, promotion
website) or threads (characteristic for Twitter, this type of
layering helps to contextualize the message). On rare
occasions, social media posts might contain more than
role
content
anchor
proof
support
paraphr.
context
illustr.
action
other
total
one extra layer (e.g., videos and photos). dia levels play in PolyFake documents. We distinguish</p>
        <p>Most importanly, only 18% of PolyFake documents are between the following roles: content (the essential part of
purely textual: adhering to the popular adage that a pic- the content is presented on the multimedia layer, whereas
ture is worth a thousand words, manipulative content the textual layer just adds minor details or suggests
opincreators use visuals for a variety of purposes, from in- ions), proof (the multimedia layer ofers a physical proof –
creasing the outreach to improving the credibility. More- cf. Example (1a)), support (the multimedia layer provides
over, the prevalence of multimedia content is way more some material to support the claim, from a reputable
critical for Facebook and Instagram – the two platforms source – cf. Example (1b)), paraphrase (the multimedia
not typically addressed by NLP practitioners. This alone layer paraphrases the claim without adding any extra
suggests that we need to pay much more attention to joint angle – cf. Example (1c)), context (while the textual claim
models and start with deeper understanding of relevant is generally self-contained, it cannot be interpreted
withphenomena. out the context given by the multimedia part (e.g., the</p>
        <p>A large percentage of documents are re-using or claim contains pronouns and the image presents their
spreading already existing information. This is true for referents)), illustration (the multimedia layer shows some
screenshots (21% in total) and links (5%), but also for objects/persons mentioned in the claim without any
conmany videos – only very few videos represent original nection to its semantics – cf. Example (1d)) and action
content. While there exist some studies on identifying (the multimedia layer suggests an appropriate reaction to
previously fact-checked claims, they are restricted to the the claim, for example, a scam website). Finally, a rather
textual content. We believe that a more complex multi- common role for videos and photos is anchor: in such
modal approach would be beneficial here. cases, the textual claim is about the multimedia itself (for</p>
        <p>For presentation issues, in what follows we merge our example, "the sharpest image of the sun ever recorded.";
underrepresented categories link, image and thread with here, the multimedia is not compromised per se and the
roughly functionally similar major categories screenshot, textual claim contains no falsehoods about the world, yet
photo and screenshot respectively. the combination might be very misleading.</p>
        <p>Layer Roles. Table 2 shows diferent roles multime- In more than half of the documents, multimodal layers
provide essential content. This is true for all the media
types (videos, photos and screenshots). We have observed
several possible factors contributing to this efect: in
general, social media users tend to repost existing "fancy"
content and not create their own texts. Even in authentic
self-created posts, the message is often put in a visual,
whereas only some emotions are added in a text. We
believe that there is a wide variety of potential reasons
for this behaviour (e.g., videos and photos get more likes,
whereas texts are mostly ignored by peers), requiring a
more specialized study.</p>
        <p>Almost one third of multimedia layers, especially
videos, supposedly present proofs. Such compromised
proofs are out of reach for the modern evidence-based
automatic fact-checking: while a fact-checking model
can provide extensive evidence to refute a claim, the user
would still trust the video/photo and not the model.
Human fact-checkers address such proofs from a diferent,
more promising, perspective: they try to explicitly
attack and debunk the proof. We believe that this is a very
important and largely unaddressed research direction.</p>
        <p>Issues with multimedia layers. Finally, we have
identified the most common unscrupulous techniques
relevant for multimedia layers. Those include: crop
(essential part(s) of the original message are omitted to
render it out of context – cf. Example (1a)); miscaption (while
the image/video is authentic, the textual claim misleads
w.r.t. some crucial details, e.g. events or timeline – cf.</p>
        <p>Example (1b)); altered/fake (the image/video has been
altered – beyond cropping – with the specialized software,
including deep fakes); misperception (the image/video is –
deliberately or not – deceiving because of its low quality,
unclear angle, optical efects etc); noproof (the – typically
long – video does not contain any components relevant
for the claim); falsehood (the video/image is authentic,
yet its content is untrue – i.e., the textual claim spreads
the original fake generated by the video/image); and
explain (the textual part explains – misleadingly – what we
are supposed to see in the video, often of a rather low
quality).</p>
        <p>Table 3 summarizes the distribution of problematic
issues across the three main multimedia types, showing
several trends. First, video layers provide more
possibilities for unscrupulous content generators: cropped,
otherwise altered or low quality videos are pervasive
in manipulative content. While most of the research
focuses on images, they do not exhibit such a variety
of manipulative strategies. Screenshots – authentic or
fake – are largely used to disseminate falsehoods. At
the same time, an increasing amount of authentic videos,
mostly originating from TikTok, is created to spread
falsehoods and promote "critical thinking" (i.e., conspiracy
theories as opposed to rational argumentation). These
remain largely understudied, despite their large impact
on the audience. Another rather unstudied area are
explanatory claims: authentic videos/photos accompanied
by misleading explanations of what we see and what it
means; in such cases, the factual component might be
non-compromised, yet the biased explanation makes the
whole message an impactful and hard to debunk
propaganda tool. Finally, unlike videos and screenshots, most
photos represent true authentic information – the textual
claims either rely on them as illustrations or use them as
building blocks to support fallacious argumentation.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusion</title>
      <p>We have presented an in-depth analysis of the
interaction between textual and multimedia components of
compromised social media documents. We have identified
several high-impact issues, insuficiently studied by the
community at the moment. These include the interaction
between diferent modalities, the role of the
multimedia part and its impact on selecting the successful
factchecking strategy, the diference between platforms and
media types (current NLP studies predominantly focus
on Twitter and images) and the importance of a more
principled approach to content re-use. We hope that this
study, motivated by human fact-checking expertise, can
sparkle a meaningful discussion and improve automatic
modeling.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>We thank the Autonomous Province of Trento for the ifnancial support of our project via the AI@TN initiative.</title>
        <p>FC label
pants-on-fire
false
mostly false
half true
mostly true
true
total</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>A. True vs. Fake content and multimedia layers</title>
      <sec id="sec-5-1">
        <title>Our dataset by construction contains mostly untrue</title>
        <p>claims: even though PolitiFact occasionally fact-checks
statements that turn out to be true, most of their
materials are "false", "mostly false" or even "pants on fire".
Moreover, even true claims often exhibit signs of user
manipulation. In this appendix, we show statistics for
fake vs. true content in PolitiFact reports (Table 4).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Anonymous</surname>
          </string-name>
          ,
          <article-title>PolyFake: Fine-grained multiperspective annotation of fact-checking reports</article-title>
          , in: Accepted for publication,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vempala</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          Preoţiuc-Pietro,
          <article-title>Categorizing and inferring the relationship between the text and image of Twitter posts, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>2830</fpage>
          -
          <lpage>2840</lpage>
          . URL: https://aclanthology.org/P19-1272. doi:
          <volume>10</volume>
          .18653/ v1/
          <fpage>P19</fpage>
          -1272.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Cheema</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hakimov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sittar</surname>
          </string-name>
          , E. MüllerBudack, C. Otto,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ewerth</surname>
          </string-name>
          ,
          <article-title>MM-claims: A dataset for multimodal claim detection in social media, in: Findings of the Association for Computational Linguistics: NAACL 2022, Association for Computational Linguistics</article-title>
          , Seattle, United States,
          <year>2022</year>
          , pp.
          <fpage>962</fpage>
          -
          <lpage>979</lpage>
          . URL: https: //aclanthology.org/
          <year>2022</year>
          .findings-naacl.
          <volume>72</volume>
          . doi:
          <volume>10</volume>
          . 18653/v1/
          <year>2022</year>
          .findings-naacl.
          <volume>72</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Biamby</surname>
          </string-name>
          , G. Luo,
          <string-name>
            <given-names>T.</given-names>
            <surname>Darrell</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Rohrbach, TwitterCOMMs: Detecting climate, COVID, and
          <article-title>military multimodal misinformation, in: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , Seattle, United States,
          <year>2022</year>
          , pp.
          <fpage>1530</fpage>
          -
          <lpage>1549</lpage>
          . URL: https://aclanthology. org/
          <year>2022</year>
          .naacl-main.
          <volume>110</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          . naacl-main.
          <volume>110</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bondielli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dell'Oglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lenci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Marcelloni</surname>
          </string-name>
          , L. Passaro,
          <article-title>Dataset for multimodal fake news detection and verification tasks, Data in Brief 54 (</article-title>
          <year>2024</year>
          )
          <article-title>110440</article-title>
          . URL: https://www.sciencedirect.com/ science/article/pii/S2352340924004098. doi:https: //doi.org/10.1016/j.dib.
          <year>2024</year>
          .
          <volume>110440</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mubashara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Michael</surname>
          </string-name>
          , G. Zhijiang,
          <string-name>
            <given-names>C.</given-names>
            <surname>Oana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Elena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Andreas</surname>
          </string-name>
          ,
          <source>Multimodal automated factchecking: A survey</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2305</volume>
          .
          <fpage>13507</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>