<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Challenges in Disinformation Difusion Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paolo Papotti papotti@eurecom.fr EURECOM</string-name>
        </contrib>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>THE NEED FOR BETTER DIFFUSION NETWORKS</title>
      <p>
        Social media enable fast and widespread dissemination of
information that can be exploited to efectively spread disinformation
by bad actors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We refer to disinformation as the malicious
and coordinated spread of inaccurate content for manipulation
of narratives1. It has been showed that social media
disinformation has efectively reached millions of people in state-sponsored
campaigns2.
      </p>
      <p>
        Several computational solutions have been proposed for the
identification of coordinated campaign on a single platform [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
They study how content is disseminated across a network of
inter-connected users. However, two main practical challenges
limit the impact of such approaches.
      </p>
      <p>First, existing approaches focus on a single sources, such as
Twitter or Reddit. Unfortunately, misinformation campaigns span
multiple platforms and there is a recognized need to jointly
analyze the difusion of content across diferent sources, such as
social networks, online forums (e.g., Reddit), and traditional news
outlets (e.g., comments in reputable sites).</p>
      <p>
        Second, the content difusion graphs that are currently
generated from social network APIs are limited in quality. For example,
only the information about content re-posting (e.g., re-tweets) of
a user is directly provided. But information is disseminated also
by manipulating the original content to add bias, “evidence”, or
propaganda material. Moreover, fine granularity of the re-posting
is not available, with the recognized problem of star-efect for
re-tweets that can heavily degrade the quality of the network
model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>Consider the example in Figure 1 that shows the coordinated
sharing of the same initial piece of content (say, a textual news)
by three users over diferent platforms. With the current
infrastructure and APIs, a journalist or a fact-checker willing to study
the difusion network would look at each network in isolation.
S/he would be able to follow content across users (nodes) only
when they re-post explicitly (full edges across nodes).</p>
      <p>
        In this example, the information would not be enough for the
early identification of the coordinated campaign started by the
three users. Looking at only one source with limited information
does not enable the analytics, neither in terms of scope nor
evidence, that we need to identify and understand how false and
biased content is used in online campaigns [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        To overcome this limitation, recent approaches explore
evidence across users and platforms, such as coordinated link
sharing [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. While this signal has proven to be useful, we believe this
1The observations in this work apply also for misinformation, where actors spread
incorrect content unintentionally.
2E.g., https://transparency.twitter.com/en/reports/information-operations.html
is just one example of the richer kind of metadata that is needed
for better difusion networks.
      </p>
      <p>In fact, to tackle the first challenge, difusion networks should
be heterogeneous, covering multiple platforms, with the ability to
recognize the same content and the same users across services.
In Figure 1, users that refer to the same real world person are
annotated with nodes having the same color. Also, to handle the
second challenge, the edges should be typed with fine-grained
metadata that model diferent actions in the spread of the content.</p>
      <p>
        We believe that data here plays a role as important as the
algorithms used for the analysis and therefore more attention is
deserved to the problem of creating such richer difusion networks.
Their creation can lead to better identification of coordinated
eforts [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and ultimately allow the analysis of disinformation
campaigns in terms of actors, space, and time.
      </p>
      <p>The goal is therefor to develop methods for modeling and
creating the rich difusion networks from the existing platform
APIs. The resulting networks can be exploited to assist users,
such as fact-checkers and journalists, in
(1) monitoring sources at scale and recognize misleading
information (in terms of false or biased textual content) on
social networks and forum websites;
(2) tracking the spread and difusion of the content in terms
of time and actors;
(3) generating visualizations that support the fight against
misinformation and related literacy eforts.</p>
      <p>This network generation is indeed challenging, as the desired
metadata is not available and hard to profile automatically in an
accurate way. We discuss next two research directions that we
identify as critical to tackle these challenges.</p>
    </sec>
    <sec id="sec-2">
      <title>2 RESEARCH DIRECTIONS</title>
      <p>The goal is to develop methods for the automatic modeling of
content manipulation and difusion across time and diferent
media sources, such as social networks, forums, and news outlets.</p>
      <p>Not only we want the difusion graph for a given content to be
across sources and very well described in terms of information,
we also want it (i) to preserve precisely the provenance of the
data (who created and shared, how and when) and (ii) to be as
much as possible automatic in its creation, both to handle the
Web scale and to not put additional burden on the users. There
are therefore several challenges that we need to overcome
• diferent sources do not contain any readily available
information to connect users/content across networks and
the automatic matching is a dificult task in both cases;
• labeling the content in terms of being false, manipulated,
and biased requires deep understanding of the language
and of the reference and background information;
• Web scale implies massive ingestion from heterogeneous
data sources, but we would prefer tools that can be used
by end users on their machines for confidentiality;
• support for diferent languages as we aim at helping users
across diferent countries.</p>
      <p>Given the challenges above, a natural first line of work is to
conduct data integration research to generate a unified
representation from heterogeneous, non-aligned sources. A second line of
work is to deploy natural language processing (NLP) techniques
to profile the content and enrich the graph with typed nodes and
edges.</p>
      <p>
        Data integration. In the first line of research, the aim is the
online creation of a dissemination network for a given textual
content. Given a textual article, for example, the first task is the
identification of its citations and appearance across sources
(online articles, boards in forums, social posts) and time. This is not
trivial as one requirement is to go beyond the identification of
content by links, which act as unique identifiers. For this goal,
one promising direction is to exploit text-matching literature [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
to identify also manipulated texts that express the original input
content. The goal is to have one difusion network, as in Figure 1
for every given content to analyze, such as a web page, a social
message post, or a generic textual claim. The linking and merging
of actors across sources, nodes in the graphs, is also important.
This can be modelled as an entity resolution problem, from a data
integration perspective, for example by using deep learning
techniques [
        <xref ref-type="bibr" rid="ref15 ref4">4, 15</xref>
        ]. However, the task is especially challenging in real
settings where we drop assumptions about trusted information
about the user accounts.
      </p>
      <p>
        Example 1. Consider a textual article  about a new vaccine
that circulates on social platform 1. Existing APIs allow the
tracing of the difusion of the specific content  on platform 1 across
users 11, . . . , 1 , but the same content may be circulating in a
diferent form ( ′ ) and across a diferent platform 2 by users
12, . . . , 2 . We aim at identifying that the two posts refer to the
same content ( = ′ ) and at matching the subset of users that
are sharing the article across the two networks (e.g., 31 = 2).
6
Metadata from text. Existing NLP tools should be extended and
integrated to characterize the nature of the interactions across
actors w.r.t. the specific content. This can lead to labelling the
edges in the graph with information (metadata) about the
interaction between the nodes (actors). Possible metadata for such edges
include the nature of the connection between two users, if it is
based on friendship or topic afinity [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], if an node is likely to be
a bot [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], if the content has been manipulated by inserting false
claims or bias in the language [
        <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
        ]. In this line of research, it
seems promising to exploit both linguistic analysis of the text and
external knowledge. The latter could be modelled as reference
information in relational datasets [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], knowledge-graphs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
or check corpora [
        <xref ref-type="bibr" rid="ref14 ref8">8, 14</xref>
        ]. Recent results show that
transformerbased language models and query generation techniques can
automatically detect text containing false claims3 and therefore
provide valuable metadata to enrich the network.
      </p>
      <p>Example 2. Consider again the article example from Example 1.
When it is shared across users, some of them introduce incorrect
statistics about its impact (“it works only for young people”), or
facts that are not supported by any source (“it will cost 100$ per
dose”). We aim to enrich the network edges by recognizing how
the content goes from its form  to a new form ∗ when it is
shared by a certain user .</p>
      <p>
        We believe that an efective solution to the problem of
creating difusion networks for textual content across heterogeneous
sources would enable better disinformation campaigns detection.
The resulting graph with typed nodes (persons, organizations)
and typed relationships (copy or manipulation in terms of
content or form) can be then analyzed with existing methods such
as clustering and geometric deep learning [
        <xref ref-type="bibr" rid="ref12 ref9">9, 12</xref>
        ], or with novel
methods that take full advantage of the new information and
better identify emerging coordinated campaigns.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <fpage>2020</fpage>
          .
          <article-title>Cross-platform disinformation campaigns: lessons learned and next steps. The Harvard Kennedy School (HKS) Misinformation Review (</article-title>
          <year>2020</year>
          ). https: //doi.org/10.37016/mr-2020
          <source>-002</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Naser</given-names>
            <surname>Ahmadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Joohyung</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Papotti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>Saeed</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Explainable Fact Checking with Probabilistic Answer Set Programming</article-title>
          .
          <source>In TTO.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Barbieri</surname>
          </string-name>
          , Francesco Bonchi, and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Manco</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Who to follow and why: link prediction with explanations</article-title>
          .
          <source>In SIGKDD. ACM</source>
          ,
          <volume>1266</volume>
          -
          <fpage>1275</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Riccardo</given-names>
            <surname>Cappuzzo</surname>
          </string-name>
          , Paolo Papotti, and
          <string-name>
            <given-names>Saravanan</given-names>
            <surname>Thirumuruganathan</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks</article-title>
          .
          <source>In SIGMOD. ACM</source>
          ,
          <volume>1335</volume>
          -
          <fpage>1349</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Minje</given-names>
            <surname>Choi</surname>
          </string-name>
          , Luca Maria Aiello, Krisztián Zsolt Varga, and
          <string-name>
            <given-names>Daniele</given-names>
            <surname>Quercia</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Ten Social Dimensions of Conversations and Relationships</article-title>
          .
          <source>In WWW. ACM / IW3C2</source>
          ,
          <fpage>1514</fpage>
          -
          <lpage>1525</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Emilio</given-names>
            <surname>Ferrara</surname>
          </string-name>
          , Onur Varol, Clayton A.
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>Filippo</given-names>
          </string-name>
          <string-name>
            <surname>Menczer</surname>
            , and
            <given-names>Alessandro</given-names>
          </string-name>
          <string-name>
            <surname>Flammini</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The rise of social bots</article-title>
          .
          <source>Commun. ACM 59</source>
          ,
          <issue>7</issue>
          (
          <year>2016</year>
          ),
          <fpage>96</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Giglietto</surname>
          </string-name>
          , Nicola Righetti, Luca Rossi, and
          <string-name>
            <given-names>Giada</given-names>
            <surname>Marino</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Coordinated Link Sharing Behavior as a Signal to Surface Sources of Problematic Information on Facebook</article-title>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Naeemul</given-names>
            <surname>Hassan</surname>
          </string-name>
          , Fatma Arslan,
          <string-name>
            <given-names>Chengkai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Tremayne</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <string-name>
            <given-names>Toward</given-names>
            <surname>Automated</surname>
          </string-name>
          Fact-Checking:
          <article-title>Detecting Check-worthy Factual Claims by ClaimBuster</article-title>
          .
          <source>In KDD.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Sameera</given-names>
            <surname>Horawalavithana</surname>
          </string-name>
          , Kin Wai Ng, and
          <string-name>
            <given-names>Adriana</given-names>
            <surname>Iamnitchi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Twitter Is the Megaphone of Cross-platform Messaging on the White Helmets</article-title>
          . In Social, Cultural, and Behavioral Modeling.
          <fpage>235</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Georgios</surname>
            <given-names>Karagiannis</given-names>
          </string-name>
          , Mohammed Saeed, Paolo Papotti, and
          <string-name>
            <given-names>Immanuel</given-names>
            <surname>Trummer</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Scrutinizer: A Mixed-Initiative Approach to Large-Scale, DataDriven Claim Verification</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .
          <volume>13</volume>
          ,
          <issue>11</issue>
          (
          <year>2020</year>
          ),
          <fpage>2508</fpage>
          -
          <lpage>2521</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Franziska</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Keller</surname>
            , David Schoch,
            <given-names>Sebastian</given-names>
          </string-name>
          <string-name>
            <surname>Stier</surname>
            , and
            <given-names>JungHwan</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Political Astroturfing on Twitter: How to Coordinate a Disinformation Campaign</article-title>
          .
          <source>Political Communication</source>
          <volume>37</volume>
          ,
          <issue>2</issue>
          (
          <year>2020</year>
          ),
          <fpage>256</fpage>
          -
          <lpage>280</lpage>
          . https: //doi.org/10.1080/10584609.
          <year>2019</year>
          .1661888
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Federico</surname>
            <given-names>Monti</given-names>
          </string-name>
          , Fabrizio Frasca, Davide Eynard, Damon Mannion, and
          <string-name>
            <surname>Michael</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bronstein</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Fake News Detection on Social Media using Geometric Deep Learning</article-title>
          . CoRR abs/
          <year>1902</year>
          .06673 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Francesco</surname>
            <given-names>Pierri</given-names>
          </string-name>
          , Carlo Piccardi, and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Ceri</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Topology comparison of Twitter difusion networks reliably reveals disinformation news</article-title>
          .
          <source>Sci. Rep</source>
          .
          <volume>10</volume>
          (
          <year>2020</year>
          ). https://doi.org/10.1038/s41598-020-58166-5
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Shaden</surname>
            <given-names>Shaar</given-names>
          </string-name>
          , Nikolay Babulkov, Giovanni Da San Martino, and
          <string-name>
            <given-names>Preslav</given-names>
            <surname>Nakov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>That is a Known Lie: Detecting Previously Fact-Checked Claims</article-title>
          . In ACL.
          <volume>3607</volume>
          -
          <fpage>3618</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Saravanan</surname>
            <given-names>Thirumuruganathan</given-names>
          </string-name>
          , Nan Tang, Mourad Ouzzani, and
          <string-name>
            <given-names>AnHai</given-names>
            <surname>Doan</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Data Curation with Deep Learning</article-title>
          .
          <source>In EDBT</source>
          .
          <volume>277</volume>
          -
          <fpage>286</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>