<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>New Frontiers for Fake News Research on Social Media in 2021 and Beyond? Invited Talk - Extended Abstract</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Ben-Gurion University</institution>
          ,
          <addr-line>Beer-Sheva 8410501</addr-line>
          ,
          <country country="IL">Israel</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the past ve years, the research community made impressive strides in quantifying the dissemination and reach of fake news, understanding the cognitive mechanisms underlying belief in falsehoods and failure to correct it, and developing new methods for limiting its spread. Yet, there are many open challenges that must be addressed to ensure the integrity of the democratic process and the health of our online information ecosystem. In this talk, I focus on areas of research that are critical for advancing our understanding of fake news on social media: going beyond representative samples and convenience samples, detecting emerging fake news sources and developing new kinds of benchmark datasets for its detection, studying cross-platform impacts of platform interventions and delivering more ecologically valid experiments on social media. Progress on these fronts is necessary in order to study the pockets of society that are most heavily hit by fake news, to limit its impact, and to devise mitigations.</p>
      </abstract>
      <kwd-group>
        <kwd>Fake news Social media Research agenda</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Key takeaways from ve years of research</title>
      <p>
        Strictly focusing on the experience of voters on social media reveals four themes
in the academic literature. First, it is evident that a lot of fake news, at least in
the 2016 U.S. presidential election circulated on social media. Groundbreaking
reporting from Buzzfeed's Craig Silverman revealed that the top 20 fake news
stories outperformed the top 20 real news stories in terms of the number of
engagements on Facebook in the months leading up to the elections [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Consist
with that, Guess et al. show that a considerable amount of visits to fake news
sites by voters originate from Facebook [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and our work estimated that about
5% of political content owing to voters on an average day before the election
on Twitter came from sources of fake news [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Somewhat in contrast to the rst theme, a second theme in the literature
nds that the average voter saw and shared very little fake news content in
2016. Allcott and Gentzkow conclude that \the average US adult might have
seen perhaps one or several news stories in the months before the election" [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Guess et al. report that over 90% of voters in their sample linked to Facebook
data did not share any links to fake news sources during the survey period [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Our work estimated that the average U.S. voter on Twitter had in their Timeline
only slightly more than 1% of political content coming from fake news sources [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        A third theme \settles" the con ict of the two themes: The ample amounts
of fake news on social media are concentrated in a small part of the population.
Our work nds that exposure to and sharing of content from fake news sources is
extremely concentrated { only 1% of voters on Twitter accounted for 80% and a
mere 0.1% of voters shared 80% of it [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Others have noted this concentration on
Facebook [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and in online browsing to fake news sources [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], though at slightly
less extreme levels. Moreover, consumption and sharing vary considerably based
on individuals' political orientation, interest in and engagement with politics,
and age [
        <xref ref-type="bibr" rid="ref1 ref5 ref6 ref7">1, 5, 6, 7</xref>
        ].
      </p>
      <p>
        Finally, we now know considerably more about the psychological mechanisms
behind belief in misinformation and the e ectiveness of various interventions. In
a recent review, Pennycook and Rand synthesize this burgeoning literature and
highlight, for example, how contrary to common belief a lack of critical thinking
is more dominant in falling for fake news than motivated reasoning [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. There is
also a growing body of work that describes how interventions such as attaching
warning labels to content a ect people's perception of veracity and intentions
to share it further on social media [
        <xref ref-type="bibr" rid="ref11 ref19 ref4">4, 11, 19</xref>
        ]. A comprehensive review of this
nascent area of research is outside the scope of this talk. However, it is important
to note that much of this literature is based on surveys and lab experiments
that decontextualize and decouple information from the social context they are
naturally experienced on social media, it relies on a sample of the population that
may or may not engage with fake news under normal circumstances, and that
depends, to some degree, on self-reported measures rather than actual behavior.
      </p>
    </sec>
    <sec id="sec-2">
      <title>New frontiers for fake news research</title>
      <p>I focus on three research areas that can signi cantly move the entire eld
forward. In particular, I discuss challenges and opportunities in sampling, fake news
detection, and experimentation.</p>
      <p>
        Sampling: Given the knowledge we now have about the concentration of fake
news in the population and the e orts to manipulate public opinion on social
media, we need new sampling methodologies to address these issues. The
rarity of engagement with fake news, of less than one in a thousand participants,
makes representative surveys underpowered when it comes to studying this
phenomenon, even when large surveys of thousands of people are involved. Moreover,
combined with the fact that fake news sources tend to share their audience with
other fake news sources [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the inability to generate a meaningful-size sample of
people consuming fake news leads to a potential blind spot { missing the
communities that engage most heavily with fake news. Of course, sampling social
media users based on observed behavior online can alleviate these concerns, but
it brings along issues of representativeness, validity, and reliability. In
particularly, this is problematic because we know that a considerable amount of political
activity on social media is generated by bots, trolls, foreign actors, and other
entities [
        <xref ref-type="bibr" rid="ref15 ref9">9, 15</xref>
        ], who are not eligible to vote.
      </p>
      <p>
        To overcome these challenges, the research community needs to adopt new
sampling techniques. Salganik o ers two research designs that are particularly
appropriate for addressing these issues [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The rst is Ampli ed Asking, which
refers to the use of big data to extrapolate and predict survey responses for
individuals who did not take the survey. This approach is particularly useful
for extending existing survey responses to the rest of the online platform. The
second approach proposed by Salganik is called Enriched Asking and refers to
the combination of large survey or administrative data with big observational
data through the use of record linkage. By linking the responses of a set of
individuals with their online behavior one can attain large samples while keeping
the contamination from bots, trolls, and other accounts relatively low. Building
such resources is a costly e ort, but once built it can be re-used multiple times
and result in many new insights about the online behavior of diverse populations.
Detection: Considerable amount of research has focused on the development
of machine learning models and algorithms for the automatic detection of fake
news (see Shu et al. for a comprehensive review [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). While fully automated
detection may be the ultimate goal, it is still far from substituting human
factcheckers. Most existing datasets for training machine learning models focus on
veracity directly, have varying levels of granularity and de nitions of fake news,
and are limited in size [
        <xref ref-type="bibr" rid="ref10 ref2">2, 10</xref>
        ]. Moreover, there is little to guarantee that a model
trained on the past falsehoods will reliably detect the lies of tomorrow. These are
fundamental issues that machine learning models might solve one day, but they
require considerably more training data and domain knowledge than current
models possess.
      </p>
      <p>
        To make steady progress, we need to work more closely with fact-checkers
(rather than attempting to substitute them) and build computational models
that support smaller tasks in the process of the fact-checking. For example, the
research community has largely ignored the important, time-consuming, and
non-trivial task of identifying claims that have already been fact-checked [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Benchmark datasets for this purpose are still very much absent. Another
challenge that has been largely overlooked involves the retrieval of relevant evidence
for supporting a given claim [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In addition, more research is needed to help
fact-checkers direct their e orts toward claims that are feasible to check (and not
popular) as well as identifying stable and persistent characteristics of credibility
(e.g. audience composition of a source) that are di cult to manipulate.
Interventions: Ultimately, to reduce the spread of fake news on social media
action must be taken, but the academic community cannot leave
experimentation with these actions solely at the hands of platform providers. The academic
community must be able to assess the e ectiveness of interventions
independently and free of any commercial interests. Of course, academics can try to
replicate the organic experience of social media users in lab settings, but this
comes at a cost to the ecological validity of the experiment and particularly
its social elements. One approach that has not been su ciently explored is the
instrumentation and manipulation of social media apps and web interfaces for
consenting individuals. For example, nothing stops academics from building a
Twitter clone app that interacts with organic content on the actual Twitter
platform while allowing researchers to introduce interventions to the user experience.
Consenting individuals could be asked to use such an app instead of the regular
app, and perhaps platform providers would o er such a service to the research
community.
      </p>
      <p>
        Another avenue for impactful research is the study of cross-platform e ects.
Interventions on one platform are not necessarily limited to just that platform
and can spill over to other platforms. For example, in May 2020 the social media
platform Parler gained a signi cant number of new users, allegedly due to
Twitter labeling President Trump's tweets as glorifying violence [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Therefore, it is
no longer su cient to study the impact of certain interventions in the context
of one platform, but cross-platform research is necessary to understand the full
impact of those interventions.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In recent years, fake news captured the attention of both the public and the
research community. We now know considerably more about the scale and scope
of fake news on social media and its distribution in the population. Building on
these ndings, this talk portrayed a path forward that calls for innovation in
sampling techniques, a greater focus on detection tasks that aid fact-checkers,
and methodology for evaluating intervention independently of social media
platforms.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Allcott</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gentzkow</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Social media and fake news in the 2016 election</article-title>
          .
          <source>Journal of Economic Perspectives</source>
          <volume>31</volume>
          (
          <issue>2</issue>
          ),
          <volume>211</volume>
          {
          <fpage>36</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Augenstein</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lima</surname>
            ,
            <given-names>L.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonsen</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          :
          <article-title>MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims</article-title>
          .
          <source>arXiv preprint (Oct</source>
          <year>2019</year>
          ), http://arxiv.org/abs/
          <year>1909</year>
          .03242
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Dimensions.
          <article-title>ai: Publications containing \fake news" in title or abstract in the years 2017-2021 (inclusive</article-title>
          )., https://www.dimensions.ai/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ecker</surname>
          </string-name>
          , U.K.,
          <string-name>
            <surname>O'Reilly</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>E.P.:</given-names>
          </string-name>
          <article-title>The e ectiveness of short-format refutational fact-checks</article-title>
          .
          <source>British Journal of Psychology</source>
          <volume>111</volume>
          (
          <issue>1</issue>
          ),
          <volume>36</volume>
          {
          <fpage>54</fpage>
          (
          <year>2020</year>
          ), publisher: Wiley Online Library
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Grinberg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joseph</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swire-Thompson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lazer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Fake news on Twitter during the 2016 U.S. presidential election</article-title>
          .
          <source>Science</source>
          <volume>363</volume>
          (
          <issue>6425</issue>
          ),
          <volume>374</volume>
          (Jan
          <year>2019</year>
          ). https://doi.org/10.1126/science.aau2706
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Guess</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tucker</surname>
          </string-name>
          , J.:
          <article-title>Less than you think: Prevalence and predictors of fake news dissemination on Facebook</article-title>
          .
          <source>Science advances 5(1)</source>
          (
          <year>2019</year>
          ). https://doi.org/10.1126/sciadv.aau4586, https://advances.sciencemag.org/content/5/1/eaau4586
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Guess</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyhan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rei</surname>
            <given-names>er</given-names>
          </string-name>
          , J.:
          <article-title>Selective exposure to misinformation: Evidence from the consumption of fake news during the 2016 US presidential campaign</article-title>
          .
          <source>European Research Council</source>
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <volume>4</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lewinski</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          :
          <article-title>Parler App's #Twexit Movement Recruits Twitter Users In Wake Of Trump Social Media Wars</article-title>
          . Forbes (May
          <year>2020</year>
          ), https://www.forbes.com/sites/johnscottlewinski/2020/05/29/parler-apps
          <article-title>-twexitmovement-recruits-twitter-users-in-wake-of-trump-social-media-wars/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Linvill</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boatwright</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grant</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warren</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          :
          <article-title>\THE RUSSIANS ARE HACKING MY BRAIN!" investigating Russia's internet research agency twitter tactics during the 2016 United States presidential campaign</article-title>
          .
          <source>Computers in Human Behavior</source>
          <volume>99</volume>
          ,
          <issue>292</issue>
          {
          <fpage>300</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. N rregaard, J.,
          <string-name>
            <surname>Horne</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adali</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <string-name>
            <surname>NELA-GT-2018. Harvard Dataverse</surname>
          </string-name>
          (
          <year>Jun 2019</year>
          ). https://doi.org/10.7910/DVN/ULHLCB
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pennycook</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bear</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collins</surname>
            ,
            <given-names>E.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rand</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          :
          <article-title>The implied truth e ect: Attaching warnings to a subset of fake news headlines increases perceived accuracy of headlines without warnings</article-title>
          .
          <source>Management Science</source>
          <volume>66</volume>
          (
          <issue>11</issue>
          ),
          <volume>4944</volume>
          {
          <fpage>4957</fpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pennycook</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rand</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          :
          <article-title>The psychology of fake news</article-title>
          .
          <source>Trends in cognitive sciences (</source>
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Salganik</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>Asking quesetions</article-title>
          . In:
          <article-title>Bit by bit: Social research in the digital age</article-title>
          . Princeton University Press (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Shaar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martino</surname>
            ,
            <given-names>G.D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babulkov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>That is a Known Lie: Detecting Previously Fact-Checked Claims</article-title>
          . arXiv preprint (May
          <year>2020</year>
          ), http://arxiv.org/abs/
          <year>2005</year>
          .06058
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciampaglia</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varol</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>K.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The spread of low-credibility content by social bots</article-title>
          .
          <source>Nature communications 9(1)</source>
          , 1{
          <issue>9</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Liu, H.: Disinformation, Misinformation, and Fake News in Social Media. Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Silverman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>This analysis shows how viral fake election news stories outperformed real news on Facebook</article-title>
          .
          <source>BuzzFeed News (Nov</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumgartner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korn</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Relevant document discovery for factchecking articles</article-title>
          .
          <source>In: Companion Proceedings of the The Web Conference</source>
          <year>2018</year>
          . p.
          <volume>525</volume>
          {
          <fpage>533</fpage>
          . WWW '18,
          <string-name>
            <given-names>International</given-names>
            <surname>World Wide Web Conferences Steering Committee</surname>
          </string-name>
          (
          <year>2018</year>
          ), https://doi.org/10.1145/3184558.3188723
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Yaqub</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kakhidze</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brockman</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Memon</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patil</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>E ects of credibility indicators on social media news sharing intent</article-title>
          .
          <source>In: Proceedings of the 2020 CHI conference on human factors in computing systems</source>
          . pp.
          <volume>1</volume>
          {
          <issue>14</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>