<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Narrative Trends of COVID-19 Misinformation</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Nitin Agarwal</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Thomas Marcoux</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Arkansas at Little Rock</institution>
          ,
          <addr-line>Little Rock AR 72204</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>The COVID-19 crisis has seen the rise of many harmful online narratives. These narratives have seeped into the real world and pose tangible health risks. For this reason, we have leveraged existing techniques and developed tools to further our understanding of online misinformation and its dynamics. This provides policy makers with more tools to sift through otherwise impossibly large data sets. To ensure this, we worked closely with the Arkansas Attorney General.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>LDA topic model due to its widespread use and proved performances [BNJ03]. Following recent research claiming
that the use of custom stop-words adds little benefits [SMM17], we followed the researchers’ recommendation and
removed common words after the model had been trained. Our model choice has seen use in previous research
using LDA for short texts, specifically for short social media texts such as tweets [AAAH+20, CMVM20, ZML17]
or hashtags [ARL17] - to provide further context to topic models. We also tested our methodology on a secondary
data set using a hierarchical Dirichlet process model (HDP) [TJBB06].
3</p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>For our first set, we used data from a manually curated corpus of 243 unique misinformation narratives spanning
from January 2020 to June 2020, and gathered from various news aggregators. Topic modeling revealed
various latent narratives. For instance, some topics captured the words “purposeful”, “creators”, “bill gates”, etc.,
leading us to associate that topic with conspiracy theories suggesting the virus is man-made. As speculations
on the origin of the virus dwindled, we saw a rise in various attempts at taking advantage of vulnerable citizens
through scams or online identity theft. That topic included the words “scam”, “phishing”, “giveaways”. One
other dominant topic that was revealed is included the words “government”, “control”, “citizens”, and
“predicted”. This highlighted narratives suggesting the virus stems from a government e↵ort. The main takeaway is
that misinformation items attempting to spread fear about a potential COVID-19 vaccine and phishing scams
remained prominent during June. During the month of July, the main themes of the misinformation items shifted
back to attempts to downplay the deadliness of the novel coronavirus. Another prominent theme in July were
attempts to convince the public that COVID-19 testing is inflating the results.</p>
      <p>Although a variety of misinformation themes were identified, particular dominant themes stood out. These
themes were considered as dominant based on a simple sum of their frequency of occurrence in our data set.
During the month of March, the prominent misinformation theme was the promotion of remedies and techniques
to supposedly prevent, treat, or kill the novel coronavirus. During the month of April, the prominent themes still
included the promotion of remedies and techniques, but additional prominent themes began to stand out. For
example, several misinformation stories attempted to downplay the deadliness of the novel coronavirus. Others
discussed the anti-malaria drug hydroxychloroquine. Others promoted the idea that the virus was a hoax meant
to defeat President Donald Trump. Others consisted of various attempts to attribute false claims to high-profile
people, such as politicians and representatives of health organizations. Also in April, although first signs of
these were seen in March, the idea that 5G caused the novel coronavirus began to become more prevalent.
During the month of May, the prominent themes shifted to predominantly false claims made by high-profile
people, followed by attempts to convince citizens that face masks are either more harmful than not wearing
one, or are ine↵ective at preventing COVID-19, and how to avoid rules that required their use. The number
and variety of identity theft phishing scams also increased during May. Misinformation items attempting to
attribute false claims to high-profile people continued throughout May. Also becoming prominent in May were
misinformation items attempting to spread fear about a potential COVID-19 vaccine, and items promoting the
use of hydroxychloroquine. During the month of June, the prominent theme shifted significantly to attempts
to convince citizens that face masks are either more harmful than not wearing one, and how to avoid rules
that required their use. Phishing scams also remained prominent during June. During the month of July, the
dominant themes of the misinformation items shifted back to attempts to downplay the deadliness of the novel
coronavirus. Another prominent theme in July were attempts to convince the public that COVID-19 testing is
inflating the results.</p>
      <p>For our second data set; after seeing promising results while using small samples of data, we experimented
with YouTube data: the video titles and associated comments of videos coming up when searching YouTube
for relevant keywords: “covid”, “outbreak”, “virus”, etc. Since the data was not as finely curated for YouTube
videos, the resulting topics were unsurprisingly not as detailed. However we obtained some promising results,
especially when studying YouTube comments. With a much larger set (652,120 comments), we found that a
HDP model outperform LDA models insofar as it was able to identify a probable topic for misinformation.
When applied to our comments set, our LDA model mostly found general terms while also successfully isolating
non-English comments. Figure 1 shows two topics of interests detected by our LDA model. One is Topic “7”,
characterized by the keywords “china”, “virus”, and “made”. While discussion of China has been on a downward
trend since the start of the pandemic, the mention of the term “virus” along with “china” suggests toxic behavior.
The other Topic, “17”, includes toxic language and keywords that could be used in a hostile way or communicate
further sinophobic sentiments - for example, “dumb”, and “’bats”.</p>
      <p>Our LDA model, however, behaved as expected and was able to identify major topics, mostly news videos,
as well as what we suspect to be a vehicle of misinformation. On this very large set, our HDP model somewhat
outperformed LDA for our purposes as it was able to identify a probable topic for misinformation. When applied
to our comments set, our LDA model mostly found general terms while also successfully isolating non-English
comments. The model did identify a topic with some toxic language and some that could be used in a hostile
way or communicate sinophobic sentiments. While discussion of China has so far been on a downward trend
since the start of the pandemic, the mention of the term “virus” along with “china” suggests toxic behavior.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>This study has highlighted some of the narratives that surfaced during the COVID-19 pandemic. From January
2020 to July 2020, we collected 243 unique misinformation narratives and proposed a tool to observe their
evolution. We have shown the potential of using topic modeling visualization to get a bird’s eye view of the
fluctuating narratives and an ability to quickly gain a better understanding of the evolution of individual stories.
We have seen that the tool is ecient to chronologically represent actual narratives pushed to various outlets,
as confirmed by the ground truth observed by our misinformation curating team and independent international
organizations. Working with the Arkansas Oce of the Attorney General, this study illustrates a relatively
quick technique for allowing policy makers to monitor and assess the di↵usion of misinformation on online social
networks in real-time, which will enable them to take a proactive approach in crafting important theme-based
communication campaigns to their respective citizen constituents. We have made most of our findings available
online to support this e↵ort. We then scaled up our data and repeated our methodology on social media data:
YouTube video titles and comments. Over concerns of our LDA topic model becoming dicult to scale, we
experimented with a HDP (Hierarchical Dirichlet Process) model, which attempts to infer the number of topics.
We found promising but unsurprisingly less precise results. We notice that HDP was able to isolate a probable
subset of polarizing comments. One possible way forward would be to use HDP to identify these subsets, filter
out irrelevant comments, then apply the LDA model. This may reveal various narratives, some of them spreading
misinformation, and further automate the process of identifying online misinformation in uncontrolled spaces.
4.0.1</p>
      <p>Acknowledgements
This research is funded in part by the U.S. National Science Foundation (OIA-1946391, OIA-1920920,
IIS1636933, ACI-1429160, and IIS-1110868), U.S. Oce of Naval Research (N00014-10-1-0091, N00014-14-1-0489,
N00014-15-P-1187, N00014-16-1-2016, N00014-16-1-2412, N00014-17-1-2675, N00014-17-1-2605,
N68335-19-C0359, N00014-19-1-2336, N68335-20-C-0540, N00014-21-1-2121), U.S. Air Force Research Lab, U.S. Army
Research Oce (W911NF-17-S-0002, W911NF-16-1-0189), U.S. Defense Advanced Research Projects Agency
(W31P4Q-17-C-0059), Arkansas Research Alliance, the Jerry L. Maulden/Entergy Endowment at the University
of Arkansas at Little Rock, and the Australian Department of Defense Strategic Policy Grants Program (SPGP)
(award number: 2020-106-094). Any opinions, findings, and conclusions or recommendations expressed in this
material are those of the authors and do not necessarily reflect the views of the funding organizations. The
researchers gratefully acknowledge the support.
[TJBB06]
[ZML17]</p>
      <p>Qian Liu, Zequan Zheng, Jiabin Zheng, Qiuyi Chen, Guan Liu, Sihan Chen, Bojia Chu, Hongyu Zhu,
Babatunde Akinwunmi, Jian Huang, Casper J P Zhang, and Wai-Kit Ming. Health communication
through news media during the early stage of the covid-19 outbreak in china: Digital topic modeling
approach. J Med Internet Res, 22(4):e19118, 4 2020.</p>
      <p>Thomas Marcoux, Esther Mead, and Nitin Agarwal. The ebb and flow of the covid-19
misinformation themes. 5th International Workshop on Mining Actionable Insights from Social Networks
Special Edition on Dis/Misinformation Mining from Social Media (MAISON 2020), 10 2020.
Gordon Pennycook, Jonathon McPhetres, Yunhao Zhang, Jackson G. Lu, and David G.
Rand. Fighting COVID-19 Misinformation on Social Media: Experimental Evidence for a
Scalable Accuracy-Nudge Intervention. Psychological Science, 31(7):770–780, 2020. eprint:
https://doi.org/10.1177/0956797620939054.</p>
      <p>Hao Sha, Mohammad Al Hasan, George Mohler, and P. Je↵rey Brantingham. Dynamic topic
modeling of the covid-19 twitter narrative among u.s. governors and cabinet executives, 2020.
A. Schofield, M. Magnusson, and D. Mimno. Pulling out the stops: Rethinking stopword removal
for topic models. In 15th Conference of the European Chapter of the Association for Computational
Linguistics, volume 2, page 432–436. Association for Computational Linguistics, 2017.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [AAAH+20]
          <string-name>
            <given-names>Alaa</given-names>
            <surname>Abd-Alrazaq</surname>
          </string-name>
          , Dari Alhuwail, Mowafa Househ, Mounir Hamdi, and
          <string-name>
            <given-names>Zubair</given-names>
            <surname>Shah</surname>
          </string-name>
          .
          <source>Top Concerns of Tweeters During the COVID-19 Pandemic: Infoveillance Study. J Med Internet Res</source>
          ,
          <volume>22</volume>
          (
          <issue>4</issue>
          ):e19016,
          <year>April 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[ARL17] [BLS09] [BNJ03] Md. Hijbul Alam</source>
          ,
          <string-name>
            <surname>Woo-Jong Ryu</surname>
            , and
            <given-names>SangKeun</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Hashtag-based topic evolution in social media</article-title>
          .
          <source>World Wide Web</source>
          ,
          <volume>20</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1527</fpage>
          -
          <lpage>1549</lpage>
          ,
          <year>November 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>David M.</given-names>
            <surname>Blei</surname>
          </string-name>
          , John D. La↵erty, and
          <string-name>
            <surname>Ashok</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Srivastava</surname>
          </string-name>
          . Text Mining: Classification, Clustering, and Applications. CRC Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>David M.</given-names>
            <surname>Blei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Andrew Y.</given-names>
            <surname>Ng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          :
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [CMVM20]
          <string-name>
            <given-names>Ranganathan</given-names>
            <surname>Chandrasekaran</surname>
          </string-name>
          , Vikalp Mehta, Tejali Valkunde, and
          <string-name>
            <given-names>Evangelos</given-names>
            <surname>Moustakas</surname>
          </string-name>
          .
          <article-title>Topics, trends, and sentiments of tweets about the covid-19 pandemic: Temporal infoveillance study</article-title>
          .
          <source>J Med Internet Res</source>
          ,
          <volume>22</volume>
          (
          <issue>10</issue>
          ):e22624,
          <year>Oct 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [KAJK+20]
          <string-name>
            <surname>Ramez</surname>
            <given-names>Kouzy</given-names>
          </string-name>
          , Joseph Abi Jaoude, Afif Kraitem,
          <string-name>
            <surname>Molly B El Alam</surname>
            , Basil Karam, Elio Adib, Jabra Zarka, Cindy Traboulsi, Elie W Akl, and
            <given-names>Khalil</given-names>
          </string-name>
          <string-name>
            <surname>Baddour</surname>
          </string-name>
          .
          <source>Coronavirus Goes Viral: Quantifying the COVID-19 Misinformation Epidemic on Twitter. Cureus</source>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <fpage>e7255</fpage>
          -
          <lpage>e7255</lpage>
          ,
          <year>March 2020</year>
          . Publisher: Cureus.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[LZZ+20] [MMA20] [PMZ+20] [SHMB20] [SMM17] Yee Whye Teh</source>
          ,
          <string-name>
            <surname>Michael I. Jordan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthew J.</given-names>
            <surname>Beal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>David M.</given-names>
            <surname>Blei</surname>
          </string-name>
          .
          <article-title>Hierarchical Dirichlet Processes</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          ,
          <volume>101</volume>
          (
          <issue>476</issue>
          ):
          <fpage>1566</fpage>
          -
          <lpage>1581</lpage>
          ,
          <year>2006</year>
          . Publisher: Taylor &amp; Francis.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W. Mao, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>Modeling topic evolution in social media short texts</article-title>
          .
          <source>In 2017 IEEE International Conference on Big Knowledge (ICBK)</source>
          , pages
          <fpage>315</fpage>
          -
          <lpage>319</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>