<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigating Online Toxicity in Users Interactions with the Mainstream Media Channels on YouTube</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sultan Alshamrani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohammed Abuhamad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahmed Abusnaina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Mohaisen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Loyola University Chicago</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Saudi Electronic University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Central Florida</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social media has become an essential platform and source for most mainstream news channels, and many works have been dedicated to analyzing and understanding user experience and engagement with the online news on social media in general, and on YouTube in particular. In this study, we investigate the correlation of di erent toxic behaviors such as identity hate, and obscenity with di erent news topics. To do that, we collected a large-scale dataset of approximately 7.3 million comments and more than 10,000 news video captions, utilized deep learning-based techniques to construct an ensemble of classi ers tested on a manually-labeled dataset for label prediction, achieved high accuracy, uncovered a large number of toxic comments on news videos across 15 topics obtained using Latent Dirichlet Allocation (LDA) over the captions of the news videos. Our analysis shows that religion and crime-related news have the highest rate of toxic comments, while economy-related news has the lowest rate. We highlight the necessity of e ective tools to address topic-driven toxicity impacting interactions and public discourse on the platform.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        People around the globe adopt social media as an
essential part of their daily routine, not only for
socializing with each other, but also as a major source of
news. Among the di erent social media platforms,
the video-sharing platform \YouTube" has witnessed
a massive growth in contents, measured by the number
of published videos, as well as their popularity, with a
viewership of more than 2 billion monthly users [
        <xref ref-type="bibr" rid="ref22">21</xref>
        ]).
This massive growth has attracted publishers to
deliver their content through video-sharing platforms for
a fast delivery of content to viewers, and to enable the
social interaction with their viewers, which is enabled
by the comment section of videos.
      </p>
      <p>
        A major feature of video-sharing platforms such as
YouTube used for delivering news stories is the
interactive experience of the audience. However, users may
misuse such a feature by posting toxic comments or
spreading hate and racism. To improve the user
experience and facilitate positive interactions, numerous
e orts have been made to detect inappropriate
comments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Despite the e orts focused on detecting
inappropriate comments, the associations between
various types of toxicity and topics covered in news videos
from mainstream media remains an unexplored
challenge. This work provides an in-depth analysis of the
relationship of such toxic comments and the topics
presented on the news. Discovering topics in news videos
requires accessing, processing, and modeling the script
(i.e., caption) at a ne granularity, to allow the
detection of all news topics. Relying on the YouTube
categorization feature does not accurately capture the
topics of the video. For instance, YouTube has categorized
87.3% of the collected videos as news &amp; politics. To
this end, we explored and established topics using the
Latent Dirichlet Allocation (LDA) topic-modeling
approach that allowed assigning videos to speci c topics.
Our analysis shows that religion- and
violence/crimerelated news derive the highest rate of toxic comments
constituting 24.8%, and 25.9% of the total comments
posted on videos covering these topics, while
economyrelated news shows the lowest rate of toxic comments
with 17.4% of the total comments.
      </p>
      <p>Contribution. This work investigates the online
toxicity observed in the comments posted on mainstream
K
7
9K 68 054K 304K 400K
0
5</p>
      <p>K
9
6
0
K 2
0
12K K 48
6 43K 480
3
234K 816K
4
8
4
RHTuffPostCGTN FAolxJazeerMaSNBBloComberg CNNNDSTkVyNews CBC BBC ABC RETuÉronews</p>
      <p>Channel</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        With the growing popularity of online platforms in
delivering news [
        <xref ref-type="bibr" rid="ref6 ref8">8, 6</xref>
        ], the comment section of these
platforms has become an important feature where users
interact with the contents, contents providers, and each
other, to express their opinions on the published
contents. The convenience of expressing opinions through
the non-restrictive medium of online social platforms
may result in misusing such a medium by posting toxic
comments [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This has led many researchers to
investigate di erent inappropriate behaviors in the
comment section of di erent websites. The majority of the
prior research work, however, has focused on designing
classi cation or detection mechanisms for
inappropriate comments, while a few have focused on user
experience and engagement, as outlined below.
      </p>
      <p>
        Toxic Comment Classi cation. Despite various
efforts on analyzing toxic contents, identifying distinct
behaviors and patterns in this space is a challenge,
especially when (1) providing directions for prevention
and detection methods, and (2) establishing an
association with the comment/content topics. However,
there are numerous studies that explored several
aspects of toxicity, hate speech, and bias in online social
interactions [
        <xref ref-type="bibr" rid="ref1 ref17 ref19 ref4">18, 16, 4, 1</xref>
        ].
      </p>
      <p>
        User Engagement and Interactivity. Another
major area in studying user's behavior is using the
comments to identify users' engagement with the
online news and comments [
        <xref ref-type="bibr" rid="ref18 ref20 ref9">17, 9, 19</xref>
        ]. Diakopoulos et
1
5
8
9 3
676 237
1
      </p>
      <p>
        7
97 97 190
77 10 93 10
6 4
37 69 2 7
5 5 55 42 37
ABC CNN RT BBC Fox NDATlJVazeeraCGTN CBCMSNBBloCombSekrgyNews RTEÉuronewHsuffPost
Channel
al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] investigated the relationship between the quality
of the comments and both the consumption and
production of news on SacBee.com, including users
motivation for both reading and writing news comments.
Ksiazek et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposed a framework to distinguish
between users commenting on contents and those
replying to other users to better understand engagement.
In this work, and in the same space, we study the
correlation between the topic of the news and the type of
inappropriate comments, e.g., obscenity and identity
hate.
      </p>
      <p>
        Other noteworthy works that have been
conducted on behavioral modeling of YouTube content
include [
        <xref ref-type="bibr" rid="ref10 ref12 ref13">13, 10, 12</xref>
        ], although not particularly addressing
ne-grained toxicity analysis of mainstream news.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>This section describes the methods used for data
collection and representation, toxicity detection, and
topic modeling.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Collection and Measurements</title>
        <p>
          The data used in this study consists of comments
posted on news videos from YouTube, as well as the
captions of these videos. We collected more than
7.3 million comments posted on roughly 14,500 news
videos from popular 30 news channels. The collected
comments are distributed from early 2007 until
October 2019. We were able to extract video captions from
only 10,883 videos, as the remaining videos do not
include captions. Moreover, we extended our data
collection with the annotated ground truth dataset from
the Conversation AI team [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] for the purpose of
comment toxicity analysis task.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>YouTube News Channels. We collected comments</title>
        <p>
          on YouTube videos published by the most viewed
mainstream media based on Ranker [
          <xref ref-type="bibr" rid="ref15">14</xref>
          ]. We
extended our list of mainstream media channels from a
Wikipedia list of the most viewed news channels [
          <xref ref-type="bibr" rid="ref21">20</xref>
          ].
The nal list includes 30 English-speaking news
channels from 16 countries.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Data Statistics and Measurements. We collected</title>
        <p>a total of 7.3 million comments posted by 2,992,273
unique users, and published in the past 13 years (2007
to 2019) where most of the videos were published in
2019, as the trend shows an increase in news video
popularity in recent years.</p>
        <p>The popularity of the channels used in our study
can be seen in the average number of views as shown
in Figure 1 for the top-15 most-viewed channels. For
instance, videos collected from channels such as ABC,
CNN, and RT have a considerably high number of
views (i.e., with an average exceeds one million views
per video). Intuitively, as the number of views
increases, the number of comments is more likely to
increase. The average number of comments posted on
videos from the most popular mainstream media
channels on YouTube is very high as shown in Figure 2.
Here, the videos published by CNN, ABC, and Fox
news have the highest average number of comments
per video which are 6,622, 4,243, and 3,581
respectively. Generally, most of the top-15 channels maintain
an average of more than 500 comments per video.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Toxicity-related Annotated Datasets. To study</title>
        <p>
          users' behavior in the comment section, we utilized
two ground truth datasets to train a machine
learningbased ensemble classi er for toxic comment detection
and classi cation: (i) Wikipedia comments created by
Conversation AI team [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and (ii) our own
manuallyannotated YouTube comments.
        </p>
        <p>• Wikipedia Ground Truth: 160,000 comments from
Wikipedia Talk pages, manually-annotated by the
Conversation AI team, with 143,000 comments
labeled as safe, 15,294 toxic, 8,449 obscene, and
1,405 identity hate comments. The labels may
overlap, allowing the assignment of more than one
label to a toxic comment.
• YouTube Ground Truth Dataset: This is an
inhouse dataset that we created by manually
annotating 5,958 random YouTube comments, rst
into either toxic or safe. The toxic (general class)
comments are then mapped to either (i.e., toxic,
obscene, or identity hate). The nal dataset had
1,832 safe, 4,126 toxic, 2,367 obscene, and 788
identity hate comments.
3.2</p>
      </sec>
      <sec id="sec-3-5">
        <title>Data Preprocessing</title>
        <p>For proper data analysis, we initially removed all
nonEnglish contents across all datasets and eliminated
irrelevant characters, tokens, and stop-words. We also
removed frequent words appearing in more than 50%
of the captions.
3.3</p>
      </sec>
      <sec id="sec-3-6">
        <title>Data Representation</title>
      </sec>
      <sec id="sec-3-7">
        <title>Comments Data Representation. We utilized</title>
        <p>
          the pre-trained Word2Vec model from Gensim [
          <xref ref-type="bibr" rid="ref16">15</xref>
          ].
Word2Vec maps words to numerical vectors, and
words occurring in a similar context are mapped
into similar vectors. Capturing such relationships is
possible when acquiring enough data, enabling the
Word2Vec model to accurately predict the word
meaning based on past appearances from the provided
context. The comment is then represented as word
vectors of size n 300, where n is the number of words
in the comment, with an upper limit of 50 words per
comment, as most comments have less than 50 words.
Captions Data Representation. Investigating the
topic/comments associations requires de ning and
understanding the topics raised in videos where the
comments are observed. This understanding of topics can
be done using topic modeling on captions extracted
from videos. For the topic modeling task and topics
assignment to videos, we extracted and pre-processed
captions from the videos, i.e., transforming captions
to lowercase, tokenization, and eliminating irrelevant
tokens such as stopwords, punctuation, and words
containing less than three characters. After the
preprocessing phase, captions are represented using bags
of words, in which, words are assigned a unique
identier. To reduce the dimensionality of the bag-of-words,
we selected the top 10,000 words to be the caption data
representation.
3.4
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>Toxicity Detection Models</title>
        <p>The rst task of this study is to detect and classify
different toxic behaviors of comments, in order to further
investigate their association with the topics covered in
the news of which the comments are collected. We
inspected comments for three categories of toxicity:
toxic, obscene, and identity hate. We utilized a neural
network-based ensemble of three models for classifying
the three toxic categories.</p>
      </sec>
      <sec id="sec-3-9">
        <title>Deep Neural Network (DNN)-based Architec</title>
        <p>ture. DNN is a supervised learning method that can
1.0
discover both linear and non-linear relationships
between the input and the output. Comments
represented as sequences of word embeddings are fed to
the DNN-based models for labeling. The DNN model
used in this study consists of (1) an input layer of size
(50 300), similar to the shape of the embeddings of
the Word2Vec representation, (2) two fully connected
hidden layers of size 128 with ReLU activation
function, and (3) the output layer with one sigmoid.</p>
      </sec>
      <sec id="sec-3-10">
        <title>Dataset Handling and Splitting. Using the two</title>
        <p>ground truth datasets, we utilized two di erent
approaches to split the datasets for training and
evaluating the models. (1) We adopted a 50/50
splitting method for the training and testing of our models
using our YouTube ground truth comments datasets.
Since the manually-annotated comments dataset is
relatively small, the training process is initially done
using Wikipedia ground truth comments dataset. Then,
each model was ne-tuned using the 50% training
dataset of the manually-annotated YouTube
comments. (2) We also used 50/50 training/testing splits
of the Wikipedia ground truth comments dataset for
exploring the e ects of di erent experimental settings.
We note that comments can be categorized into
multiple toxic categories, e.g. one comment can be toxic,
obscene, and implies identity hate. Therefore,
comments that imply multiple toxic behaviors can be used
for training and evaluating multiple models.
3.5</p>
      </sec>
      <sec id="sec-3-11">
        <title>Topic Modeling using LDA</title>
        <p>Topic modeling is an unsupervised statistical machine
learning technique that processes a set of documents
and detects word and phrase patterns across
documents to cluster them based on their similarities.
Fine-grained Topics Extraction. We studied the
associations between a speci c toxic behavior (e.g.
obscenity) and an extracted topic from videos of
mainstream media channels. To do so, we conducted a topic
modeling to assign topics to videos based on their
caption. This is a challenging task since YouTube
categorization is generic and lacks speci cation of topics
covered in the video script. We observed that most
videos (87.3%) published by the news channels are
categorized as News &amp; Politics. Based on our analysis
of topics appeared in news videos, a variety of
topics were captured including war/attack/refugees,
violence/crime, sports/games, politics, economy.
LDA Model Settings and Evaluation. The LDA
operates using the bag of words representation of
caption segments. The topic model receives input vectors
of 10,000 bag-of-word representation and assigns
topics for each segment. This process includes a training
phase that requires setting several parameters such as
the number of topics, alpha (the segment-topic
density), and beta (topic-word density). To examine the
e ect of di erent parameters on the modeling task, we
conducted a grid search mechanism to obtain the best
con guration of the LDA model that allows for the
highest coherence score possible. For the number of
topics, we explored the e ects of changing the number
of targeted topics from 10 to 40 with an increase of 5
topics each iteration. For tuning alpha and beta
parameters, we vary the values from 0.01 to 1 with an
increment of 0.3 at each step. The LDA-model achieves
the best performance using the following settings:
numberof topics 20; alpha 0:61; beta 0:31 with
a coherence score of 0.55.</p>
        <p>We manually inspected the frequent keywords of the
best-performing LDA output and assigned names and
descriptions to them, resulting in various
consolidations, and producing 15 distinct topics.
0.2
AfriUcaSnFAofrfeaiigrCsnliPmoalitceCy/EoCnneoflriucgrtyt//PLreogteasltSysteEmconoEmEdyuurcoFapateiomannilyU/TFnioilomFnp/oEiocvde/nDtiset/FarmRSepliogritosn/GamUeSVsPioWaleratnirec/Aset/tCarcikm/Reefuges
1 Toxic Comments: Figure 3(a) shows the
performance of the toxic-behavior detection model
in terms of TPR and TNR using di erent
classi cation probability thresholds. We selected the
threshold of 0:520 as the best TPR/TNR trade-o
with a TPR of 86.2% and a TNR of 71.2%. This
model shows that 22.4% of the comments are
classi ed as toxic with a total of 1,648,345 comments.
2 Obscene Comments: The model with a
decision threshold of 0:27 achieves a high TPR of
86.6% and TNR of 88.8% for detecting obscene
comments. Figure 3(b) shows the results of
adopting di erent thresholds. Applying the model
allows the classi cation of 7.43% of the comments
as obscene with a total of 547,222 comments.
3 Identity Hate Comments: Figure 3(c) shows
the outstanding performance of the specialized
model for detecting identity hate. Using a
decision threshold of 0:140, the model achieves a TPR
of 74.8% and a TNR of 98.4%. The model shows
that 7.03% of the comments are classi ed as
identity hate with a total of 518,213 comments.
4.2</p>
      </sec>
      <sec id="sec-3-12">
        <title>Toxicity and Topics Associations</title>
        <p>The detection of toxic behaviors and access to the
topic categorization of videos allow us to conduct
toxicity/topic analyses. Such associations show whether
speci c toxicity is topic-driven or derived by other
factors. Based on our topic model and ensemble classi er,
we examined the presence of toxic, obscene and
identity hate comments on each topic of our LDA model.
1 Toxic Comments: Figure 6 shows that the
videos discussing topics related to religions or
violence/crime have the highest rate of toxic
comments, with roughly 25% of the comments are
toxic. On the other hand, economy-related news
shows the lowest rate of toxic comments with 17%
of the total number of comments.
2 Obscene Comments: The
violence/crimerelated news had the highest number of obscene
comments; 10% of the total comments. News
covering the United States foreign policy had the
least number of obscene comments, with only 3%,
as shown in Figure 4.
3 Identity Hate Comments: Among the 15
topics, African a airs and religion news had the
highest ratio of identity hate comments; 20% of the
comments. While news related to climate/energy
and the United States foreign policy have the least
number of identity hate comments with about 4%
of total comments as shown in Figure 3(c).
Content-related Toxicity. We note that toxic
comments can be posted due to several factors and may
not be totally driven by the covered topics. In an
attempt to relate speci c toxic comments with the topics
content, we conducted a statistical analysis to measure
the commonalities between comments and the content
of the caption. For videos of each topic, we obtained
the average number of common terms and expressions
to be the baseline of indicating the relationship
between the topic and the toxic comment. We note that
this might not always hold. However, we observed that
comments containing a number of common terms with
the caption that is higher than the average of common
terms in a target topic are more likely to be related
to the topics covered in the caption. This analysis
produced similar ratios of di erent toxic behaviors in
di erent news topics.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We designed and evaluated an ensemble of models to
detect various types of toxicity in comments posted
on YouTube mainstream media channels. By
analyzing 7 million YouTube comments, posted on 14,506
YouTube news videos, we detected and classi ed toxic
comments with high accuracy, and demonstrated that
despite countless e orts in comment moderation taken
by YouTube, 69% of the collected videos contained
toxic comments. We investigated the correlation
between the content of news videos and di erent toxic
behaviors across 15 topics, showing that religion and
violence/crime-related news have the highest rate of
toxic comments, while economy-related news have the
lowest rate of toxic comments. While interesting in
its own right from a behavioral standpoint, this study
highlights the need for more e ective moderation.
Acknowledgement. Work was done while all
authors were at the University of Central Florida, and is
supported by NRF grant 2016K1A1A2912757 (Global
Research Lab). S. Alshamrani was supported by a
scholarship from the Saudi Arabian Cultural Mission.</p>
      <p>Ac</p>
      <p>Accessed:</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Brassard-Gourdeau</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Khoury</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Impact of sentiment detection to recognize toxic and subversive online comments</article-title>
          . CoRR abs/
          <year>1812</year>
          .01704 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] ConversationAI. https://conversationai.github.io/,
          <year>2019</year>
          . cessed:
          <fpage>2019</fpage>
          -10-03.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Diakopoulos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Towards quality discourse in online news comments</article-title>
          .
          <source>In Proc. of the ACM Conference on Computer Supported Cooperative Work</source>
          ,
          <string-name>
            <surname>CSCW</surname>
          </string-name>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D</given-names>
            <surname>'Sa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            ,
            <surname>Illina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            , and
            <surname>Fohr</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Towards non-toxic landscapes: Automatic toxic comment detection using DNN</article-title>
          . CoRR abs/
          <year>1911</year>
          .08395 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Ernst</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmitt</surname>
            ,
            <given-names>J. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rieger</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beier</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vorderer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bente</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Roth</surname>
          </string-name>
          , H.-J.
          <article-title>Hate beneath the counter speech? a qualitative content analysis of user comments on youtube related to counter speech videos</article-title>
          .
          <source>Journal for Deradicalization</source>
          ,
          <volume>10</volume>
          (
          <year>2017</year>
          ),
          <volume>1</volume>
          {
          <fpage>49</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>GEIGER</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Key ndings about the online news landscape in america</article-title>
          .
          <source>tinyurl.com/y44m63xu</source>
          ,
          <year>2019</year>
          . Accessed:
          <fpage>2020</fpage>
          -16-04.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ksiazek</surname>
            ,
            <given-names>T. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lessard</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>User engagement with online news: Conceptualizing interactivity and exploring the relationship between online news videos and user comments</article-title>
          .
          <source>New media &amp; society 18</source>
          ,
          <issue>3</issue>
          (
          <year>2016</year>
          ),
          <volume>502</volume>
          {
          <fpage>520</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Locklear</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>More people get their news from social media than newspapers</article-title>
          . https://tinyurl.com/y8ht3ubr,
          <year>2018</year>
          . Accessed:
          <fpage>2020</fpage>
          -16-04.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Cong</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>Topic-driven reader comments summarization</article-title>
          .
          <source>In Proc. of 21st ACM International Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mariconti</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suarez-Tangil</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blackburn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cristofaro</surname>
            ,
            <given-names>E. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kourtellis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leontiadis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stringhini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>"you know what to do": Proactive detection of youtube videos targeted by coordinated hate attacks</article-title>
          .
          <source>Proc. ACM Hum. Comput. Interact. 3</source>
          ,
          <string-name>
            <surname>CSCW</surname>
          </string-name>
          (
          <year>2019</year>
          ),
          <volume>207</volume>
          :1{
          <fpage>207</fpage>
          :
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Massaro</surname>
            ,
            <given-names>T. M.</given-names>
          </string-name>
          <article-title>Equality and freedom of expression: The hate speech dilemma</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Papadamou</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papasavva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zannettou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blackburn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kourtellis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leontiadis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stringhini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sirivianos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Disturbed youtube for kids: Characterizing and detecting disturbing content on youtube</article-title>
          . arXiv:
          <year>1901</year>
          .
          <volume>07046</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Papadamou</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zannettou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blackburn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cristofaro</surname>
            ,
            <given-names>E. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stringhini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sirivianos</surname>
            ,
            <given-names>M. Understanding</given-names>
          </string-name>
          <article-title>the incel community on youtube</article-title>
          . CoRR abs/
          <year>2001</year>
          .08293 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          www.ranker.com,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Ranker</surname>
          </string-name>
          .
          <fpage>2019</fpage>
          -
          <volume>09</volume>
          -09.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Rehurek</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sojka</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          .
          <source>In Proc. of the Workshop on New Challenges for NLP Frameworks</source>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Shtovba</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shtovba</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Petrychko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Detection of social network toxic comments with usage of syntactic dependencies in the sentences</article-title>
          .
          <source>In Proc. of the 2nd International Workshop on Computer Modeling and Intelligent Systems</source>
          , CMIS (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Sil</surname>
            ,
            <given-names>D. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sengamedu</surname>
            ,
            <given-names>S. H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bhattacharyya</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Supervised matching of comments with news article segments</article-title>
          .
          <source>In Proc. of the 20th ACM Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mondal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Correa</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benevenuto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Analyzing the targets of hate in online social media</article-title>
          .
          <source>In Proc. of the 10th International Conference on Web and Social Media</source>
          ,
          <string-name>
            <surname>ICWSM</surname>
          </string-name>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tsagkias</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weerkamp</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and de Rijke,
          <string-name>
            <surname>M. Predicting</surname>
          </string-name>
          <article-title>the volume of comments on online news stories</article-title>
          .
          <source>In Proc. of 18th ACM Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM</source>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20] wikipedia. https://tinyurl.com/y5oyytc8,
          <year>2019</year>
          . Accessed:
          <fpage>2019</fpage>
          -09-09.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21] YouTube. https://tinyurl.com/y9nmv95q,
          <year>2020</year>
          . Accessed:
          <fpage>2020</fpage>
          -04-29.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>