<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Participant profiling on Twitch based on chat activity and message content</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jari Lindroos</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ida Toivanen</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaakko Peltonen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tanja Välisalo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raine Koskimaa</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sami Äyrämö</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Archives of Finland</institution>
          ,
          <addr-line>Rauhankatu 17, 00170, Helsinki</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tampere University</institution>
          ,
          <addr-line>Kalevantie 4, 33100 Tampere</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Jyväskylä</institution>
          ,
          <addr-line>Seminaarinkatu 15, 40014, Jyväskylä</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <fpage>267</fpage>
      <lpage>282</lpage>
      <abstract>
        <p>Identifying participant profiles based on commenting activity in esports livestream platforms enhances our understanding of esports audiences and patterns of viewer engagement. Active chat participants represent a core audience of esports viewers due to their high level of engagement. In this study, we identify participant profiles based on chat data collected from Twitch livestreams of CS:GO Majors tournaments from 2022 and 2023. Profiling was conducted based on two types of features: chat activity and message content. We performed clustering to both sets of features to get insights about the communication patterns of chat participants and the contents of messages they sent during matches. The results show that livestreaming chat data even on a larger scale can be harnessed to understand participation in livestream chats from multiple viewpoints. Combining both of these approaches can give us a comprehensive way of analyzing and forming participant profiles based on chat participants' message behavior.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;esports</kwd>
        <kwd>Twitch</kwd>
        <kwd>clustering</kwd>
        <kwd>chat</kwd>
        <kwd>participant profiling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Chat is an essential part of esports livestreams in multiple ways. The chat in esports livestreams in
Twitch has been described as a significant part of the esports experience, “a proxy for noise, transmitting
afects compelling continued viewing and consumption“ [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Chat participation has been seen as part of
‘audience work’, which is essential for the esports economy via advertising, sponsorship, and various
other revenue channels [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In-depth analysis of esports chat discussions is crucial for understanding
audience’s behavior, preferences, and opinions [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. Analyzing livestream chat messages helps us
to gain a better understanding of what the viewers are interested in, what they dislike, and how they
engage with the content, the stream provider and.
      </p>
      <p>
        Recently, there has been a growing amount of machine learning research using massive chat data to
discuss the chat cultures in livestream platforms like Twitch [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref5 ref6 ref7 ref8 ref9">5, 6, 7, 8, 9, 10, 11, 12, 13</xref>
        ]. Manually going
through massive amounts of chat data is a resource-heavy process, which has given us an incentive
to focus on research that introduces large-scale datasets and makes primary use of machine learning
solutions in the context of Twitch chats.
      </p>
      <p>The chat data used in this study was collected during two Counter-Strike: Global Ofensive
Major Championships tournaments (CS:GO Majors) in 2022 and 2023, from two diferent broadcasters,
Pelaajat.com and Yle on Twitch.tv. Using this data, we investigate audience engagement with esports
livestreams by studying the behavior of esports spectators who participate in Twitch chat. To do
this, we construct features in two ways: 1) describing the variation of chat participants’ chat activity
over time and matches, and 2) categorizing chat messages by content types. We refer to the first as
“activity-based” and the second as “content-based”. We use machine learning methods of exploratory
9th International GamiFIN 2025 (GamiFIN 2025), April 1-4, 2025, Ylläs, Finland.
$ jari.m.m.lindroos@jyu.fi (J. Lindroos); ida.m.toivanen@jyu.fi (I. Toivanen); jaakko.peltonen@tuni.fi (J. Peltonen);
tanja.valisalo@jyu.fi (T. Välisalo); raine.koskimaa@jyu.fi (R. Koskimaa); sami.ayramo@jyu.fi (S. Äyrämö)
0000-0003-2275-3866 (J. Lindroos); 0000-0003-0440-7337 (I. Toivanen); 0000-0003-3485-8585 (J. Peltonen);
0000-0001-8678-4683 (T. Välisalo); 0000-0002-1492-4074 (R. Koskimaa); 0000-0002-7532-2771 (S. Äyrämö)
© 2025 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
data analysis, particularly clustering and dimensionality reduction, to discover and analyze distinct
participant profiles within the esports livestream audience based on their commenting behavior. We
also examine whether the same structure of participant profiles is found from both tournaments.</p>
      <p>The primary contributions of this work are:
1. We provide a method for forming participant profiles via clustering by identifying and categorizing
features of chat participant activity and the types of content in those chat messages.
2. We propose that the participants’ chat behavior can vary significantly – by forming six diferent
activity-based participant profiles and 11 diferent chat content-based participant profiles. We
then demonstrate the variety of these behaviors.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Previous research</title>
      <sec id="sec-2-1">
        <title>2.1. Audience engagement in livestreaming</title>
        <p>
          Research on audience engagement in livestreaming is largely focused on audience surveys or
questionnaires [
          <xref ref-type="bibr" rid="ref14 ref15 ref16 ref17 ref18 ref19 ref20 ref21 ref22 ref23 ref24">14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24</xref>
          ], rather than the analysis of massive the chat data
the audience produces. However, computational methods have also been employed to investigate this
phenomenon, ofering new ways for understanding the dynamics of audience engagement [
          <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
          ].
        </p>
        <p>
          Studies on chat activity itself have often focused on small datasets and qualitative methods exclusively.
Having a larger span of adjacent viewers is important in discerning which communicative characteristics
can be found from massive chat data. For example, in [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], Twitch chat datasets consisting of 50-message
segments from livestreams with up to 10,000 concurrent viewers were analyzed by hand. The study
suggested that the unique property of the type of communication found in massive chats could be
described as "crowdspeak". Conversational norms usually found in small-scale interactions dissipate
in a setting of massive chats, making repetitive and shorter speech more prominent to the degree of
the chat appearing to be chaotic and serving no communicative purpose. However, the researchers
found that in large-scale chat setting bricolage, shorthanding and voice-taking help in having more
coherence in the communication. While bricolage refers to having a small set of elements that are
re-used and re-arranged for further communicative use, shorthanding describes the deliberate choice to
ift speech into a smaller frame. Sharing viewpoints and mannerisms in online communities is referred to
as voice-taking. Conveying meaning has been found to be disrupted in massive online group settings in
a similar study [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. The researchers discovered that communication in massive chats tends to resemble
a cacophony, which is represented by repetitive and information-poor messages, as well as lower per
capita participation.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Machine learning in livestream chat analysis</title>
        <p>
          Machine learning solutions have been employed in the analysis of Twitch livestream chat data to
mitigate the limitations of time and resources of manual processing when going through massive
amounts of data. This can be seen in studies regarding automatic chat bot detection [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] or identifying
toxic language [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and spam [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ], for instance.
        </p>
        <p>
          To gain a deeper understanding of the livestream chat dynamics, machine learning has been used to
predict viewer engagement, and also its impact on the popularity of a stream. For example, shallow
artificial neural networks were used in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] to predict low or high engagement based on gameplay
events in PlayerUnknown’s Battlegrounds livestreams. Their analysis encompasses both chat logs
and game telemetry data, with specific gameplay features (e.g., player health status, in-game choices,
and map location). The results demonstrate the possibility of accurately predicting continuous viewer
engagement based exclusively on key gameplay events. Predicting the viewer count of a Twitch stream
was investigated in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The authors categorized these reactions into textual features, such as n-grams
and sentiment, and non-textual features, including chat frequency and the active time of how long
a viewer engages on the chat. They found that while textual features are important for predicting
popularity when analyzing the entire chat log, non-textual features become more crucial for early
prediction, within the first 15 minutes of a stream. More recently a novel machine learning based
approach was proposed in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for analyzing emotional responses in Spanish video game streams on
Twitch. A unique corpus of Spanish Twitch chat messages was created and manually annotated for
polarity and emotions. The results show that a BERT-based model achieved the highest accuracy in
detecting polarity (78%) and emotions (68%), outperforming other methods.
        </p>
        <p>
          Clustering has also been used to analyze the diferences between first-time visitors and regular
participants in Twitch chats [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. In the study, it was found that subscribers and regular users were the
most common participants followed by moderators, leaving the new and other participants, like bots,
to be the smallest group of participants. This may tell us that normative communication behavior in
massive chats is maintained by those who have the most experience in interacting in said environments.
Certain behaviors, such as messaging repetitively and outside typical streaming time, mentioning
irrelevant topics and other channels, and having no response in chat interactions clearly stood out
among the human-generated input and were linked to bot-generated content.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], k-means clustering was applied for word vectors to reveal more about the innate qualities
of chat data of 10,000 tokens collected from Twitch. They found that clustering word vector data is
possible but unfolds odd shapes, depending on the chosen word vector method (e.g., skip-gram with
negative sampling; see more in [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]). Features like streamer popularity were found to correlate with the
two clusters found in the study. One cluster was interpreted as having more game-specific terminology
and the other as likely consisting more of general speech and terms used throughout diferent Twitch
channels.
        </p>
        <p>
          One method frequently applied to chat data is topic modeling. For example, the Twitter-LDA topic
modeling algorithm was employed in [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] to investigate the message content of chat participants
from two Dota 2 tournament broadcasts from May 2016 and December 2016. The results suggested
that various stages of the broadcast elicit distinct types of communication behavior among viewers
depending on the content types. Active game sessions tended to include shorter and more emotional
expressive messages, whereas during breaks and inactive moments more analytical discussions and
social interactions occurred. Crowd behavior was also investigated in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to distinguish more meaningful
and coherent communication among thousands of chat participants that contributed to upholding
“crowdspeak” in chat. The authors utilized structural topic modeling and cross-correlation analysis to
examine topical and temporal patterns of chat participants’ messages during the Dota 2 International
tournament in 2017, investigating their relationship to in-game events. They showed that in-game
events significantly influence communication within large-scale chats, shaping the emergent topical
structure. During inactive periods of game activity, boredom and frustration was expressed through
emotes and recursively replicated “copypasta”. During highly anticipated phases of the in-game events
chat participants often trigger a high volume of short, emotionally charged messages and emotes (e.g.,
“ez”, “PogChamp”). Chat participants also engage in ongoing discussion regarding what is happening
on the screen – examples of this include cheering or supporting teams and players in chat messages.
This behavior occurred with less correlation to in-game event triggers.
        </p>
        <p>In summary, while prior research has applied machine learning to analyze Twitch chat from content
moderation to predicting viewer engagement, the construction of comprehensive participant profiles
integrating activity- and content-based features remains largely unexplored to our best knowledge. Our
study aims to address these gaps, ofering a more comprehensive view of engagement patterns and
participant behavior on Twitch.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>This study used chat data from Twitch.tv to explore audience engagement dynamics. Twitch is a
livestreaming platform owned by Amazon with a focus on games and related activities. Since its launch
in 2011, Twitch has significantly shaped how esports content is consumed and understood. By using
machine learning methods to perform exploratory data analysis, we aim to analyze chat participation
patterns and model distinct participant profiles within the esports livestream audience.</p>
      <p>Our dataset consists of Twitch chat messages collected during the 2022 and 2023 CS:GO Majors
tournaments, titled PGL Major Antwerp 2022 and BLAST.tv Paris Major 2023. Both tournaments
consisted of 70 individual matches played over 12 days. The data were collected from two Finnish
Twitch channels: ’yleeurheilu’ and ’pelaajatcom’. These channels served as the primary
Finnishlanguage broadcasters for the events. Yleeurheilu is operated by the publicly funded national broadcaster
Yleisradio (abbrev. Yle) and Pelaajat.com is run by the Finnish esports media company Pelaajat.com.
Chat activity across both channels produced over 107,900 messages in 2022 and 117,500 in 2023 (see Table
1). Streaming activity was also distributed due to concurrent matches: Yle was the main broadcaster for
the entirety of the 2022 Majors, including a majority of the games featuring ENCE and "Aleksib", while
Pelaajat.com provided secondary coverage. This pattern was reversed in 2023, with streaming more
evenly distributed and Yle remaining the primary broadcaster for the finals. Importantly, the presence
of Finnish organizations, such as ENCE, or Finnish players, such as "Aleksib", was a significant factor
influencing viewership patterns on both channels. This is noticed in the high viewership observed
during matches featuring these players and teams, regardless of the broadcasting channel.</p>
      <p>For clarity, we refer to our datasets with shortened terms YLE for data collected from ’yleeurheilu’ and
PCOM for data collected from ’pelaajatcom’ so that, e.g., YLE22 refers to data collected from yleeurheilu
in 2022, PCOM23 refers to pelaajatcom data collected in 2023 etc.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <p>In this study we form participant profiles by conducting two types of clustering: activity- and
contentbased clustering. These two types of clustering are conducted to explore audience engagement based
on two diferent types of distinct features: 1) features based on chat participant activity, and 2) features
based on uniqueness of the message content produced by chat participants (see summary in Table 2).
Given the difering nature of these features, we applied two diferent clustering evaluation metrics to
ensure optimal cluster separation and interpretability for each feature type.</p>
      <sec id="sec-4-1">
        <title>4.1. Activity-based clustering</title>
        <p>To analyze the activity of chat participants within the tournament, we start by manually separating
timestamps when a match occurs. This lets us distinguish between activity occurring during active
gameplay and activity in the post and pre-match period. To further understand participants’ behavior
during the tournament, we define several features to construct participant profiles based on observable
conversational participant traits. Such traits can be categorized into two distinct aspects by
messageand time related information. These activity features are the following:
• activity_count: The total number of matches the participant has taken part of (chatting during
the matches).
• comment_count_mean: The average number of messages sent by the participant across their
participated matches.
• comment_count_std: The standard deviation of the number of comments the participant made
across matches.
• comment_count_10th_quantile: The value separating the lower 10% of a participant’s matches
in terms of number of comments made from the upper 90%.
• comment_count_90th_quantile: The value separating the lower 90% of a participant’s matches
in terms of number of comments made from the upper 10%.
• timestamp_dif_to_match_start_mean : The average time elapsed in seconds between the
start of a match and the participant’s messages.
• timestamp_dif_to_match_start_std : The standard deviation of time elapsed in seconds
between the start of a match and the participant’s messages.</p>
        <p>
          Pre-examination of the preprocessed data of the activity features showed sensitivity to outliers. Due
to this, we applied the QuantileTransformer from the scikit-learn library [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] to each feature of the
activity data prior to clustering to transform the features into a uniform distribution. Previous studies
have shown it to be efective in reducing the impact of outliers while still maintaining the distribution
of the data. For instance in [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ], the authors showed that QuantileTransformer spreads common values
more evenly, minimizing outlier influence without distorting the overall feature scale. This allows
for better scaling balance compared to standard scaling techniques, which can be disproportionately
afected by outliers.
        </p>
        <p>While the transformation in the method is non-linear and it could alter correlation between features,
we still believe that its advantages in handling outliers outweigh this potential drawback in our specific
case. Since our primary goal is to identify meaningful clusters in the presence of outliers. To further
justify this choice, we conducted a comparative analysis using the scikit-learn library, evaluating the impact
of various scaling methods (including no scaling, Normalizer, MinMaxScaler, RobustScaler,
StandardScaler, MaxAbsScaler, PowerTransformer) on the clustering results. The evaluation of the clusters was
focused on interpretability, such as cluster separation and compactness, and the QuantileTransformer
yielded the most balanced and interpretable clustering results.</p>
        <p>
          We apply k-means clustering using k-means++ initialization via scikit-learn to the processed dataset
with the activity features to identify potential groups of participants exhibiting similar patterns of
chat engagement behavior. To select the optimal number of clusters for activity-based clustering, we
employ the silhouette method [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ]. Finally to visualize the clustering results, we employ t-SNE to
perform dimensionality reduction of the activity data, projecting the participants from the original
high-dimensional feature space into two dimensions [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] where the participants can be visualized as a
scatterplot and the discovered clusters can be depicted with diferent colors.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Content-based clustering</title>
        <p>
          As another framework for creating participant profiles with clustering, we use deep learning based
content detection [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ] that gives us features based on the uniqueness of message content generated by
chat participants. After describing the deep learning model used to create these features, we explain
how we are forming participant profiles based on the uniqueness features.
        </p>
        <p>Features for the second clustering approach, which is based on the participants’ chat content, are
constructed by quantifying each participant’s total messages as a distribution in terms of uniqueness.
Before any data processing, we filter out known channel bots, such as Moobot and Nightbot, to focus
the content on authentic participants. We utilize four distinct features:
• unique: messages which stand out as original, informative, and directly relevant to the contents
of the stream;
• non-unique: repetitive content such as emotes, game slang and copypasta;
• essentially non-unique but human-written: Formulaic content based on variations of certain
patterns such as reactions and cheers;
• commands and replies: commands written to a channel bot and the responses of the bot.</p>
        <p>
          For each chat participant, we calculated the share of their messages for each feature (values are in
between 0 and 1). No scaling is applied. We apply k-means clustering also using k-means++ initialization
via scikit-learn to the uniqueness features to identify potential groups of participants exhibiting similar
chat engagement patterns. To select the optimal number of clusters for content-based clustering, we
employ the elbow method [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ]. Finally to visualize the clustering results, we employ t-SNE to perform
dimensionality reduction, projecting the participants from the high-dimensional feature space of content
features into two dimensions [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ], where we then visualize participants as a scatterplot with clusters
shown as colors of the participants.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In this section we compare the results of the two diferent types of clustering, activity-based (features
based on chat participant activity) and content-based (features based on uniqueness of the message
content produced by chat participants).</p>
      <sec id="sec-5-1">
        <title>5.1. Activity-based clustering</title>
        <p>
          Looking at the number of clusters determined by the silhouette score for the activity-based features, we
get varying values across datasets: YLE22 with six clusters (silhouette score = 0.5529), YLE23 with four
clusters (silhouette score = 0.5631, six clusters score = 0.5506), PCOM22 with four clusters (silhouette
score = 0.6164, six clusters score = 0.5905), and PCOM23 with six clusters (silhouette score = 0.5910).
According to [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ], an average silhouette score of 0.5 or higher provides good evidence that the clusters
are clearly distinguishable. For the sake of interpretability and consistency, we decided to define the
number of clusters as six for all datasets (see Figure 1). While the silhouette scores indicated diferent
optimal sizes for some datasets, all scores were above 0.5 when using six clusters, providing suficient
evidence of well-separated clusters across the datasets. Defining six clusters allowed us to maintain
a coherent structure for comparing participant profiles. From these clusters, we observe a similar
underlying structure for participant profiles across diferent datasets, reinforcing the comparability of
the results.
        </p>
        <p>Altogether, we obtain six diferent participant profiles across the datasets (Table 3). These findings
indicate that participants’ chat behavior can be categorized into six diferent types. Profiles 5 and 6
describe the “come and go” type of chat participants who only leaves a single footprint during the
whole tournament either earlier (profile 5) or later (profile 6) during an ongoing match. Then there are
more active groups of participants in terms of messaging activity (profiles 1-4). Profile 4 consists of
chat participants who send exactly one message per match they participate in. Profile 3 on the other
hand consists of chat participants who leave at least one message in a match. Profiles 1 and 2 consist of
even more active chat participants who write up to 100 messages per single match. The main diference
between them is that chat participants from profile 1 always participate in at least two matches and
send at least three messages. Chat participants in profile 2 mostly participate in a single match in
the tournament, but when they participate in more than one match (four matches at maximum), they
write the same amount of comments in each match they participate in. Profiles also exhibit distinct
diferences in interaction rate which means how many of the participant’s messages are replies to or
mention other participants. With activity-based clustering, we can diferentiate between participants
who have distinct attitudes to chat participation and use diferent amounts of efort to chat engagement.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Content-based clustering</title>
        <p>We employed k-means clustering to identify distinct groups of participants based on four unique
features within each dataset. In the cluster analysis, the number of clusters was determined using the
elbow method by calculating the within-cluster-sum-of-square (WCSS) values for the uniqueness-based
features, and resulted in values ranging from 8 to 10. The datasets YLE22 and PCOM22 both have 9
clusters, PCOM23 has 10 clusters and YLE23 has 8 clusters (see Figure 2).</p>
        <p>In order to further analyze the clusters and form participant groups that span all datasets, we began
by calculating the mean values across all participants for the four unique features for each identified
cluster. We then used cosine similarity with a strong threshold of 0.95 to compare these mean feature
values between clusters across diferent years and channels, identifying pairs of clusters that exhibit
highly similar characteristics in terms of chat content (see Figure 3). Clusters with high similarity across</p>
        <p>Moderate 27%
– High
Moderate 56%
Low</p>
        <p>6.9%
Low
Low
6.4%
6.9%
3
2
1
2
1
1
datasets are considered potential candidates for forming participant profiles.</p>
        <p>Out of all the participant profiles (Table 4), there are six profiles (1-6) that are present in every dataset
(i.e., there is one cluster fitting this profile in each dataset), and two participant profiles (7-8) that are
present in three datasets (i.e., a cluster representing this profile is present in YLE22, YLE23 and PCOM22
datasets, but is missing a cluster from PCOM23 dataset). One participant profile (9) is also present in
two datasets (YLE22 and PCOM22). This leaves out some clusters from the PCOM23 dataset that were
not classified into any of the aforementioned participant profiles. The cluster pcom23_3 was classified
into participant profile 10 and fairly similarly distributed clusters pcom23_0, pcom23_6, and pcom23_8
were grouped together into 11a, 11b and 11c.</p>
        <p>Unlike other analyzed datasets, participant profiles involving PCOM23 showed a great deal of variety
among participants by their message content in “commands and replies”. Examining participant profiles
11a, 11b, 11c and participant profile 2, they all exhibit the majority of content focused on ‘commands
and replies’. Together they account for over 48% of all participants in PCOM23. Comparatively, this</p>
        <p>Profile</p>
        <p>We are here Unique and essentially
to stay non-unique messages.</p>
        <p>(43% and 45% respectively)
Very active in chat and
mentions.</p>
        <p>Communicating Primarily engage with channel
with the bots for giveaways.
system (98% of messages)</p>
        <p>Minimal social interaction.</p>
        <p>We like to gg Dominated by non-unique
and LUL messages,
primarily repetitive content
and emotes.</p>
        <p>Limited direct social
interaction.</p>
        <p>Less unique Primarily contain
members of essentially non-unique
the crowd messages.</p>
        <p>Unique Primarily contain unique
members of messages.
the crowd Largest group for PCOM22 and</p>
        <p>YLE datasets,
but only third largest in</p>
        <p>PCOM23.</p>
        <p>We keep the Contain both unique and
conversation essentially non-unique
flowing messages.</p>
        <p>Most active group in terms of
messages and interactions.</p>
        <p>Non original Contain both non-unique
discussants and essentially non-unique
messages.</p>
        <p>Contain both unique
and non-unique messages.</p>
        <p>Primarily contain commands
and replies,
with a significant portion of
unique messages.</p>
        <p>Active only in the 2022 Majors.</p>
        <p>Contain a mix of non-unique,
non-unique but human-written,
and unique messages.</p>
        <p>Only PCOM23-specific.</p>
        <p>Relies heavily on "commands
and replies,"
with varying secondary content
types (unique, non-unique,
essentially non-unique).</p>
        <p>Only PCOM23-specific groups.
We like to
LUL but also
keep it
original
Commands
are cool, but
so are chat
contents
Jack of non
giveaways
11(a,b,c) Giveaways
first, content
second</p>
        <p>Moderate
6-8</p>
        <p>4-5
pattern is less pronounced in other datasets: 22.7% of participants in PCOM22 (participant profiles 2
and 9), 10.5% of participants in YLE22 (same profiles), and only 5.7% of participants in YLE23 (profile 2).
To understand why this behavior is so diferent between the channels we are next going to analyze this
content type further.</p>
        <p>While clustering may give us an idea about participant profiles appearing in chat data, chat behavior
may also be described with tracking the most active moments in the timeline of the tournament (see
Figure 4).</p>
        <p>
          The distribution of the most active moments, overall and by content types, was visualized by
percentages along a timeline of the tournament. This metric signifies the starting point of an 8-second window
containing the highest concentration of activity. Most active moments overall across the channels
generally concentrate towards the end of the timeline for both YLE22 and YLE23, but for PCOM22 and
PCOM23 there is also some peak activity concentrated around the midle of the match. The most active
moments vary between the chat content types. The non-unique messages gravitate greatly towards the
end of the match, which makes sense in the form of the volume of “gg” (short for “good game”) messages
when a match ends. The highest activity for unique messages occurs slightly past the midpoint of the
match for every dataset, which could be explained by discussion regarding in-game activity. For the
‘human-written non-unique’ messages no clear trend was found regarding the activity. This could
indicate that these messages are triggered by various events on-screen rather than regular, repetitive
events such as the beginning or ending of a match. Similar observations were noticed in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The
most active moment in ‘commands and replies’ reinforces the previous analysis on the diferences of
Pelaajat.com and Yle channels regarding interaction using these types of messages. For both of the YLE
datasets, the chat activity greatly concentrated on the early timeline. This concentration could indicate
that at the start of each match participants use informative commands, for example, in order to get
further information about the ongoing match. For Pelaajat.com we see density throughout the whole
match rather than a single spike at the start like for Yle. This indicates consistent behavior regarding
the use of commands during a match.
        </p>
        <p>The analysis of top commands used across channels provides insight into the distinct priorities of
the participants in diferent livestreams (see Table 5). On the Pelaajat.com channel, a great number
of commercially oriented commands dominate the stream. Commands like "!grandiosa" (popular
brand of frozen pizza in Finland), "!beefmode" (partnership with a beef jerky brand), and "!gigantti"
(Finnish electronics store) strongly suggest frequent giveaways and sponsor-driven promotions keeping
participants engaged with the channel, especially during the 2023 tournament. In contrast, the chat
participants on the Yle channel, which is lacking the commercial promotional campaigns, are primarily
focused on core game details (as a public broadcast company, YLE does not run commercial advertisement
campaigns). Commands such as "!maps", "!casters", and "!results" are the most common. Participants
demonstrate a clear desire for additional context about the ongoing or upcoming match regarding the
map and teams going against each other, and the commentators that are casting the match. The frequent
use of the Finnish equivalents of these commands such as “!kartat” or “!selostajat” (Finnish for !maps or
!casters) further points to an audience invested in understanding the specifics of the match in question.</p>
        <p>Overall, these command usage patterns indicate a diference in audience’s interests and preferences
on these channels.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and conclusions</title>
      <p>In this paper, we aspired to use two types of clustering, activity and content-based, to classify the chat
behavior of audiences on livestream platform Twitch for the two consecutive CS:GO Majors 2022 and
2023, focusing on Finnish language channels. The typical commenting behaviors were used to identify
participant profiles. After carefully conducting a pair of clustering and analysis on the participant
profiles they compose, we ended up with six participant profiles using activity-based clustering, and 11
participant profiles using content-based clustering. The disparity in the number of participant profiles
between the two approaches suggests that when using the activity features the participant profiles are
clean-cut (same profiles are visible in diferent channels and times), but using content features there
is a lot more variety in the commenting behaviour. In content-based clustering, we first used a
deeplearning based approach to classify chat messages into four distinct categories by "uniqueness" using a
transformer-based chat content detection model to create content-based features for all participants.
This approach made it possible for us to focus on the commenting style of participants, ofering a chance
to uncover deeper qualities of user-generated content, and provided insights into engagement patterns
that are generalizable across diferent contexts.</p>
      <p>Another observation in our analysis is that content-based clustering produced more clusters compared
to activity-based clustering. This diference is likely due to the nature of the features used in each
clustering approach. The features in activity-based clustering rely on message- and time related
information, while content-based clustering incorporates features of the uniqueness of a message,
which capture the diferences in participants’ communication styles in livestream chats. This leads to a
wider range of participant profiles, reflected in the larger number of clusters.</p>
      <p>Based on the t-SNE figures (see Figure 1 and 2), activity-based clustering shows more clearly defined
clusters than content-based clustering, where for each dataset there are around four circle-shaped
clusters lying outside of the other clusters that form a united group in the middle. These circle-shaped
clusters situated mostly on the outside belong to participant profiles that are predominantly presenting
only one content based feature. In activity-based clustering the participants that prefer sending a single
message can be found in the string-like formations. Other clusters in activity-based clustering are
distinctively far away from other clusters. Both types of clustering show us original properties of the
data, complementing each other.</p>
      <p>Moreover, we find our approach useful to understand participation in livestream chats from multiple
viewpoints. Activity-based participant profiling gives information about the amount of efort the
participants bring into the chatting, while content-based participant profiling describes the type of
content participants create. Combining both approaches allows forming a comprehensive understanding
of a particular participant’s chat behavior, rather than preferring one type of clustering over another.</p>
      <p>Besides diferences between the two types of clustering, our analysis also reveals substantial disparities
in the participant profiles between the Yle and Pelaajat.com channels, especially within the PCOM2023
data, based on chat content. Throughout the 2023 tournament Pelaajat.com exhibited four distinct
participant profiles out of which three were only present on this channel, each characterized by a
majority of ‘commands and replies’ within messages, consisting between 50% to 98% of their total
message distribution and accounting for over 48% of all the active participants. These disparities can be
explained by the frequent giveaways and sponsored promotions within the channel. This influence
on participant behavior could also explain why, for instance, the cluster shapes for PCOM23 difer
from other datasets in Figure 1. Therefore it seems that commercial promotions can substantially afect
audience behavior in livestream chats. This information could provide insights for esports organizations
aiming to enhance audience engagement. For example, knowing which clusters tend to respond to
promotional content or giveaways could help businesses tailor their engagement strategies to fit specific
participant profiles, leading to more efective interaction tactics and potentially improving viewer
retention during these events.</p>
      <p>Naturally, there are limitations to our work. These limitations include using data from only Finnish
broadcasters. In this paper, we decided to focus mainly on a singular geo-political region, Finland, and
its popular broadcasters, Yle and Pelaajat.com. Our study could be expanded by including livestream
chat data from international broadcasters as well, or to compare the participant profiles created based
on this data and data from international broadcasters.</p>
      <p>The participant profiles in our analysis showed mostly ’unique’ message content for the total messages
of both of the tournaments combined across both channels. However, these patterns may not generalize
well to the even more massive livestreams. There may very well be diferences, e.g., in the type of
uniqueness that permeates the chats of non-Finnish livestreams. Is there a distinct type of audience in
these smaller national channels that difers from the larger "sports crowd" associated with the main
international streams of the tournaments? Further research is needed to investigate potential diferences
in chat participants’ behavior on an international scale based on their message content types.</p>
      <p>Another limitation lies in the shortage of methods and features that could have been used. To keep
the scope of our study reasonable, we only focused on the active part of the tournament and left out
metadata that was originally collected from livestream chat data. In future research, the inactivity
period between games as well as metadata, like information about banned participants’ messages, could
be included to examine chat behavior. The research we presented here could also be expanded by using
several diferent clustering, or other unsupervised, methods and by making more technical analysis
on the implications each method brings. There can also be some constraints related to the basis of
the unique content detection model, because of the constantly evolving subculture of emotes and the
emergence of new instances of text-based memes. For example, third-party browser based plugins
like 7TV or BetterTTV allows integration of custom emotes into any word. This can add another
layer of complexity to interpret the message, making it dificult to distinguish between the intentional
use of a custom emote or a text-based meme. This poses an ongoing challenge for classification of
these types of messages. Future research could focus on refining the chat content model to enable
more enhanced classification of the proposed content types, as well. Potential examples include the
direct categorization of ‘non-unique’ messages into distinct subcategories, such as ’emotes’, ’gaming
slang’, and ’copypasta’. This fine-grained classification would provide deeper insights into the nature
of communication patterns within chat participants’ messages, especially as livestream subcultures
continue to evolve.</p>
      <p>
        The identification of distinct participant profiles in online livestream chat communities, as
demonstrated in this study of Finnish-language Twitch users in esports streams, opens new areas for
understanding these dynamic social spaces. By combining both activity-based and content-based clustering,
this research ofers a more broad understanding of user engagement. The resulting participant profiles
not only describe audiences based on the volume of their chat activity (e.g., prolific vs. infrequent
chatter) but also based on the nature of their content (e.g., unique contributors vs. repetitive posters).
While we focused more on engagement patterns for broader applicability, future research might build
upon these findings by examining thematic clustering within these chat messages, using approaches
such as topic modeling, to explore specific conversation topics and community dynamics in more depth.
While this methodological approach has only been applied here specifically into the Twitch context, it
still holds promise for a more broad application in the analysis of online communities across diferent
livestreaming platforms. It could give a diferent point of view to research made on chat data from
other livestream platforms, like Youtube [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ]. Additionally, instead of using an existing framework or
theory for explaining participant behavior or communication styles and basing our profiling research
on a set of assumptions [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], we decided to utilize an unsupervised method (clustering) to bring
out qualities that are not necessarily readily visible to the researcher but exist as a groupable substance.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This research was supported by the FIN-CLARIAH infrastructure project (The Research Council of
Finland 358726).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Carter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Egliston</surname>
          </string-name>
          ,
          <article-title>The work of watching twitch: Audience labour in livestreaming and esports</article-title>
          ,
          <source>Journal of Gaming &amp; Virtual Worlds</source>
          <volume>13</volume>
          (
          <year>2021</year>
          )
          <fpage>3</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N. T.</given-names>
            <surname>Taylor</surname>
          </string-name>
          ,
          <article-title>Now you're playing with audience power: The work of watching games</article-title>
          ,
          <source>Critical Studies in Media Communication</source>
          <volume>33</volume>
          (
          <year>2016</year>
          )
          <fpage>293</fpage>
          -
          <lpage>307</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bulygin</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Musabirov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Suvorova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Konstantinova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Okopnyi</surname>
          </string-name>
          ,
          <article-title>Between an arena and a sports bar: Online chats of esports spectators</article-title>
          , arXiv preprint arXiv:
          <year>1801</year>
          .
          <volume>02862</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Backes</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          <article-title>Games and beyond: Analyzing the bullet chats of esports livestreaming</article-title>
          ,
          <source>in: Proceedings of the International AAAI Conference on Web and Social Media</source>
          , volume
          <volume>18</volume>
          ,
          <year>2024</year>
          , pp.
          <fpage>761</fpage>
          -
          <lpage>773</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Niranjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lamba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kumaraguru</surname>
          </string-name>
          ,
          <article-title>Characterizing and detecting livestreaming chatbots</article-title>
          ,
          <source>in: Proceedings of the 2019 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>683</fpage>
          -
          <lpage>690</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nicolaou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <article-title>Twitchchat: A dataset for exploring livestream chat</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment</source>
          , volume
          <volume>16</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>259</fpage>
          -
          <lpage>265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Flores-Saviaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Flores</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seering</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reeves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Savage</surname>
          </string-name>
          ,
          <article-title>Audience and streamer participation at scale on twitch</article-title>
          ,
          <source>in: Proceedings of the 30th ACM Conference on Hypertext and Social Media</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>277</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Merayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cotelo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Carratalá-Sáez</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. J. Andújar,</surname>
          </string-name>
          <article-title>Applying machine learning to assess emotional reactions to video game content streamed on spanish twitch channels</article-title>
          ,
          <source>Computer Speech &amp; Language</source>
          <volume>88</volume>
          (
          <year>2024</year>
          )
          <fpage>101651</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wakamiya</surname>
          </string-name>
          , E. Aramaki,
          <article-title>Ofensive language detection on video live streaming chat</article-title>
          ,
          <source>in: Proceedings of the 28th international conference on computational linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1936</fpage>
          -
          <lpage>1940</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Yousukkee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wisitpongphan</surname>
          </string-name>
          ,
          <article-title>Analysis of spammers' behavior on a live streaming chat</article-title>
          ,
          <source>IAES International Journal of Artificial Intelligence</source>
          <volume>10</volume>
          (
          <year>2021</year>
          )
          <fpage>139</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Janet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nikam</surname>
          </string-name>
          , et al.,
          <article-title>Real time malicious url detection on twitch using machine learning</article-title>
          ,
          <source>in: 2022 International Conference on Electronics and Renewable Systems (ICEARS)</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>1185</fpage>
          -
          <lpage>1189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Melhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gravina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. N.</given-names>
            <surname>Yannakakis</surname>
          </string-name>
          ,
          <article-title>Moment-to-moment engagement prediction through the eyes of the observer: Pubg streaming on twitch</article-title>
          ,
          <source>in: Proceedings of the 15th International Conference on the Foundations of Digital Games</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cha</surname>
          </string-name>
          ,
          <article-title>Learning how spectator reactions afect popularity on twitch</article-title>
          ,
          <source>in: 2020 IEEE International Conference on Big Data and Smart Computing (BigComp)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>154</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Gros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wanner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hackenholt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zawadzki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Knautz</surname>
          </string-name>
          ,
          <article-title>World of streaming. motivation and gratification on twitch</article-title>
          ,
          <source>in: Social Computing and Social Media. Human Behavior: 9th International Conference, SCSM</source>
          <year>2017</year>
          ,
          <article-title>Held as Part of HCI International 2017</article-title>
          , Vancouver, BC, Canada, July 9-
          <issue>14</issue>
          ,
          <year>2017</year>
          , Proceedings,
          <source>Part I 9</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Why do audiences choose to keep watching on live video streaming platforms? an explanation of dual identification framework</article-title>
          ,
          <source>Computers in human behavior 75</source>
          (
          <year>2017</year>
          )
          <fpage>594</fpage>
          -
          <lpage>606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hilvert-Bruce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Neill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sjöblom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hamari</surname>
          </string-name>
          ,
          <article-title>Social motivations of live-streaming viewer engagement on twitch</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>84</volume>
          (
          <year>2018</year>
          )
          <fpage>58</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Robb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mirza-Babaei</surname>
          </string-name>
          ,
          <article-title>Profiling livestream spectators</article-title>
          ,
          <source>in: Extended Abstracts of the 2020 Annual Symposium on Computer-Human Interaction in Play</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>403</fpage>
          -
          <lpage>407</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-J. Choe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , G.-Y. Noh,
          <article-title>The role of wishful identification, emotional engagement, and parasocial relationships in repeated viewing of live-streaming games: A social cognitive theory perspective</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>108</volume>
          (
          <year>2020</year>
          )
          <fpage>106327</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>F.</given-names>
            <surname>Neus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nimmermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schramm-Klein</surname>
          </string-name>
          ,
          <article-title>Diferences and similarities in motivation for ofline and online esports event consumption</article-title>
          ,
          <source>Proceedings of the 52nd Hawaii International Conference on System Sciences</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Beres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klarkowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mirza-Babaei</surname>
          </string-name>
          ,
          <article-title>Exploring esports spectator motivations</article-title>
          ,
          <source>in: CHI Conference on Human Factors in Computing Systems Extended Abstracts</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Schuck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Altmeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krüger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lessel</surname>
          </string-name>
          ,
          <article-title>Viewer types in game live streams: questionnaire development and validation, User Modeling and User-Adapted Interaction 32 (</article-title>
          <year>2022</year>
          )
          <fpage>417</fpage>
          -
          <lpage>467</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Kairam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Mercado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Sumner</surname>
          </string-name>
          ,
          <article-title>A social-ecological approach to modeling sense of virtual community (sovc) in livestreaming communities</article-title>
          ,
          <source>Proceedings of the ACM on human-computer interaction 6</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kucek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Toepfer</surname>
          </string-name>
          ,
          <article-title>Active within structures: Predictors of esports gameplay and spectatorship</article-title>
          ,
          <source>Communication &amp; Sport</source>
          <volume>10</volume>
          (
          <year>2022</year>
          )
          <fpage>195</fpage>
          -
          <lpage>215</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>A systematic review of literature on user behavior in video game live streaming</article-title>
          ,
          <source>International journal of environmental research and public health 17</source>
          (
          <year>2020</year>
          )
          <fpage>3328</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cauteruccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kou</surname>
          </string-name>
          ,
          <article-title>Investigating the emotional experiences in esports spectatorship: The case of league of legends</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>60</volume>
          (
          <year>2023</year>
          )
          <fpage>103516</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>V.</given-names>
            <surname>Diwanji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferchaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seibert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Weinbrecht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sellers</surname>
          </string-name>
          ,
          <article-title>Don't just watch, join in: Exploring information behavior and copresence on twitch</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>105</volume>
          (
          <year>2020</year>
          )
          <fpage>106221</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gardner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Horgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsaasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rickman</surname>
          </string-name>
          ,
          <article-title>Chat speed op pogchamp: Practices of coherence in massive twitch chat</article-title>
          ,
          <source>in: Proceedings of the 2017 CHI conference extended abstracts on human factors in computing systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>858</fpage>
          -
          <lpage>871</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nematzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. L.</given-names>
            <surname>Ciampaglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-Y.</given-names>
            <surname>Ahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Flammini</surname>
          </string-name>
          , Information overload in group communication:
          <article-title>From conversation to cacophony in the twitch chat</article-title>
          ,
          <source>Royal Society open science 6</source>
          (
          <year>2019</year>
          )
          <fpage>191412</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>J.</given-names>
            <surname>Seering</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hammer</surname>
          </string-name>
          , G. Kaufman, D. Yang,
          <article-title>Proximate social factors in first-time contribution to online communities</article-title>
          ,
          <source>in: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>26</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>K.</given-names>
            <surname>Konstantinova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bulygin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Okopny</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Musabirov</surname>
          </string-name>
          ,
          <article-title>Online communication of esports viewers: Topic modeling approach</article-title>
          , in: Advances in Computer Entertainment Technology: 14th International Conference,
          <string-name>
            <surname>ACE</surname>
          </string-name>
          <year>2017</year>
          , London, UK, December
          <volume>14</volume>
          -
          <issue>16</issue>
          ,
          <year>2017</year>
          , Proceedings 14, Springer,
          <year>2018</year>
          , pp.
          <fpage>608</fpage>
          -
          <lpage>613</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , E. Duchesnay,
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>T. M. Alenezi</surname>
            ,
            <given-names>T. H.</given-names>
          </string-name>
          <string-name>
            <surname>Sulaiman</surname>
            ,
            <given-names>A. M. AbdelAziz,</given-names>
          </string-name>
          <article-title>Applying machine learning models to electronic health records for chronic disease diagnosis in kuwait</article-title>
          .,
          <source>International Journal of Advanced Computer Science &amp; Applications</source>
          <volume>14</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Rousseeuw</surname>
          </string-name>
          ,
          <article-title>Silhouettes: a graphical aid to the interpretation and validation of cluster analysis</article-title>
          ,
          <source>Journal of computational and applied mathematics 20</source>
          (
          <year>1987</year>
          )
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>L.</given-names>
            <surname>Van der Maaten</surname>
          </string-name>
          , G. Hinton,
          <article-title>Visualizing data using t-sne.</article-title>
          ,
          <source>Journal of machine learning research 9</source>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lindroos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peltonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Välisalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Koskimaa</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Toivanen</surname>
          </string-name>
          ,
          <article-title>From PogChamps to insights: Detecting original content in twitch chat</article-title>
          ,
          <source>Proceedings of the 58th Hawaii International Conference on System Sciences</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bholowalia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Ebk-means: A clustering technique based on elbow method and k-means in wsn</article-title>
          ,
          <source>International Journal of Computer Applications</source>
          <volume>105</volume>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>D. T.</given-names>
            <surname>Larose</surname>
          </string-name>
          ,
          <article-title>Data mining and predictive analytics</article-title>
          , John Wiley &amp; Sons,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>C.</given-names>
            <surname>Liebeskind</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liebeskind</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yechezkely</surname>
          </string-name>
          ,
          <article-title>An analysis of interaction and engagement in youtube live streaming chat</article-title>
          ,
          <source>in: 2021 IEEE SmartWorld, Ubiquitous Intelligence &amp; Computing</source>
          ,
          <string-name>
            <given-names>Advanced &amp; Trusted</given-names>
            <surname>Computing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Scalable</given-names>
            <surname>Computing</surname>
          </string-name>
          &amp;
          <article-title>Communications, Internet of People and Smart City Innovation (SmartWorld/SCALCOM</article-title>
          /UIC/ATC/IOP/SCI), IEEE,
          <year>2021</year>
          , pp.
          <fpage>272</fpage>
          -
          <lpage>279</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>