<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting Tomorrow's Headline using Twitter Deliberations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roshni Chakraborty</string-name>
          <email>roshni.pcs15@iitp.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abhijeet Kharat</string-name>
          <email>abhijeet.mtcs17@iitp.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Apalak Khatua</string-name>
          <email>apalak@xlri.ac.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sourav Kumar Dandapat</string-name>
          <email>sourav@iitp.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joydeep Chandra</string-name>
          <email>joydeep@iitp.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IIT Patna</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>XLRI Jamshedpur</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Predicting the popularity of a news article is a challenging task. Existing literature mostly focused on article contents and polarity to predict the popularity. However, existing research has not considered the users preference towards a particular article. Understanding users preference is an important aspect for predicting the popularity of news articles. Hence, we consider social media data, from the Twitter platform, to address this research gap. In our proposed model, we have considered the users involvement as well as the users reaction towards an article to predict the popularity of the article. In short, we are predicting tomorrows headline by probing todays Twitter discussion. We have considered 300 political news articles from the New York Post, and our proposed approach has outperformed other baseline models.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Gone are those days when an o ce going New Yorker
used to board the subway with a folded newspaper in
his hand. Reading morning newspapers on New Yorks
subway is becoming outdated. Things have changed
Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
BY 4.0).
drastically in recent times. Todays millennial
generation is not only emotionally but also physically tied to
their smartphones and tablets. This has severely
affected the newspaper industry. All leading newspapers
across the globe have reported a sharp drop in their
print circulation. So, the future of this industry lies on
the digital platform. The competition in this
newspaper industry is not anymore about sending the print
version to the remotest corner of the country. The
challenge of this digital platform is to understand the
latent psychological aspects of the users. If a
newspaper fails to satisfy the user, then within the next few
seconds she will switch to another news-related app.
This will directly impact the ad revenue of a news
outlet. Between various news related apps, and
various social media platforms, users these days are spoilt
for choice. Customer loyalty is a concept of a bygone
era in this digital age, and the customers preferences
are also not homogeneous. In brief, the phenomenal
growth of online news consumption and innumerable
news sources has signi cantly increased the
competition among news media outlets. Further, the
continuous in ux of newsworthy events further aggravates
the situation. Thus, for media outlets, the need of
the hour is to develop an automated system that can
help them to predict which of the todays headlines will
maintain its popularity tomorrow.</p>
      <p>Existing literature has attempted to address this.
However, this stream of research mostly explored
various features and contents of the articles and the
title of the articles [FVC15, LWZ+17]. Prior
studies considered the subjectivity and polarity of
contents [FVC15, KWHR16], the sentiment of the
headline [RBdM+15], and so on. In other words, these
works focus broadly into the articulation aspects of
an article. Few studies also considered the
importance of an event to predict the popularity of news
article [SAMA17]. One of the major shortcomings of
the above approach is that hypothetically, two articles
might have similar feature and polarity, but the
reaction of readers might be di erent. It would be similar
to comparing an apple to an orange even though they
might be nearly similar in shape and weight. We argue
that probing the social media platform can hint which
is orange, and which is apple. Social media platforms
can hint about the users preference.</p>
      <p>Nowadays social media platforms, such as Twitter,
generate an enormous amount of user-generated data,
and many times this social media platforms become
the mirror of the society. Existing literature has
successfully explored the Twitter data to predict the
election outcome [KKGC15], to understand social
movements [KK16a], to tackle disasters and epidemic
outbreaks [KK16b]. Therefore, we argue that
understanding the ner nuances of Twitter deliberation can be
bene cial to predict the popularity of news article on
digital platforms. This paper attempts to address this
research gap.</p>
      <p>Prior studies noted that analyzing social media
platform could shed light regarding the
popularity [KHGPS16] and the life cycle of various news
article [Cas13]. However, these studies have considered
tweets, which have exclusively mentioned the URL of
news article. In fact, these studies failed to probe the
richness of the Twitter platform by restricting them to
a very small sub-sample of tweets with the URL of a
speci c news article. On the contrary, our study takes
a more holistic approach than these studies and
considers both users involvement with a news article and
user reaction towards a speci c news article. However,
the biggest challenge for this approach is to identify the
relevant tweets for a speci c news article. Therefore,
we have developed an iterative and adaptive algorithm
that considers both textual and semantic attributes to
identify the relevant tweets. Our user involvement
aspect considers various count measures, such as a
total number of tweets and average number of retweets,
count of hashtags, the cumulative number of unique
users as well as in uential users, and so on. Also, our
user reaction indices consider linguistic aspects of the
Twitter discussion, such as variances in sentiment and
emotion for a particular news article. For the sake of
robustness, we have considered various machine
learning algorithms and an exhaustive set of baseline
models based on prior studies. Our ndings on the basis of
300 news items strongly suggest that patterns of
Twitter deliberations can outperform other baseline models
in predicting the popularity of the news articles.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>Predicting the popularity of news article is a
wellresearched area. However, the genesis of this
research lies in the prior works on news recommender
system. News recommender system research mostly
focused on the personal preference of an
individual user [LXG+14]. So, understanding the
userlevel latent political leaning, or bias towards a
certain sport or a team, can help to predict the
suitable news article for an individual user. For
instance, prior studies considered users historical
preferences [WLC+10], social network data [DFMGL12,
AGHT11], user feedback [SBZ11] or combination of
both user preferences and user feedback [LWL+11,
LCLS10, MGARLGMM13], but this stream of studies
struggled due to lack of adequate data. Moreover, the
users interest can vary over time, and historical data is
not available for new users. Thus, prior studies have
attempted to mitigate the challenges by considering
opinions of social in uencers [LXG+14], topic or
temporal [XXLZ14] relationships between news items and
users [LL13] or analyzing user communities [ZLHL13].
These approaches yield better results in comparison to
initial studies, but still, the accuracy of ltering news
articles for a newspaper is not satisfactory.</p>
      <p>In comparison to news recommender system for an
individual user, predicting the popularity of a news
article is a complex task because an e cient
prediction model needs to account for the heterogeneity of
users. More importantly, the summation of
individual user preference would not be the proxy for
societal acceptance For instance, it is easy to predict what
a Democrat or Republican will prefer to read on the
digital platform but the task becomes complex if we
try to predict what political news will engage both
Democrats and Republicans. So, a dominant stream of
prior works focused on the content of the news article
and article headline and employed machine learning
algorithms [SAM+16] to predict the popularity of a news
article. For instance, existing literature considered
different features of an article [VCLDD17, KWHR16],
such as textual [FVC15] and temporal features, to
predict its popularity. Prior studies noted that
article content features, such as the length of the article,
the time of publishing, category or genre of the
article, the author of the article and so on, can predict
the popularity of a news article [LWZ+17]. Another
set of studies also considered the linguistic attributes
of an article [KMJO16, KYS+17] or the presence of
important entities within an article [SS16] to
investigate the issue. Existing literature also noted that the
headline of an article itself [KFKN15] and the polarity
within the headline [RBdM+15] could be important
input variables for the predicting the popularity of a
news article.</p>
      <p>Another set of works highlighted the event
importance to predict the popularity of a news article. For
instance, Setty et al. [SAMA17] ranked news articles
by linking them to a chain of recent news events.
Similarly, other studies tried to explore the event
importance by combining articles using topic similarity from
Wikipedia [MB16] or by considering the causal
relationships [KVW14]. However, this approach has
limitations for new upcoming events or for an event which
is losing relevance among readers. In these scenarios,
we argue that probing users behavioral pattern on
social media platform can hint about the popularity of
news article.</p>
      <p>To the best of our knowledge, there is hardly
any study which considered social media platform for
predicting the popularity of news article. Some of
the prior studies considered the users behavior on a
news media outlet and argued that engagement of
users could predict the popularity of a news
article [TADAF14, TLA+11]. However, it is worth
noting that the news media outlet represents a minuscule
of digital platform readers. This is one potential
research gap in the existing literature. Popular social
media platforms, such as Twitter, not only provides
users an option to share their views but also allows to
reply or endorse the views of others by retweeting. In
other words, Twitter provides a platform for its users
to engage in a deliberation. Interaction of users on the
Twitter platform can shed light about users
engagement with a particular news article [OCDA15]. Thus,
this paper attempts to predict the popularity of news
article using Twitter data. Some of the earlier works
consider initial twitter reactions [MTR14, CEHPS14]
or content and structural features [LZZ15]. However,
as we mentioned, they have only considered tweets
that has news related URLs. Thus, none of the prior
studies considered the richness of Twitter data. So,
this paper not only considers the user-level
involvement (by using count measures of tweets, retweets,
number of unique users and other parameters) but also
probes user-level reaction towards a particular news
article (by understanding the various linguistic aspects
of their tweets).
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data Collection</title>
      <p>To address our research problem, we have considered
the front-page political news of the New York Post,
which is one of the most popular newspapers in the
United States. It has experienced a whopping 500%
growth in the last ve years with 331 million page
views in March 2018. Predicting the popularity of
political news is the most challenging in comparison to
other genres of news. For instance, a news article on
climate change will uniformly a ect all users.
Therefore, it is easy to predict the reaction of users.
However, the political news might not uniformly engage
and a ect all users because of their ideological
heterogeneity. In this study, we have considered 300 political
news, from the New York Post, during the period July
2016 to September 2016.</p>
      <p>We have considered the Twitter platform for
collecting the social media data. Twitter allows free access
to approximately 1% of total tweets (in a random
fashion) using the streaming API. To probe our research
question, we have considered the tweets related to a
particular news article. Extracting the relevant tweets
for particular political news is a challenging task. To
address this, we have developed an adaptive algorithm
that has considered both content (similar keyword
mapping) features and context (same hashtag)
features of tweets to extract the related tweets of a news
article. Following prior studies [CBDC17], as an initial
step, we have considered a set of preliminary
hashtags that have threshold keywords overlap with the
representative (by top 10 TF-IDF) keywords within
the news article. In other words, these preliminary
hashtags, which we have initially considered to crawl
tweets for a particular news article, are a bag of
hashtags on the basis of the articles seed tweets [CBDC17].
However, there are certain limitations to this approach
because users can use multiple hashtags for a
particular news article on the social media platform but not
all of them might be unique to that particular news
article. For instance, social media users have used
multiple hashtags for the following news article titled
GOP blasts Obama 400 million dollars secret ransom
paid to Iran as follows: #whitehouse, #trump2016,
#chicago, #irandeal, #obamabetrayus and so on (as
shown in Table 3). The last two hashtags are more
speci c about the news article in comparison to others.
Hence, we need to consider this in our data collection
as well as analysis.</p>
      <p>To address the above concern, we have collected all
hashtags related to all political news articles published
in the previous one month (with respect to the
publication date of the article we are considering in our
analysis). This process has generated a bag of
hashtags. From this bag of hashtags, we have identi ed a
set of hashtags, which were frequently used by Twitter
users and labeled them as generic hashtags.
Consequently, we have labeled #whitehouse, #trump2016,
#chicago as generic hashtags for the above article
because these hashtags were used by Twitter users for
other issues/news article also.</p>
      <p>Next, we have developed an automated system for
identifying hashtags speci c to a news article. We
have ltered out the preliminary hashtags as the
hashtags those were mentioned in a tweet T, and
ful1
2
3
4
lled the threshold criteria of keywords matching with
the news article [CBDC17]. We de ne article-speci c
hashtags as those hashtags that were mentioned in T
but not in our list of generic hashtags. For instance,
Thus, the article 1 (as shown in Table 3), we have
labeled #irandeal and #obamabetrayus as
articlespeci c hashtags. Similarly, the article-speci c
hashtag for the news article 4 (as shown in Table 3), i.e.,
Obamacare hikes has families struggling to a ord
insurance was #repealobamacare.</p>
      <p>To check the accuracy of this approach, we have
provided around 130 news articles along with all the
hashtags to three annotators. We have labeled a hashtag
as an article-speci c hashtag if the majority of
annotators have marked that particular hashtag as speci c to
that article, or otherwise labeled it as a generic
hashtag. We observed that our proposed approach yields
an accuracy of 89%in identifying an article speci c
hashtags. (as shown in Table 3) reports a few sample
news articles and corresponding article-speci c (HA)
and generic (HG) hashtags. After identifying the
article speci c hashtags, we use these hashtags to extract
further tweets related to that news article.</p>
      <p>Next, we have extracted the user level information.
In other words, we have extracted the information
related to users who had participated in the political
discussion related to any of these 300 news articles.
We have extracted the users name from our Twitter
corpus and identi ed around 1 million unique users
who have tweeted at least once for our sample of 300
news article. Next, we have crawled their last 3200
tweets and pro le-related information. In our model,
we have considered whether a user is in uential or not.
If a user has more than 1000 followers, then we have
considered them as an in uential user. To sum up,
we have considered 1:8 million tweets for our 300 news
articles made by around 1 million unique users.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Proposed Approach</title>
      <p>For predicting the popularity of a news article, we have
considered two categories of social media data namely,
user involvement indices and user reaction indices. We
argue that the popularity of a news article among the
social media users (which is a proxy for digital
platform readers) can be captured by analyzing the
attention that a news article is receiving and the linguistic
content of the discussion on the Twitter platform on
the very day of its publication. In brief, the former
category considers various user-level tweet statistics, and
the latter employs natural language processing
techniques to understand the linguistic aspects of the social
media discussions. The following sections narrate how
we have operationalized the involvement and reaction
indices.
We capture the attention of Twitter users for a
particular news article by considering the user involvement
through three aspects: tweet statistics, user statistics
and hashtag statistics. Under the tweet statistics
category, we have considered the number of tweets, the
number of retweets and the number of favorites
received by a particular news article on the very day
of publication. These three statistics represents the
user response towards a particular article. We
observe that there is a signi cant variance in user
involvement. Some news article receives hundreds of tweets
and retweets whereas another news article merely
receives ten to twenty tweets. Thus, we have normalized
the number of tweets received by a particular article
by dividing it with the maximum number of possible
tweets that an article can have in a day.</p>
      <p>Intuitively, the number of users get involved with a
particular news article is a good predictor of the
popularity of the news article. Furthermore, we note that
some users get more involved with a particular news
article, and they tweet multiple times in a day.
However, it is worth noting that 10 tweets from 10 di erent
users, in comparison to 10 tweets from 1 particular
user, is a better proxy to gauge the popularity of a
news article. On the contrary, if a user tweets, about
a particular news article, for more than once, then it
also indicates his high involvement of that user with
that particular news article. So, we have considered
these ner variances in our analysis. We have
considered the fraction of users, who have tweeted more than
once for a particular news article, as a ected users.</p>
      <p>Subsequently, it is also important to note whether
a user is in uential on the social media platform or
not. In other words, an in uential person can be an
opinion leader on a social media platform. For
instance, if personalities, such as Barack Obama or
Donald Trump, endorse a particular news article through
their personal/o cial twitter handle then immediately
that tweet will get retweeted by hundreds of their
followers. more than 1000 followers. To sum up, in
addition to generic user statistics we have also considered
the fraction of a ected users and in uential users for
each news article as an input variable in our model.</p>
      <p>Next, we considered the number of article-speci c
hashtags on the Twitter platform as a metric to gauge
the involvement of social media users. Intuitively, it
can be argued that higher user involvement with a
news article would generate higher number of
articlespeci c hashtags. For instance, the news article Trump
gives Post Columnist a shout-out in economic speech
didnt generate article speci c hashtags. On the
contrary, the news article Trump to propose big tax breaks
in economic plan has created 8 article speci c hashtags
(as shown in Table 2). Thus, we have considered the
total number of article-speci c hashtags as an input
variable in our proposed model.
4.2</p>
      <sec id="sec-4-1">
        <title>User Reaction Indices</title>
        <p>As we mentioned earlier, natural language
processing (NLP) techniques allowed us to go beyond various
count-based user-level measures and to probe the
linguistic content of Twitter deliberations to understand
the cognitive involvement of users with a particular
news article. This cognitive involvement of users can
be a good proxy to predict the popularity of a news
article. So, we have employed NLP techniques, such
as sentiment and emotion analysis, to gauge the user
reactions towards a particular news article. We have
considered three indicators to capture the user
reaction namely, sentiment variance, emotion variance and
argumentativeness index.</p>
        <p>We argue that di erences of opinion would lead to
higher debates and discussion on the Twitter platform.
For instance, most social media users would agree with
a news article such as Global warming would be a
serious threat in the coming decades and it might create
a discussion but not debates. However, a
hypothetical news article such as President Trump is failing
to take appropriate policy measures to control global
warming would probably lead to a debate between the
Democrats and Republicans. Republican will try to
discard this view, whether Democrats will try to
justify this view. Consequently, the popularity of this
particular news article will also go up. We are
attempting to capture this in our proposed model.</p>
        <p>Following prior studies, such as Vader Sentiment
Analyzer [HG14] and TextBlob [LKH+14], we have
calculated the average sentiment score of a tweet. We
have identi ed all tweets speci c to a particular news
article and classi ed whether the tweet is positive or
negative. Next, we have considered the sentiment
variance of all tweets related to a particular news article
to understand the di erences in opinions. We have
calculated the sentiment variation (SV) as follows:
SV = 1
j (P C N C) j
j (P C + N C) j</p>
        <p>PC is the number of positive tweets for a news
article, and NC is the number of negative tweets for the
same news article. The sentiment variance is highest
when the count of positive tweets and negative tweets
are equal for a news article, and the sentiment
variance decreases when there is only (or higher number
of) positive/negative tweets. In other words, having
an equal number of positive and negative sentiment
indicates that users are from two ideologically
opposite camps. On the contrary, only positive or negative
tweets indicate that users are ideologically
homogeneous.</p>
        <p>Next, we also considered the emotional content of a
tweet. We employ the NRC emotion lexicon[MT10,
MT13] to classify a tweet among various emotion
classes such as anger, anticipation, trust, disgust, fear,
joy and surprise. Similar to our sentiment variance
analysis, we have considered all tweets speci c to a
particular news article and classify them into
various categories of emotions. Intuitively, high emotional
variance indicates that users are displaying di erent
emotions towards a news article. We have calculated
the emotional variance (EV) as follows:</p>
        <p>EV =</p>
        <p>P8
i=1(e(i)
n
m(e))2</p>
        <p>In the above formula, n is the number of emotion
categories which is 8 [MT10, MT13], e(i) is the fraction
of tweets with ithemotion; andthevalueof m(e)isN 8
where N is the total number of tweets related to a
particular article. Here, the highest emotion variance
indicates that tweet corpus for a particular news article
represents multiple emotion categories. For instance,
in response to the immigration issue related news, a
Republican, who believes that strong immigration law
would protect American jobs, might display joy. On
the contrary, a social activist, who thinks otherwise,
might display her anger to the same news article.
5
5.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Data Analysis</title>
      <sec id="sec-5-1">
        <title>Preparation of Gold Standard</title>
        <p>For our analysis, we need to know whether a
particular news article of the nth day is followed by another
subsequent article on (n + 1)th day. It is important to
note that on (n+1) the day the title or the content of
the subsequent article can di er signi cantly from the
previous day. For instance, on nth day, the
hypothetical title of a news article can be: Why Brexit matters
for the American Corporate Sector? However, on the
(n + 1)th day the issue will continue, but the title can
be: American Corporates are reluctant to invest in the
UK. So, it requires a contextual understanding to
prepare the database for our analysis. Thus, we employed
three annotators and provided them with a particular
news article of nth day and all the news articles of
(n + 1)th day for manual annotation. We have asked
our annotators to mark a news article either as 1 if the
same news gets covered on the subsequent or (n + 1)th
day and 0 otherwise. For our analysis purpose, we
have considered the labeling on the basis of the
majority of the annotators. We have done this for all 300
news articles that we considered for our nal analysis.</p>
        <sec id="sec-5-1-1">
          <title>List of Features/Input Variables</title>
          <p>no of words in the article; the rate of non-stop words; day of
the week on which it got published; published on weekend
or not;no of entities in the news article; average word length
of the article
Polarity score of the article; the rate of positive and
negative words per 100 words; the rate of positive and negative
words per 100 words with non-neutral words, the average
polarity of positive and negative words; min. and max. the
polarity of positive and negative words
no of words in the title; the rate of non-stop words in the
title;no of entities in the title; the average word length of
the title
Polarity score of the title; the rate of positive and negative
words in the title; the rate of positive and negative words
with non-neutral words in the title; average polarity of
positive and negative words in the title; min. and max. the
polarity of positive and negative words in the title
no of days a news article related to the event was published,
no of articles of the event was published
As we discussed in our literature review, a plethora of
studies have tried to predict the popularity of news
article [KMJO16, KFKN15, KYS+17, RBdM+15, SS16].
However, this stream of literature is broadly classi ed
into ve categories as follows: article content and
polarity, title content and polarity, and event importance
categories(as shown in Table 4). We have considered
all these ve prediction models as our baseline models.</p>
          <p>Following the prior studies [KFKN15, RBdM+15,
KVW14], we have extracted both the content and
polarity of the article to predict whether the article will
get published on the next day or not. Here, we
employed NLP techniques to the understand the overall
sentiment of the article, usages of positive or negative
words within the article, the length of the article, and
so on. Similarly, we have also considered the content
and polarity of the title of the article to predict its
popularity. Here, our analysis is restricted only to the
title of the article. Prior studies portray certain
differences in terms of the number of features. However,
in our studies, we have tried to consider an exhaustive
set of features for rst four baseline models: article
content and polarity, title content and polarity. The
nal baseline model is the event importance [SAMA17].
The event importance tries to capture the dominance
of a topic/issue in comparison to others. So, the event
importance of a news article is calculated by the
numbers of similar articles that get published on
consecutive days. A pair of news article will be considered
as a similar article if it crosses the threshold among
the list of entities and bi-grams features between two
articles. Prior studies noted that event importance is
also an indicator to predict the popularity of a news
article.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Results and Discussions</title>
      <p>We have employed Random Forest Classi ed (RFC),
Support Vector Machine (SVM), Gradient Boosting
Classi er (GBC), and Classi cation and Regression
Trees (CART) algorithms for our analysis. We have
applied these four classi ers on our article dataset both
for the ve baseline models as well as our proposed
models. We have considered ten-fold cross-validation
for our analysis. We have repeated our experiments
multiple times and found our results are consistent.
We have reported the same in Table 5. Our proposed
model has outperformed all ve baseline models for
all three classi ers. Our F1-score for SVM and RFC
classi ers are marginally better than the CART
classi er. Broadly, the SVM classi er has outperformed
other classi ers not only for our proposed model but
also for baseline models.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>The advent of information and communication
technology has a ected the newspaper industry severely in
last few decades. Seamlessly connected various
communication channels are generating a huge volume of
information. Moreover, the digital platform is
becoming a crowded place. Multiple news outlets are
struggling to grab a larger share of this platform. Therefore,
selecting a potentially popular news article is
becoming a daunting task for the journalists. This leads to
the requirement of an automated system, which can
e ciently select the news article that will most likely
draw the maximum attention of users on the digital
platform. To address this, prior studies focused mostly
on the content and polarity of the news article to
predict the popularity of news articles. However, these
studies failed to capture the latent psychological
aspects of users. Thus, our proposed approach is trying
to gauge the users perception from the social media
discussions. We have consideredthe Twitter platform
for our study. Our proposed model has incorporated
users involvement and reaction towards a particular
news article. In short, as our title suggests that we
are trying to predict tomorrows popular headline by
considering todays discussion on Twitter platform.We
have employed various machine-learning algorithms to
test the accuracy of our proposed approach. We
observe that our proposed approach ensures higher
accuracyin comparison to other baseline models.
Considering Twitter discussion for predicting the popularity
of news article is the core contribution of this study.
However, there are certain shortcomings of our
proposed mode which future research needs to address.
Firstly, we have considered a small sample of 300-news
article for a relatively shorter period. Future studies
in this area should consider a larger sample and longer
time horizon. Secondly, we have considered the
political news of the New York Post. In other words,
we have tested our proposed approach in the political
sphere of the United States. So, future studies need
to probe the e cacy of our model for other genres of
news in other countries. The biggest challenge will be
to extrapolate this approach to a context where native
language is not English. Finally, we have considered a
few fundamentalmachine-learning algorithms. Future
studies need to consider advanced deep learning based
models to see the accuracy of our model.
[AGHT11]
[Cas13]
[CBDC17]
[CEHPS14]
[DFMGL12]</p>
      <sec id="sec-7-1">
        <title>Fabian Abel, Qi Gao, Geert-Jan Houben, and Ke Tao. Analyzing user modeling on twitter for personalized news recommendations.</title>
        <p>User Modeling, Adaption and
Personalization, pages 1{12, 2011.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Carlos Castillo. Tra c predic</title>
        <p>tion and discovery of news via
news crowds. In Proceedings of the
22nd International Conference on
World Wide Web, pages 853{854.
ACM, 2013.</p>
      </sec>
      <sec id="sec-7-3">
        <title>Roshni Chakraborty, Maitry</title>
        <p>Bhavsar, Sourav Dandapat, and
Joydeep Chandra. A network
based strati cation approach for
summarizing relevant comment
tweets of news articles. In
International Conference on Web
Information Systems Engineering,
pages 33{48. Springer, 2017.</p>
      </sec>
      <sec id="sec-7-4">
        <title>Carlos Castillo, Mohammed El</title>
        <p>Haddad, Jurgen Pfe er, and Matt
Stempeck. Characterizing the life
cycle of online news stories using
social media reactions. In
Proceedings of the 17th ACM
conference on Computer supported
cooperative work &amp; social computing,
pages 211{223. ACM, 2014.</p>
      </sec>
      <sec id="sec-7-5">
        <title>Gianmarco De Francisci Morales,</title>
        <p>Aristides Gionis, and Claudio
Lucchese. From chatter to
headlines: harnessing the real-time web
for personalized news
recommendation. In Proceedings of the fth
ACM international conference on
Web search and data mining, pages
153{162. ACM, 2012.</p>
      </sec>
      <sec id="sec-7-6">
        <title>Joon Hee Kim, Amin Mantrach,</title>
        <p>Alejandro Jaimes, and Alice Oh.
How to compete online for news
audience: Modeling words that
attract clicks. In Proceedings of the
22nd ACM SIGKDD International
Conference on Knowledge
Discovery and Data Mining, pages 1645{
1654. ACM, 2016.</p>
      </sec>
      <sec id="sec-7-7">
        <title>Erdal Kuzey, Jilles Vreeken, and</title>
        <p>Gerhard Weikum. A fresh look on
knowledge bases: Distilling named
events from news. In Proceedings
of the 23rd ACM International
Conference on Conference on
Information and Knowledge
Management, pages 1689{1698. ACM,
2014.</p>
      </sec>
      <sec id="sec-7-8">
        <title>Yaser Keneshloo, Shuguang Wang,</title>
        <p>Eui-Hong Han, and Naren
Ramakrishnan. Predicting the
popularity of news articles. In
Proceedings of the 2016 SIAM
International Conference on Data
Mining, pages 441{449. SIAM, 2016.</p>
      </sec>
      <sec id="sec-7-9">
        <title>Nagendra Kumar, Anusha</title>
        <p>Yadandla, K Suryamukhi, Neha
Ranabothu, Sravani Boya, and
Manish Singh. Arousal prediction
of news articles in social media.
In International Conference on
Mining Intelligence and
Knowledge Exploration, pages 308{319.
Springer, 2017.</p>
      </sec>
      <sec id="sec-7-10">
        <title>Lihong Li, Wei Chu, John Lang</title>
        <p>ford, and Robert E Schapire.
A contextual-bandit approach to
personalized news article
recommendation. In Proceedings of the
19th international conference on
World wide web, pages 661{670.
ACM, 2010.</p>
      </sec>
      <sec id="sec-7-11">
        <title>Steven Loria, P Keen, M Honnibal, R Yankovsky, D Karesh, E Dempsey, et al. Textblob: simpli ed text processing. Secondary</title>
        <p>TextBlob: Simpli ed Text
Processing, 2014.
[MGARLGMM13] Alejandro Montes-Garc a,
Jose Mar a Alvarez-Rodr guez,
Jose Emilio Labra-Gayo, and
Marcos Mart nez-Merino. Towards a
journalist-based news
recommendation system: The wesomender
approach. Expert Systems with</p>
      </sec>
      <sec id="sec-7-12">
        <title>Saif M Mohammad and Peter D</title>
        <p>Turney. Emotions evoked by
common words and phrases: Using
mechanical turk to create an
emotion lexicon. In Proceedings of the
NAACL HLT 2010 workshop on
computational approaches to
analysis and generation of emotion in
text, pages 26{34. Association for
Computational Linguistics, 2010.</p>
      </sec>
      <sec id="sec-7-13">
        <title>Saif M Mohammad and Peter D</title>
        <p>Turney. Crowdsourcing a word{
emotion association lexicon.
Computational Intelligence, 29(3):436{
465, 2013.</p>
      </sec>
      <sec id="sec-7-14">
        <title>Nuno Moniz, Lu s Torgo, and</title>
        <p>F Rodrigues. Improvement of news
ranking through importance
prediction. In Proc. KDD Workshop
on Data Science for News
Publishing (NewsKDD), page 6, 2014.</p>
      </sec>
      <sec id="sec-7-15">
        <title>Alexandra Olteanu, Carlos</title>
        <p>Castillo, Nicholas Diakopoulos,
and Karl Aberer. Comparing
events coverage in online news and
social media: The case of climate
change. In Proceedings of the
Ninth International AAAI
Conference on Web and Social Media,
number EPFL-CONF-211214,
2015.</p>
      </sec>
      <sec id="sec-7-16">
        <title>Julio Reis, Fabricio Benevenuto,</title>
        <p>P Vaz de Melo, Raquel Prates,
Haewoon Kwak, and Jisun An.
Breaking the news: First
impressions matter on online news. In
ICWSM15: Proceedings of The
International Conference on Weblogs
and Social Media, 2015.</p>
        <p>R Shreyas, DM Akshata, BS
Mahanand, B Shagun, and CM
Abhishek. Predicting popularity of
online articles using random
forest regression. In Cognitive
Computing and Information
Processing (CCIP), 2016 Second
International Conference on, pages 1{5.
IEEE, 2016.
[SBZ11]
[TADAF14]
[TLA+11]
[VCLDD17]
[WLC+10]</p>
      </sec>
      <sec id="sec-7-17">
        <title>Vinay Setty, Abhijit Anand,</title>
        <p>Arunav Mishra, and Avishek
Anand. Modeling event
importance for ranking daily news
events. In Proceedings of the Tenth
ACM International Conference
on Web Search and Data Mining,
pages 231{240. ACM, 2017.</p>
      </sec>
      <sec id="sec-7-18">
        <title>Yanir Seroussi, Fabian Bohnert, and Ingrid Zukerman. Personalised rating prediction for new users using latent factor models. In</title>
        <p>Proceedings of the 22nd ACM
conference on Hypertext and
hypermedia, pages 47{56. ACM, 2011.</p>
      </sec>
      <sec id="sec-7-19">
        <title>Pedro Saleiro and Carlos Soares. Learning from the news: Predicting entity popularity on twitter. In</title>
        <p>International Symposium on
Intelligent Data Analysis, pages 171{
182. Springer, 2016.</p>
      </sec>
      <sec id="sec-7-20">
        <title>Alexandru Tatar, Panayotis Antoniadis, Marcelo Dias De Amorim, and Serge Fdida. From popularity prediction to ranking online news.</title>
        <p>Social Network Analysis and
Mining, 4(1):174, 2014.</p>
      </sec>
      <sec id="sec-7-21">
        <title>Alexandru Tatar, Jeremie Leguay,</title>
        <p>Panayotis Antoniadis,
Arnaud Limbourg, Marcelo Dias
de Amorim, and Serge Fdida.
Predicting the popularity of online
articles based on user comments.
In Proceedings of the International
Conference on Web Intelligence,
Mining and Semantics, page 67.
ACM, 2011.</p>
      </sec>
      <sec id="sec-7-22">
        <title>Steven Van Canneyt, Philip Ler</title>
        <p>oux, Bart Dhoedt, and Thomas
Demeester. Modeling and
predicting the popularity of online
news based on temporal and
content-related features.
Multimedia Tools and Applications, pages
1{28, 2017.</p>
      </sec>
      <sec id="sec-7-23">
        <title>Jia Wang, Qing Li, Yuanzhu Peter</title>
        <p>Chen, Jiafen Liu, Chen Zhang, and
Zhangxi Lin. News
recommendation in forum-based social media.
In AAAI, 2010.
[ZLHL13]</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [SS16]
          <string-name>
            <given-names>Zhengyou</given-names>
            <surname>Xia</surname>
          </string-name>
          , Shengwu Xu, Ningzhong Liu, and
          <string-name>
            <given-names>Zhengkang</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>Hot news recommendation system from heterogeneous websites based on bayesian model</article-title>
          .
          <source>The Scienti c World Journal</source>
          ,
          <year>2014</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Li</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Lei</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Wenxing</given-names>
            <surname>Hong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Tao</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Penetrate: Personalized news recommendation using ensemble hierarchical clustering</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>40</volume>
          (
          <issue>6</issue>
          ):
          <volume>2127</volume>
          {
          <fpage>2136</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>