<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>All news</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Biased News Data In uence on Classifying Social Media Posts</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Marija Stanojevic, Jumanah Alshehri, Eduard Dragut, Zoran Obradovic Center for Data Analytics and Biomedical Informatics (DABI) Temple University Philadelphia</institution>
          ,
          <addr-line>Pennsylvania</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>53</volume>
      <issue>2</issue>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A common task among social scientists is to
mine and interpret public opinion using social
media data. Scientists tend to employ o
-theshelf state-of-the-art short-text classi cation
models. Those algorithms, however, require a
large amount of labeled data. Recent e orts
aim to decrease the compulsory number of
labeled data via self-supervised learning and
ne-tuning. In this work, we explore the use
of news data on a speci c topic in ne-tuning
opinion mining models learned from social
media data, such as Twitter. Particularly, we
investigate the in uence of biased news data on
models trained on Twitter data by
considering both the balanced and unbalanced cases.
Results demonstrate that tuning with biased
news data of di erent properties changes the
classi cation accuracy up to 9:5%. The
experimental studies reveal that the
characteristics of the text of the tuning dataset, such
as bias, vocabulary diversity and writing style,
are essential for the nal classi cation results,
while the size of the data is less
consequential. Moreover, a state-of-the-art algorithm
is not robust on unbalanced twitter dataset,
and it exaggerates when predicting the most
frequent label.</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>In recent years, social media platforms have
become leading channels for the exchange of
knowledge, debates, and product or opinion advertising
[PP10, WD07, Gly18, SKB12]. Social scientists
routinely use data from social media platforms to
survey public opinion on speci c topics [Mos13, CSPR16,
HBK+17, BM18] and computer scientists use the data
to improve the performance of state-of-the-art
natural language processing (NLP) algorithms [CXHW17,
ACCF16, GPCR18, ZWWL18].</p>
      <p>Social media data, while abundant, pose many
challenges in usage: 1) user demographics are rarely
available; 2) posts are short and sometimes hard to
understand without context, and 3) it is challenging to
label millions of posts manually in short time. One
may overcome the rst challenge by selecting only
information from users where demographic information
is available using multiple social platforms. However,
this may bias the data. In order to solve the other
two problems, we need systems that classify data into
di erent opinion classes with limited human
involvement.
Numerous algorithms have been proposed to cope
with large amounts of short text [ZZL15, ZQZ+16,
LQH16, XC16, CSBL16, YYD+16, MGB+18]. All
these algorithms are supervised in nature, and
therefore, require hundreds of thousands of labels in
order to achieve adequate performance levels. In the
last two years, algorithms such as CoVe [MBXS17],
ELMo [PNI+18], ULMFiT [HR18] and OpenAI GPT
learning rate, slanted triangular learning rates are used
for every layer to improve the accuracy of the model
[HR18]. First, the learning rate sharply linearly
increases so that the model can learn fast from the rst
examples. Once learning rate achieves the L, it slowly
linearly declines as shown in the top-right corner of
Figure 2.</p>
      <p>In the third step (Figure 2c), layers are trained
gradually. First, only the top layer is trained with
labeled data for one epoch while other layers are frozen.
In each new epoch, the next frozen layer from the top
is added to the training.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>Experiments are conducted using Twitter data on
USA midterm elections 2018 and news data from USA
elections 2016.</p>
      <p>Twitter data is collected by searching for posts
published between November 4th and 7th 2018 which
have one of the hashtags: "#vote", "#trump",
"#election", "#midtermelection", "#democrats",
"#republicans" and "#2018midterms". In total, we
accrue 936,462 tweets. Most of the posts are retweets,
which appear multiple times in the corpus. After
retweets removal, 244,320 distinct posts remained, and
we pre-process their text by removing all characters,
except alphanumerics.</p>
      <p>Out of those posts, we label 1,526 examples with 0,
1 or 2. Label 0 is assigned to examples that support
or promote the left political spectrum or denounce the
right point of view. Label 1 is given to politically
neutral posts (e.g., posts that encouraged voting). Label
2 is assigned to examples that support or advertise
the right political spectrum or condemn the left point
of view. We discard 500 examples ( 25% of posts)
because they are unrelated to elections.</p>
      <p>News data is collected from six outlets that are
perceived to have di erent political partisanship,
ranging from the left-oriented to right-oriented outlets
based on media bias fact check website (Table 1).
Articles published between October 2015 and May 2017
that contain words "election", "ballot", "republican",
"GOP", or "democrat" are selected. The news
articles di er substantially in writing style, content
diversity, bias, number of articles and number of words
(Table 1). As with the tweets, news articles do not
always discuss the U.S. elections. Sometimes, they
debate Brexit or elections in France and other
countries worldwide. In pre-processing, we remove all
nonalphanumeric characters from news articles.</p>
      <p>Experiments settings. We use the pre-trained
WT103 token-vectors in the rst ULMFiT step.
WT103 has 103 million tokens from Wikipedia texts
for training, 217K tokens for validation and 245K
tokens for testing [MXBS16]. Our system is trained
using the architecture in Figure 2a. The vocabulary has
267K unique tokens. In this paper word and token
have interchangeable meanings.</p>
      <p>For the ne-tuning step, we explore ten di erent
settings: 1) "all news" text with the data from all
outlets + tweets text; 2) only the tweets; 3) text
from "left-biased" outlets + tweets text; 4) text from
"right-biased" outlets + tweets text. Remaining six
experiments contain text from one outlet and tweets
text. We randomly permute examples in a ne-tuning
dataset before usage.</p>
      <p>In the third step, experiments test two settings of
labeled Twitter data. Mix 1 (balanced mix) contains
380 examples with label 0 (left), 323 examples with
label 1 (neutral) and 323 examples with label 2 (right).
Mix 2 (unbalanced mix) contains 380 examples with
label 0 (left), 823 examples with label 1 (neutral) and
323 examples with label 2 (right). We randomly split
labeled data into three disjoint parts: test (200
examples), validation (200 examples) and training (626
examples in Mix 1 and 1126 examples in Mix 2). Each
experiment is repeated four times and accuracy mean,
and the standard deviation is reported for each of the
ten settings.</p>
      <p>We do not clean Twitter, and news data of
nonrelevant examples in order to emulate the real-world
situation. The data retrieval process is
intentionally simple to mirror the information extraction
process often used in research papers [Mos13, CSPR16,
HBK+17, BM18]. Those experiments test the
robustness of the model to the bias and noise in data and
robustness to the unbalanced classes.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and discussion</title>
      <p>We repeat each experiment four times, and we report
the accuracy mean and standard deviation in Table
2. High standard deviation (1:2 5:3%) indicates the
model's sensitivity to the order of examples in the
netuning data and a need for more labeled examples.</p>
      <p>Results provide evidence that the model is not
robust to unbalanced datasets. When Mix 1 and Mix 2
results are compared, the model always achieved
better results for Mix 2 (Table 2) which has 54% of
neutral labels as compared to 31.5% of neutral labels in
Mix 2
(380 : 823 : 323)
59:4 3:7%
66:6 2:5%
61:1 3:3%
63:0 3:2%
62:7 3:0%
60:7 1:4%
64:1 2:7%
64:2 1:8%
60:0 4:3%
61:9 3:3%
Mix 1. As evident from Figure 4, 80 90% of
predicted labels for Mix 2 are neutral. Therefore, better
results for Mix 2 are achieved because the algorithm
exaggerates the most frequent (neutral) label in the
imbalanced dataset (which contains 54% of examples
of that class).</p>
      <p>The classi cation accuracy di erence between Mix
1 and 2 is the largest (11.9%) when "left-biased news"
is used for ne-tuning. In this case, the accuracy on
both Mix 1 and Mix 2 decreases compared to when
"No news" is present. However, outlet bias has more
in uence on the accuracy of Mix 1.</p>
      <p>Figure 3 reveals that using "all news" data for
netuning achieves the best balance among predicted
labels for Mix 1. However, almost half of predicted labels
are wrong, so accuracy is low.</p>
      <p>Labeled Twitter data demonstrate diversity among
posts with label "left". They often talk only about one
particular issue and have fewer hashtags that support
the left political spectrum. Additionally, the diversity
of people and entities mentioned is more prominent in
the posts labeled as "left" than those labeled "right"
(which mainly mention president Trump). Hence, the
best performance for Mix 1 is achieved when
netuning with "CNN" data because the model is trained
to focus more on left-relevant contexts.</p>
      <p>The next best results for Mix 1 are achieved when
ne-tuning with news articles from The Wall Street
Journal because its articles often discuss both sides in
detail (sometimes even in the same sentence). Hence,
when the model is trained with data from this
outlet, it understands relevant phrases and predicts "left"
and "right" labels with higher accuracy. On the other
hand, "The Wall Street Journal" ne-tuned
experiment predicts much more often "right" label for
"leftlabeled" example than the other experiments.</p>
      <p>The confusion matrices created for each experiment
and Figure 4 reveal that the algorithm recognizes the
right label easier than the left label in Mix 2. A better
understanding of the right label can be explained with
the di erent writing style of left-labeled tweets, which
re ects a more diverse set of topics and entities as
discussed above. The best accuracy score for Mix 2
is achieved when "no news" data is used for the
netuning process. Most of the labels are neutral, and
news data is mainly left or right oriented/biased, so it
in uences the accuracy negatively.</p>
      <p>As hypothesized, results demonstrate that
netuning with biased news datasets can in uence
accuracy in contrasting ways. Di erent in uence of
biased news is particularly visible in the results of Mix
1 where the di erence between the best and the worst
accuracy for di erent ne-tuning settings is 9:5%. In
Mix 2 this di erence is also notable, 7:2%. In uence of
the bias is not uniform. While ne-tuning with
"leftbiased news" gives the worst result for Mix 1, its
performance for Mix 2 is average when compared to other
experiments. On the other hand, ne-tuning with "all
news" gives the worst results for Mix 2 and average
results for Mix 1.</p>
      <p>The size of the ne-tuning data does not seem to
inuence the results. "Washington Post" has the largest
amount of words, but it achieves average results in
both mixes. "CNN" is the smallest dataset, but it
achieves the best result for Mix 1. It is interesting
to notice that "all news" achieves worse results than
"no news" ne-tuning for both Mix 1 and Mix 2, even
though in literature, training with more data often
contributes to better results. This result suggests that
the content (bias) of the ne-tuning dataset is more
important than its size.</p>
      <p>Accuracy behavior in many experiments requires
further analysis in order to better understand the
inuence of ne-tuning text characteristics on the
performance. Additionally, the e ect of non-relevant text
on the accuracy should be further tested since its
frequency is high in both news and Twitter data. Since
results clearly show that this model is not robust on
bias and noise, other novel methods should be tested
similarly. It is essential to create unbalanced and
biased datasets for ne-tuning and testing of the future
models to create robust methods that would be
benecial to the real-world applications.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this work we have shown that bias, noise and text
properties need to be accounted for when constructing
data for ne-tuning language models. Text size does
not seem to be an important dimension. We performed
experiments with data collected from Twitter and six
news outlets using ULMFiT language model. Results
show that the algorithm is not robust to noise in data,
to bias in the ne-tuning dataset, or to the dataset
imbalance.</p>
      <p>While conducted experiments show weaknesses of
the existing system, further work is needed to
understand better the relationship between properties of
ne-tuning data and speci c tasks. Additionally,
better models are required that are more robust to bias
and noise in order to be able to solve challenging
realworld problems.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This research was supported in part by the NSF grants
IIS-1842183.
[BM18]
[CSBL16]
[CSPR16]</p>
      <sec id="sec-6-1">
        <title>Marco Bastos and Parametrizing brexit: ter political space constituencies.</title>
        <p>munication &amp;
2018.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Dan Mercea. mapping twitto parliamentary</title>
        <p>Information,
ComSociety, 21(7):921{939,</p>
      </sec>
      <sec id="sec-6-3">
        <title>Alexis Conneau, Holger Schwenk, Loc</title>
        <p>Barrault, and Yann Lecun. Very deep
convolutional networks for text classi cation.
arXiv preprint arXiv:1606.01781, 2016.</p>
      </sec>
      <sec id="sec-6-4">
        <title>Fabio Celli, Evgeny Stepanov, Massimo</title>
        <p>Poesio, and Giuseppe Riccardi.
Predicting brexit: Classifying agreement is better
than sentiment and pollsters. In
Proceedings of the Workshop on Computational
Modeling of Peoples Opinions,
Personality, and Emotions in Social Media
(PEO</p>
        <p>PLES), pages 110{118, 2016.
[CXHW17] Tao Chen, Ruifeng Xu, Yulan He, and
Xuan Wang. Improving sentiment
analysis via sentence type classi cation using
bilstm-crf and cnn. Expert Systems with</p>
        <p>Applications, 72:221{230, 2017.
[Gly18]</p>
        <p>Carroll J Glynn. Public opinion.
Routledge, 2018.
[GPCR18] Aitor Garc a-Pablos, Montse Cuadros,
and German Rigau. W2vlda: almost
unsupervised system for aspect based
sentiment analysis. Expert Systems with
Applications, 91:127{137, 2018.
[HBK+17] Philip N Howard, Gillian Bolsover, Bence
Kollanyi, Samantha Bradshaw, and
LisaMaria Neudert. Junk news and bots
during the us election: What were
michigan voters sharing over twitter.
Computational Propaganda Research Project,
Oxford Internet Institute, Data Memo, 1,
2017.
[LQH16]
[RNSS18]
[SKB12]
[WD07]
[XC16]
amazonaws.
com/openai-assets/researchcovers/languageunsupervised/language
understanding paper. pdf, 2018.</p>
      </sec>
      <sec id="sec-6-5">
        <title>Pawel Sobkowicz, Michael Kaschesky, and</title>
        <p>Guillaume Bouchard. Opinion mining
in social media: Modeling, simulating,
and forecasting political opinions in the
web. Government Information Quarterly,
29(4):470{479, 2012.</p>
      </sec>
      <sec id="sec-6-6">
        <title>Duncan J Watts and Peter Sheridan Dodds. In uentials, networks, and public opinion formation. Journal of consumer research, 34(4):441{458, 2007.</title>
      </sec>
      <sec id="sec-6-7">
        <title>Yijun Xiao and Kyunghyun Cho. E cient</title>
        <p>character-level document classi cation by
combining convolution and recurrent
layers. arXiv preprint arXiv:1602.00367,
2016.
[ZQZ+16] Peng Zhou, Zhenyu Qi, Suncong Zheng,
Jiaming Xu, Hongyun Bao, and Bo Xu.</p>
        <p>Text classi cation improved by
integrating bidirectional lstm with
twodimensional max pooling. arXiv preprint
arXiv:1611.06639, 2016.
[ZWWL18] Shunxiang Zhang, Zhongliang Wei, Yin
Wang, and Tao Liao. Sentiment
analysis of chinese micro-blog text based on
extended sentiment dictionary. Future
Generation Computer Systems, 81:395{403,
2018.
[ZZL15]</p>
      </sec>
      <sec id="sec-6-8">
        <title>Xiang Zhang, Junbo Zhao, and Yann Le</title>
        <p>Cun. Character-level convolutional
networks for text classi cation. In Advances
in neural information processing systems,
pages 649{657, 2015.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [ACCF16]
          <string-name>
            <given-names>Orestes</given-names>
            <surname>Appel</surname>
          </string-name>
          , Francisco Chiclana, Jenny Carter, and
          <string-name>
            <given-names>Hamido</given-names>
            <surname>Fujita</surname>
          </string-name>
          .
          <article-title>A hybrid approach to the sentiment analysis problem at the sentence level</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>108</volume>
          :
          <fpage>110</fpage>
          {
          <fpage>124</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>Universal language model ne-tuning for text classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1801.06146</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Pengfei</given-names>
            <surname>Liu</surname>
          </string-name>
          , Xipeng Qiu, and
          <string-name>
            <given-names>Xuanjing</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <article-title>Recurrent neural network for text classi cation with multi-task learning</article-title>
          .
          <source>arXiv preprint arXiv:1605.05101</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [MBXS17]
          <string-name>
            <surname>Bryan</surname>
            <given-names>McCann</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>James</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , Caiming Xiong, and Richard Socher.
          <article-title>Learned in translation: Contextualized word vectors</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>6294</fpage>
          {
          <fpage>6305</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [MGB+18]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Edouard Grave, Piotr Bojanowski,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Puhrsch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          .
          <article-title>Advances in pre-training distributed word representations</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Merity</surname>
          </string-name>
          , Nitish Shirish Keskar, and Richard Socher.
          <article-title>Regularizing and optimizing lstm language models</article-title>
          .
          <source>arXiv preprint arXiv:1708.02182</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Mohamed M Mostafa.</surname>
          </string-name>
          <article-title>More than words: Social networks text mining for consumer brand sentiments</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>40</volume>
          (
          <issue>10</issue>
          ):
          <volume>4241</volume>
          {
          <fpage>4251</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [MXBS16]
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Merity</surname>
          </string-name>
          , Caiming Xiong, James Bradbury, and Richard Socher.
          <article-title>Pointer sentinel mixture models</article-title>
          .
          <source>arXiv preprint arXiv:1609.07843</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>arXiv preprint arXiv:1802.05365</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>Twitter as a corpus for sentiment analysis and opinion mining</article-title>
          .
          <source>In LREc</source>
          , volume
          <volume>10</volume>
          , pages
          <fpage>1320</fpage>
          {
          <fpage>1326</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Alec</given-names>
            <surname>Radford</surname>
          </string-name>
          , Karthik Narasimhan, Tim Salimans, and
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <article-title>Improving language understanding by generative pre-training</article-title>
          .
          <source>URL https://s3-us-west-2.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[MKS17] [Mos13] [PNI+18] [PP10]</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>