<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Perils of Classifying Political Orientation From Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hao Yan</string-name>
          <email>haoyan@wustl.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allen Lavoie?</string-name>
          <email>allenlavoie@wustl.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanmay Das</string-name>
          <email>sanmay@wustl.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Washington University in St. Louis</institution>
          ,
          <addr-line>St. Louis</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Political communication often takes complex linguistic forms. Understanding political ideology from text is an important methodological task in studying political interactions between people in both new and traditional media. Therefore, there has been a spate of recent research that either relies on, or proposes new methodology for, the classification of political ideology from text data. In this paper, we study the effectiveness of these techniques for classifying ideology in the context of US politics. We construct three different datasets of conservative and liberal English texts from (1) the congressional record, (2) prominent conservative and liberal media websites, and (3) conservative and liberal wikis, and apply text classification algorithms with a domain adaptation technique. Our results are surprisingly negative. We find that the cross-domain learning performance, benchmarking the ability to generalize from one of these datasets to another, is poor, even though the algorithms perform very well in within-dataset cross-validation tests. We provide evidence that the poor performance is due to differences in the concepts that generate the true labels across datasets, rather than to a failure of domain adaptation methods. Our results suggest the need for extreme caution in interpreting the results of machine learning methodologies for classification of political text across domains. The one exception to our strongly negative results is that the classification methods show some ability to generalize from the congressional record to media websites. We show that this is likely because of the temporal movement of the use of specific phrases from politicians to the media.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Political discourse is a fundamental aspect of government across the world, especially
so in democratic institutions. In the US alone, billions of dollars are spent annually on
political lobbying and advertising, and language is carefully crafted to influence the
public or lawmakers [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. Matthew Gentzkow won the John Bates Clark Medal in
economics in 2014 in part for his contributions to understanding the drivers of
media “slant.” With the increasing prevalence of social media, where activity patterns are
correlated with political ideologies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], companies are also striving to identify users’
ideologies based on their comments on political issues, so that they can recommend
specific news and advertisements to them.
      </p>
      <p>
        The manner in which political speech is crafted and words are used creates
difficulties applying standard methods. Political ideology classification is a difficult task
? Now at Google Brain
even for people – only those who have substantial experience in politics can correctly
classify the ideology behind given articles or sentences. In many political ideology
labeling tasks, it is even more essential than in tasks that could be thought of as similar
(e.g. labeling images, or identifying positive or negative sentiment in text) to ensure that
labelers are qualified before using the labels they generate [
        <xref ref-type="bibr" rid="ref21 ref5">5, 21</xref>
        ].
      </p>
      <p>
        One of the reasons why classification of political texts for inexperienced people is
hard is because different sides of the political spectrum use slightly different
terminology for concepts that are semantically the same. For example, in the US debate over
privatizing social security, liberals typically used the phrase “private accounts” whereas
conservatives preferred “personal accounts” [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Nevertheless, it is well-recognized
that “dictionary based” methods for classifying political text have trouble generalizing
across different domains of text [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        Many methods based on machine learning techniques have also been proposed for
the problem of classifying political ideology from text [
        <xref ref-type="bibr" rid="ref1 ref21 ref23">1, 21, 23</xref>
        ]. The training and
testing process typically follows the standard validation rules: first split the dataset into
a training set and a test set, then propose an algorithm and train a classification model
based on the training set and finally test on the test set. These methods have been
achieving increasingly impressive results, and so it is natural to assume that classifiers trained
to recognize political ideology on labeled data from one type of text can be applied
to different types of text, as has been common in the social science literature (e.g.
Gentzkow and Shapiro using phrases from the Congressional Record to measure the
slant of news media [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], or Groseclose and Milyo using citations of different think
tanks by politicians to also measure media bias [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]). However, these papers are
classifying the bias of entire outlets (for example, The New York Times or The Wall Street
Journal) rather than individual pieces of writing, like articles. Such generalization
ability is not obvious in the context of machine learning methods working with smaller
portions of text, and must be put to the test.
      </p>
      <p>The main question we ask in this paper is whether the increasingly excellent
performance of machine learning models in cross-validation settings will generalize to the
task of classifying political ideology in text generated from a different source. For
example, can a political ideology classifier trained on text from the congressional record
successfully distinguish between liberal and conservative news articles? One immediate
problem we face in engaging this question is the absence of large datasets with political
ideology labels attached to individual pieces of writing. Therefore, we assemble three
datasets with very different types of political text and an easy way of attributing labels
to texts. The first is the congressional record, where texts can be labeled by the party
of the speaker. The second is a dataset of articles from two popular web-based
publications, Townhall.com, which features conservative columnists, and salon.com,
which features liberal writers. The third is a dataset of political articles taken from
Conservapedia (a conservative response to Wikipedia) and RationalWiki (a liberal response
to Conservapedia). In each of these cases there is a natural label associated with each
article, and it is relatively uncontroversial that the labels align with common notions of
liberal and conservative. We show that standard classification techniques can achieve
high performance in distinguishing liberal and conservative pieces of writing in
crossvalidation experiments on these datasets.</p>
      <p>
        It is tempting to assume that there is enough shared language across datasets that
one can generalize from one to the other for new tasks, for example, for detecting bias
in Wikipedia editors, or the political orientation of op-ed columnists. However, is it
really reasonable to extrapolate from any of these datasets to others? As a motivating
example, we show that the results of training bag-of-bigram linear classifiers using the
three different datasets above and then using them to identify the political biases of
Wikipedia administrators leads to wildly inconsistent results, with virtually no
correlation between the partisanship rankings of the administrators based on the three
different training sets. More generally, we show that, with one exception, the unaltered
cross-domain performance of different classifiers on these datasets is abysmal, and there
is only marginal benefit from applying a state-of-the-art domain adaptation technique
(marginalized stacked denoising autoencoders [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]). The exception is in using data from
the congressional record to predict whether articles are from Salon or Townhall,
consistent with Gentzkow and Shapiro’s results on media bias. A temporal analysis suggests
that this is because phrases move in a rapid and predictable way from the congressional
record to the news media. However, even in this domain, we provide evidence that the
underlying concepts (Salon vs. Townhall compared with Democrat vs. Republican) are
significantly different: adding additional labeled data from one domain actively hurts
performance on the other. Our results are robust to using regressions on measures of
political ideology (DW-Nominate scores [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]) rather than simple classifications of
partisanship. Our overall results suggest that we should proceed with extreme caution in
using machine learning (or phrase-counting) approaches for classifying political text,
especially in situations where we are generalizing from one type of political speech to
another.
1.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work.</title>
      <p>
        While our methods and results are general, we focus in this paper on political
ideology in the US context, since there is already a rich literature on the topic, as well as
abundant data. Political ideology in U.S. media has been well studied in economics and
other social sciences. Groseclose et al., [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] calculate and compare the number of times
that think tanks and policy groups were cited by mainstream media and congresspeople.
Gentzkow et al., [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] generate a partisan phrase list based on the Congressional Record
and compute an index of partisanship for U.S. newspapers based on the frequency of
these partisan phrases. Budak et al., [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] use Amazon Mechanical Turk to manually rate
articles from major media outlets. They use machine learning methods (logistic
regression and SVMs) to identify whether articles are political news, but then use human
workers to identify political ideology in order to determine media bias. Ho et al., [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
examine editorials from major newspapers regarding U.S. Supreme Court cases and
apply the statistical model proposed by Clinton et al., [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. All of the above research gives
us quantitative political slant measurements of U.S. mainstream media outlets.
However, these political ideology classification results are corpus-level rather than article
level or sentence level.
      </p>
      <p>
        The machine learning community has focused more on the learning techniques
themselves. Gerrish et al., [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] propose several learning models to predict voting
patterns. They evaluate their model via cross-validation on legislative data. Iyyer et al., [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
apply recursive neural networks in political ideology classification. They use
Convote [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and the Ideological Books Corpus [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. They present cross-validation results
and do not analyze performance on different types of data. Ahmed et al., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] propose
an LDA-based topic model to estimate political ideology. They treat the generation of
words as an interaction between topic and ideology. They describe an experiment where
they train their model based on four blogs and test on two new blogs. However, political
blogs are considerably less diverse than our datasets; since the articles in our datasets
are generated in completely different ways (speeches, crowdsourcing and editorials).
The results in this paper constitute a more general test of cross-domain political
ideology learning.
      </p>
      <p>
        Cross-domain text classification methods are an active area of research. Glorot et
al., [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] propose an algorithm based on stacked denoising autoencoders (SDA) to learn
domain-invariant feature representations. Chen et al., [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] come up with a marginalized
closed-form solution, mSDA. Recently, Ganin et al., [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] have proposed a promising
“Y” structure end-to-end domain adversarial learning network, which can be applied in
multiple cross-domain learning tasks.
      </p>
      <p>
        Cohen et al., [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] investigate the classification of political leaning across three
different groups (based on activity level) of Twitter users. Without any domain adaptation
methodology, they show that cross-domain classification accuracy declines significantly
compared with in-domain accuracy. Our work provides a view across much more
diverse data sources than just social media, and engages the question of domain adaptation
more substantively.
2
2.1
      </p>
      <sec id="sec-2-1">
        <title>Data and Methods</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>
        Mainstream newspapers and websites have been widely used in political ideology
research [
        <xref ref-type="bibr" rid="ref13 ref3 ref5">3, 5, 13</xref>
        ]. However, these datasets contain many non-political articles, and the
political articles in news datasets are typically non-partisan [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Therefore, we carefully
construct three datasets that we expect to be partisan: (1) The Congressional Record,
containing statements by members of the Republican and Democratic parties in the US
congress; (2) News media articles from Salon (a left-leaning website) &amp; Townhall (a
right-leaning one); and (3) Articles related to American politics from two collectively
constructed “new media” websites, Conservapedia (conservative) &amp; RationalWiki
(liberal). Details of the construction process and the resulting corpora are in the appendix.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <p>
        Text Preprocessing We perform some preprocessing on all the datasets to extract
content rather than references and metadata, and also standardize the text by lowercasing,
stemming, removing stopwords and other extremely common and venue-specific words.
Logistic Regression Models Logistic regression is a standard and useful technique
for text classification. We extract bigrams from the text and Term Frequency-Inverse
Document Frequency weighting to construct the feature representation for logistic
regression to use (and denote the overall method TF-IDFLR in what follows). We use the
implementation provided in the scikit-learn machine learning package [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>
        Marginalized Stacked Denoising Autoencoders for domain adaptation
Marginalized Stacked Denoising Autoencoders (mSDA) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] are a state-of-the-art cross-domain
text classification method [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Given bag-of-words input of text from two different
domains, mSDA provides a closed-form representation of the input, and is faster than
the original Stacked Denoising Autoencoder (SDA) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] without loss of classification
accuracy. We use TF-IDF bag-of-bigrams vectors as the input to mSDA, the original
mSDA Python package1 for the implementation of mSDA in combination with the
logistic regressions described above in our domain adaptation experiments.
Semi-Supervised Recursive Autoencoders Recently, there have been rapid advances
in text sentiment and ideology classification based on recursive neural networks. Most
of this work is based on sentence or phrase level classification. Some of these methods
use fully labeled [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] or partially labeled [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] parsed sentence trees, and some need
large numbers of parameters [
        <xref ref-type="bibr" rid="ref27 ref29">27, 29</xref>
        ]. Since we have large datasets available to use,
we use semi-supervised recursive autoencoders (RAE) [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], which do not need parse
trees, labels for all nodes in the parse trees, or a large number of parameters. We use the
MATLAB package distributed by Socher et al., [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]2. We do not transform the words
down to their linguistic roots when we apply the RAE method since we need to use a
word dictionary.
3
3.1
      </p>
      <sec id="sec-4-1">
        <title>Results</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Cross-domain consistency</title>
      <p>The first question is whether training on different domains yields consistent results in
classifying political ideology. We evaluate this on a motivating task that is exactly the
type of task that one may wish to use these types of tools for, determining ideological
bias among Wikipedia administrators.</p>
      <p>For each of the 500 most active Wikipedia administrators, we concatenate all the
strings they have added to pages on Wikipedia related to U.S. politics and classify the
resulting “body of work” of that administrator using the three different training sets (the
Congressional Record is #1, Salon/Townhall is #2, and RationalWiki/Conservapedia is
#3). Each classifier produces a ranking of these 500 administrators. Shockingly we find
that these rankings have virtually no correlation with each other (see Table 1).</p>
      <p>Somewhat more anecdotally, we can also look at the ranks of some users from each
method. We select the three most liberal users according to each of the three classifiers
and find their positions in the other two lists. The results are in Table 2 and again
demonstrate how diverse the rankings can be based on the training sets.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Consistency across time</title>
      <p>1 http://www.cse.wustl.edu/˜kilian/code/files/mSDA.zip
2 http://nlp.stanford.edu/˜socherr/codeDataMoviesEMNLP.zip</p>
      <p>The words used to describe politics change across 1
tpiomliet,icaasl adroticthlees ttohpaitcasreofdiismtapnotritnanticme.eTfrhoemrefeoarceh, 000...897
other will be less similar than those written dur- 0.6
iimsnegat hthoseidgssnaibfimyceafonpcteuriissoisndug.e WofnoertnhtoehweSasloltougndisytainwcdhreeTgtohrweersnshtihaoilnsl AUC0000....3425
dataset. We use 2006 Salon and Townhall articles 0.1
as a training set and future years (from 2007 to 20007 2008 2009 2010Year2011 2012 2013 2014
2014) as separate test sets.</p>
      <p>Figure 1 shows the AUC across time. The Fig. 1: Salon &amp; Townhall
yearAUC for 2007 is 0.872, which means that the Sa- based timeline test. The training
lon &amp; Townhall articles in 2006 and 2007 are sim- set is 2006 Salon &amp; Townhall data.
ilar enough for successful generalization of the The test sets are individual year
ideology classifier from one to the other. How- data from 2007 to 2014, also from
ever, the prediction accuracy goes down signifi- Salon &amp; Townhall.
cantly as the dates of the test set become further
out in the future, as the nature of the discourse
changes. It is now clear that our classification methods have generalization problems
both across domains and across time.
Now we turn to a more comprehensive analysis. We examine the performance of
several different methods across the three labeled datasets. We study linear classifiers and
recursive autoencoders as described above, as well as the mSDA method for domain
adaptation. In order to account for the effects of time-varying language use
demonstrated above, we restrict our methods to train and test only on data from the same year,
and then aggregate results across years.</p>
      <p>Training Set</p>
      <p>Congressional</p>
      <p>Record</p>
      <p>Table 3 shows the average AUC for each
group of experiments. The within-domain
crossvalidation results (on the diagonal) are excellent
for both the linear classifier and the RAE.
However, the naive cross-domain generalization
results are uniformly terrible, often barely above
chance. While we could hope that using a
sophisticated domain-adaptation technique like mSDA
would help, the results are disappointing: in only
one cross-domain task (generalizing from the
Congressional Record to Salon and Townhall)
does it help to achieve a reasonable level of
accuracy. The AUC score gaps between
crossvalidation and domain adaptation results indicate
that, even with a state-of-the-art domain
adaptation algorithm, cross-text domain political
ideology identification is not, at this point, able to give
reliable results. It is of note that the best
performance is in generalizing from the congressional
record to a media dataset (Salon/Townhall)
because it adds weight to the existing line of
research starting from Gentzkow and Shapiro on
how language flows from politicians to the media.
(Implementation details and parameter choices
for Sections 3.1-3.3 can be found in the appendix)
0.90
0.85
C0.80
U
A
0.75
0.70
0.65
mSDA (with Congressional Record)
TF-IDFLR (with Congressional Record)
Reduced Ngram (with Congressional Record)</p>
      <p>Cross Validation (no Congressional Record)
0.R0atio0.o1f Sa0l.o2n/To0.w3nha0l.l4Dat0a.5as T0r.a6inin0g.7Data0.8
Fig. 2: AUC on Salon/Townhall as
a function of the proportion of the
labeled (Salon/Townhall) dataset
used in training. The results show
that including labeled data from
the Congressional Record never
helps and actively hurts
classification accuracy in almost all
settings, and that restricting features
to ngrams with sufficient support
in both datasets does not help
either.
3.4</p>
    </sec>
    <sec id="sec-7">
      <title>Failure of domain adaptation, or distinct concepts?</title>
      <p>There are two plausible hypotheses that could explain these negative results. H1: The
domain adaptation algorithm algorithm is failing (probably because it is easy to overfit
labeled data from any of the specific domains), or H2: The specific concepts we are
trying to learn are actually different or inconsistent across the different datasets. We
perform several experiments to try and provide evidence to distinguish between these
hypotheses. First, we may be able to reduce overfitting by restricting the features to
ngrams that have sufficient support (operationally, at least 5 appearances) in both sets
of data (this reduces the dimensionality of the space and would lead to a greater
likelihood of the “true” liberal/conservative concept being found if there were many accurate
hypotheses that could work in any individual dataset). Second, we can examine
performance as we include more and more labeled data from the target domain in the training
set. In the limit, if the concepts are consistent, we would not expect to see any
degradation in (cross-validation) performance on the source domain from including labeled
data from the target domain in training.</p>
      <p>We focus on the Salon/Townhall and Congressional Record data sets here since
they are the most promising for the possibility of domain adaptation. We combine part
of the Salon/Townhall data with Congressional Record as training set. Then we use
the rest of the Salon/Townhall data set as the test set, increasing the percentage of the
Salon/Townhall dataset used in training from 0% to 80%, and compare with
crossvalidation performance on just the Salon/Townhall dataset.</p>
      <p>Figure 2 shows that including labeled data from the Congressional Record never
helps and, once we have at least 10% of labels, actively hurts classification accuracy
on the Salon/Townhall dataset. Restricting to bigrams that appear in both datasets at
least 5 times further degrades the performance. This demonstrates quite clearly that
the problem is not overfitting a specific dataset when there are many correct concepts
available, it is that the concept of being from Salon or Townhall is significantly different
than the concept of being from a Democratic or Republican speech. Therefore, the hope
of successful domain-agnostic classification of political orientation based on text data
is significantly diminished.
3.5</p>
    </sec>
    <sec id="sec-8">
      <title>Temporal movement of topics</title>
      <p>The silver lining so far is that there is at least some
ability to predict the political orientation of
webbased news media based on the congressional 45
record. We can further investigate this insight 40
cagLonaerldsldkeidocnetvgemedncoebnewytsstareaelx.vt,eae[mn2tht2isne]ibniuengttviwletihsetyteeingoqatfutheeetdshmtetihoadeniantttsiaemtrmeweapelmoarghamalrlveeye--. iftrrxsopebeenEmm2213305505
dia and blogs. We ask a similar question – who uN10
discusses “new” political topics in the first place 5
– coInngroersdseorrttoheanmsewdeira?this question, we exam- 00 1 2 3 Days 4 5 6 7
ine mutual trigrams in the Congressional Record Fig. 3: Distribution of median
and Salon&amp;Townhall datasets. We find all new tri- value of time lag results in each
grams in any given year (those which did not ap- experiment
pear in the previous year and appeared at least
twice in the media data and five times in the
congressional record in the given year and the next
one), and then construct the time lags between first appearance in each of the two
datasets, excluding congressional recess days. Since the congressional record is much
larger, we subsample and repeat the experiment many times to get a distribution of time
lags.</p>
      <p>In each of these bootstrapped samples, there is a median time lag between the first
appearance of a phrase in the congressional record and its first appearance in the media
dataset. Figure 3 shows the distribution of these medians. The median is never negative,
and is on average 2 days, showing a definite tendency for phrases to travel from the
congressional record to the media rather than the other way round. The entire
distribution also shows a slight bias towards the media picking up on congressional topics of
discussion after the fact. These results help to explain the relative success of domain
adaptation from the congressional record to the media dataset.
4</p>
      <sec id="sec-8-1">
        <title>Conclusion</title>
        <p>
          Text analytics is becoming a central methodological tool in analyzing political
communication in many different contexts. It is obviously very valuable to have a good way of
measuring political ideology based on text. Our work sounds a cautionary note in this
regard by demonstrating the difficulty of classifying political text across different
contexts. We provide strong evidence that, in spite of the fact that writers or speech makers
in different domains often self-identify or can be relatively easily identified by humans
as being conservative or liberal, the concepts are distinct enough across datasets (even
in just the US political context!) that generalization is extremely difficult. We note that,
while we have presented our results in the context of classification, we get identical
results when using measures of political ideology on a real-valued spectrum (the standard
DW-Nominate score [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]) as the target of a regression task (this is only feasible for the
congressional record, since the scores of congresspeople can be obtained as a function
of their voting record). Our results demonstrate the need for extreme caution in the
application of machine learning techniques to classifying political ideologies, especially
when such efforts are made across domains.
A
A.1
        </p>
      </sec>
      <sec id="sec-8-2">
        <title>Datasets</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Congressional Record.</title>
      <p>The U.S. Congressional Record preserves the activities of the House and Senate,
including every debate, bill, and announcement. We use the party affiliation of the speaker
(Democrat or Republican) as an indication of ideology (liberal or conservative). We
retrieve the floor proceedings of both the Senate and House from 2005 to 2014. We
separate the proceedings into segments with a single speaker. For each of these segments,
we extract the speaker and their party affiliation (Democrat, Republican or
independent)In order to focus on partisan language, we excluded speech from independents,
and from clerks and presiding officers.</p>
      <p>A.2</p>
    </sec>
    <sec id="sec-10">
      <title>Salon and Townhall.</title>
      <p>We collect articles tagged with “politics” from Salon, a website with a progressive/liberal
ideology, and all articles from Townhall, which mainly publishes reports about U.S.
political events and political commentary from a conservative viewpoint.
A.3</p>
    </sec>
    <sec id="sec-11">
      <title>Conservapedia and RationalWiki.</title>
      <p>Conservapedia (http://www.conservapedia.com/) is a wiki encyclopedia project
website. Conservapedia strives for a conservative point of view, created as a reaction
to what was seen as a liberal point of view from Wikipedia. RationalWiki (http:
//rationalwiki.org/) is also a wiki encyclopedia project website, which was,
in turn, created as a liberal response to Conservapedia. RationalWiki and
Conservapedia are based on the MediaWiki system. Once a page is set up, other users can revise
it. For RationalWiki, we download pages ranking in the top 10000 in number of
revisions. We further select pages whose categories contain the following word stems: liber,
conserv, govern, tea party, politic, left-wing, right-wing, president, u.s. cabinet, united
states senat, united states house. Because the Conservapedia community has more
articles than RationalWiki, we download the top 40000 pages. We apply the same political
keywords list we use for RationalWiki. We always use the last revision of any page for
a given time period.</p>
      <p>
        Table 4 shows the counts of articles in the liberal and conservative parts of each of
the three datasets by year. Our datasets have the following properties that make them
useful for political ideology learning and evaluation in the context of U.S. politics:
– The content is selected to be relevant to U.S. politics.
– The content can predictably be labeled as conservative or liberal by a somewhat
knowledgeable human. While it is true that not all speeches by Democrats are
liberal, and not all articles on Townhall conservative, since these are subjectively
defined, this is nevertheless as clean a delineation as we can hope for.
– The creation times of items in the three datasets have substantial overlap;
We also motivate our task by attempting to classify bias on Wikipedia, an important
task [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Wikipedia is the largest encyclopedia project in the world and is widely used
in both natural language processing and political science studies [
        <xref ref-type="bibr" rid="ref24 ref4">4, 24</xref>
        ]. Wikipedia is
considered to have become nonpartisan as many users have contributed to political
entries [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. We focus on edits made by admins on political topics in Wikipedia. We
download the English Wikipedia dump from March 4, 2015. To focus on US politics,
we extract all articles (with full edit history) that belong to WikiProject United States3
and satisfy the same political keywords requirement that we use for RationalWiki,
yielding 4659 articles in total. We then collect all edits added or subtracted by each active
Wikipedia admin.
      </p>
      <p>B
B.1</p>
      <sec id="sec-11-1">
        <title>Details of Experimental Methodology</title>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>Cross-domain consistency</title>
      <p>For the Congressional Record and Salon/Townhall datasets, we use data from 2005
to 2014. For the RationalWiki/Conservapedia datasets, we use the data from 2014 as
capturing a recent snapshot. For this dataset only we use feature hashing to project
the bigram features into a lower dimensional non-sparse feature space. We set the
dimension of the hashed vector n f eatures = 20000, ngram range = (2; 2), and
decode error = ignore. We use a so-called “balanced” logistic regression classifier to
deal with the problem of class imbalance. All other parameters are the defaults in the
scikit-learn package for both feature hashing vectorizer and logistic regression
classifier.</p>
      <p>B.2</p>
    </sec>
    <sec id="sec-13">
      <title>Consistency across time</title>
      <p>We use the TF-IDFLR method for this experiment. For the vectorizer, we set min df =
5, ngram range = (2; 2) and decode error = ignore. For logistic regression
classifier, we set class weight = balanced to re-weight training samples. Other parameters
are set to the default values in the scikit-learn package.
3 https://en.wikipedia.org/wiki/Wikipedia:WikiProject_United_</p>
      <p>States
The linear classifier is the TF-IDFLR method described above. The RAE algorithm
trains embeddings using sentences subsampled from the data in order to balance
conservative and liberal training sentences, and then a logistic regression classifier is used
on top of the embeddings thus trained. The marginalized stacked denoising autoencoder,
which is expected to find features that convey domain-invariant political ideology
information, is run on TF-IDF bigram features before a logistic regression is applied on
top of that feature representation. We use five-fold cross validation when the training
and testing sets are the same.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xing</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          :
          <article-title>Staying informed: Supervised and semi-supervised multi-view topical analysis of ideological perspective</article-title>
          .
          <source>In: Proc. EMNLP</source>
          . pp.
          <fpage>1140</fpage>
          -
          <lpage>1150</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bakshy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Messing</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adamic</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          :
          <article-title>Exposure to ideologically diverse news and opinion on Facebook</article-title>
          .
          <source>Science</source>
          <volume>348</volume>
          (
          <issue>6239</issue>
          ),
          <fpage>1130</fpage>
          -
          <lpage>1132</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Baum</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groeling</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>New media and the polarization of American political discourse</article-title>
          .
          <source>Polit. Comm</source>
          .
          <volume>25</volume>
          (
          <issue>4</issue>
          ),
          <fpage>345</fpage>
          -
          <lpage>365</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Wikipedia as a data source for political scientists: Accuracy and completeness of coverage</article-title>
          .
          <source>PS: Polit. Sci. &amp; Politics</source>
          <volume>44</volume>
          (
          <issue>02</issue>
          ),
          <fpage>339</fpage>
          -
          <lpage>343</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Budak</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          :
          <article-title>Fair and balanced? Quantifying media bias through crowdsourced content analysis</article-title>
          .
          <source>Public Opin. Quarterly</source>
          <volume>80</volume>
          (
          <issue>S1</issue>
          ),
          <fpage>250</fpage>
          -
          <lpage>271</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sha</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Marginalized stacked denoising autoencoders for domain adaptation</article-title>
          .
          <source>In: Proc. ICML</source>
          . pp.
          <fpage>767</fpage>
          --
          <lpage>774</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Clinton</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rivers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The statistical analysis of roll call data</article-title>
          .
          <source>Am. Polit. Sci. Rev</source>
          .
          <volume>98</volume>
          (
          <issue>02</issue>
          ),
          <fpage>355</fpage>
          -
          <lpage>370</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruths</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Classifying political orientation on twitter: It's not easy</article-title>
          ! In
          <source>: Proc. ICWSM</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavoie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magdon-Ismail</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Manipulation among the arbiters of collective intelligence: How Wikipedia administrators mold public opinion</article-title>
          .
          <source>ACM Trans. on the Web</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <volume>24</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          :
          <fpage>25</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>DellaVigna</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The Fox News effect: Media bias and</article-title>
          <string-name>
            <given-names>voting. Q. J.</given-names>
            <surname>Econ</surname>
          </string-name>
          .
          <volume>122</volume>
          (
          <issue>3</issue>
          ),
          <fpage>1187</fpage>
          -
          <lpage>1234</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Entman</surname>
          </string-name>
          , R.M.
          <article-title>: How the media affect what people think: An information processing approach</article-title>
          .
          <source>J. Politics</source>
          <volume>51</volume>
          (
          <issue>02</issue>
          ),
          <fpage>347</fpage>
          -
          <lpage>370</lpage>
          (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ganin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ustinova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ajakan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Germain</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laviolette</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marchand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lempitsky</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Domain-adversarial training of neural networks</article-title>
          .
          <source>JMLR</source>
          <volume>17</volume>
          (
          <issue>59</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Gentzkow</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shapiro</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>What drives media slant? Evidence from U.S. daily newspapers</article-title>
          .
          <source>Econometrica</source>
          <volume>78</volume>
          (
          <issue>1</issue>
          ),
          <fpage>35</fpage>
          -
          <lpage>71</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Gerrish</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.:</given-names>
          </string-name>
          <article-title>Predicting legislative roll calls from text</article-title>
          .
          <source>In: Proc. ICML</source>
          . pp.
          <fpage>489</fpage>
          -
          <lpage>496</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Glorot</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Domain adaptation for large-scale sentiment classification: A deep learning approach</article-title>
          .
          <source>In: Proc. ICML</source>
          . pp.
          <fpage>513</fpage>
          -
          <lpage>520</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Greenstein</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>Is Wikipedia biased? Am. Econ. Rev</source>
          .
          <volume>102</volume>
          (
          <issue>3</issue>
          ),
          <fpage>343</fpage>
          -
          <lpage>348</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Grimmer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            ,
            <given-names>B.M.</given-names>
          </string-name>
          :
          <article-title>Text as data: The promise and pitfalls of automatic content analysis methods for political texts</article-title>
          .
          <source>Political Analysis</source>
          pp.
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Groseclose</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milyo</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A measure of media bias</article-title>
          .
          <source>Q. J. Econ</source>
          .
          <volume>120</volume>
          (
          <issue>4</issue>
          ),
          <fpage>1191</fpage>
          -
          <lpage>1237</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Acree</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Testing the Etch-a-Sketch hypothesis: A computational analysis of Mitt Romney's ideological makeover during the 2012 primary vs</article-title>
          .
          <article-title>general elections</article-title>
          .
          <source>In: APSA Annual Meeting</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quinn</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          :
          <article-title>Measuring explicit political positions of media</article-title>
          .
          <source>Q. J. Polit. Sci. 3</source>
          (
          <issue>4</issue>
          ),
          <fpage>353</fpage>
          -
          <lpage>377</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Iyyer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enns</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boyd-Graber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Resnik</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Political ideology detection using recursive neural networks</article-title>
          .
          <source>In: ACL</source>
          . pp.
          <fpage>1113</fpage>
          -
          <lpage>1122</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Backstrom</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinberg</surname>
          </string-name>
          , J.:
          <article-title>Meme-tracking and the dynamics of the news cycle</article-title>
          .
          <source>In: Proc. KDD</source>
          . pp.
          <fpage>497</fpage>
          -
          <lpage>506</lpage>
          . ACM (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xing</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hauptmann</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>A joint topic and perspective model for ideological discourse</article-title>
          .
          <source>In: Proc. ECML-PKDD</source>
          . pp.
          <fpage>17</fpage>
          -
          <lpage>32</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>In: Proc. ICLR</source>
          (
          <year>2013</year>
          ), http://arxiv.org/abs/1301.3781
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in</article-title>
          <source>Python. JMLR 12</source>
          ,
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          (
          <year>October 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Poole</surname>
            ,
            <given-names>K.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
          </string-name>
          , H.:
          <article-title>Congress: A political-economic history of roll call voting</article-title>
          . Oxford University Press (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huval</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          :
          <article-title>Semantic compositionality through recursive matrix-vector spaces</article-title>
          .
          <source>In: Proc. EMNLP</source>
          . pp.
          <fpage>1201</fpage>
          -
          <lpage>1211</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Semi-supervised recursive autoencoders for predicting sentiment distributions</article-title>
          .
          <source>In: Proc. EMNLP</source>
          . pp.
          <fpage>151</fpage>
          -
          <lpage>161</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perelygin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chuang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potts</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Recursive deep models for semantic compositionality over a sentiment treebank</article-title>
          .
          <source>In: Proc. EMNLP</source>
          . pp.
          <fpage>1631</fpage>
          -
          <lpage>1642</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Get out the vote: Determining support or opposition from congressional floor-debate transcripts</article-title>
          .
          <source>In: Proc. EMNLP</source>
          . pp.
          <fpage>327</fpage>
          -
          <lpage>335</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>