<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>American journal of detection in Spanish texts</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.26342/2022-69-23</article-id>
      <title-group>
        <article-title>Tübingen at PoliticIT: Exploring SVMs, Pretrained Language Models, and Linguistic Transfer for Ideology Detection in Social Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Çağrı Çöltekin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Brivio</string-name>
          <email>matteo.brivio@student.uni-tuebingen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fidan Can</string-name>
          <email>ifdan.can@student.uni-tuebingen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Political Ideology Detection</institution>
          ,
          <addr-line>classification, SVM, BERT, XLM-Roberta, Transfer learning, Multitask learning</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Processing and Speech Tools for Italian</institution>
          ,
          <addr-line>Sep 7 - 8, Parma, IT</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Tübingen, Department of Linguistics</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>1</volume>
      <fpage>265</fpage>
      <lpage>272</lpage>
      <abstract>
        <p>This paper describes our approach to the EVALITA 2023 PoliticIT task on predicting political ideology and gender from Italian tweets. Furthermore, we investigate the efects of out-of-domain data (transcripts of parliamentary speeches) and cross-lingual transfer from Spanish. Overall, our simple traditional SVM classifier performed best according to the shared task evaluation, ranking first in ideology detection and second in gender identification. We also demonstrate promising results for out-of-domain data and cross-lingual transfer learning.</p>
      </abstract>
      <kwd-group>
        <kwd>(left</kwd>
        <kwd>right)</kwd>
        <kwd>(2) fine-grained political orientation ( left</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Although dificult to define, a relatively uncontroversial</title>
        <p>
          definition of political ideology is ‘a set of beliefs about the
proper order of society, and how it can be achieved’ [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ].
As the definition suggests, people’s political ideologies,
and especially those of politicians, have a significant
impact on society. Much like other personal characteristics,
such as gender, age, or native language [
          <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
          ], political
ideology can help to understand individual and social
behavior [6, 7, 8].
ing ideology is a crucial step for correctly understanding
political texts. Certain words or phrases have diferent
intended meanings depending on the ideological
position of speakers or authors [9, 10], and a message can
also be understood diferently based on the audience. In
its extreme case, so-called ‘dog whistles’ [11] may allow
politicians to send ‘coded messages’ to only a part of the
public who share their ideological position. As a result,
identifying political ideology is also important for natural
language understanding tasks.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>This paper describes our contribution to the EVALITA</title>
        <p>2023 [12] shared task [13] on predicting political
orientation and gender of politicians from social media
posts in Italian. The task is defined as three related
classification sub-tasks: (1) binary political orientation
nEvelop-O
LGOBE
0000-0003-1031-6327 (Ç. Çöltekin); 0000-0001-6273-6900
(M. Brivio); 0009-0006-8900-2038 (F. Can)
2.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Method and Experiments</title>
      <p>2.1. Data
The PoliticIT dataset comprises between 80 and 100
tweets from 1751 politicians (1298 for training and
development, 453 for testing). As usual, the test set was
released only at the end of the competition. All tweet
instances are anonymized by masking references to
politicians, political parties and other Twitter account
mentions. Further information on data collection and labeling
can be found in Russo et al. [13].</p>
      <sec id="sec-2-1">
        <title>Besides the PoliticIT dataset, we experiment with a</title>
        <p>few additional resources: transcripts of the parliamentary
speeches from the Italian section of the parliamentary
corpora collection ParlaMint [14] and the PoliticES 2023
shared task dataset [15, 16].</p>
        <sec id="sec-2-1-1">
          <title>Parliamentary speeches</title>
          <p>We use the Italian section
of the ParlaMint 3.0 pre-release, which contains speeches
both from the Senate and the Chamber of Deputies,
spanning from March 2013 to September 2022. We filter out
samples that belong to the chairperson and those that
cross validation. The input to the linear models are all
are less than 50 characters long. We also exclude speak- tweets belonging to each instance combined with an
arers without a known party afiliation, as well as those
who are afiliated with parties whose political orientation
is either not specified or reported to be ‘center’ or ‘big
tent’.1 Compared to the dataset released for the present
bitrary separator symbol inserted between each one of
them. To optimize the models trained on the ParlaMint
data, we use the complete PoliticIT training set as the
development set. Besides the predictions from the best
shared task, the political orientation in the ParlaMint
model after the hyper-parameter search, we also
experidata is more fine-grained. For binary class labels, we
ment with the majority vote of the top-n best performing
map ParlaMint orientation labels to left if the class in- hyper-parameters. As the voting ensembles performed
cludes L (left) and to
right if it includes R (right). We
worse than the single-best model, we do not report their
follow the same approach for the multi-class
classificaresults here.
tion task, but mark orientation labels as moderate if they</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Since the binary ideology classifier was more accuinclude C (center). For example, samples with original rate than the multi-class one in our initial experiments, labels CL (center-left) or</title>
        <p>CCL (center to center-left) are
we also include a post-processing step, where we adjust
mapped to moderate_left, while samples marked as R
multi-class labels when they conflict with binary labels. If
or FR (far-right) are mapped to right. The resulting cor- the multi-class classifier predicts ( moderate_)left while
pus includes 167 moderate_left, 153 moderate_right,
89 right instances, and no instances with (non-center)
left labels.
the binary classifier predicts right, we set the multi-class
label to moderate_right. Similarly, we resolve conflicts
where the binary classifier predicts left by setting the</p>
      </sec>
      <sec id="sec-2-3">
        <title>Compared to social media posts, parliamentary</title>
        <p>multi-class label to moderate_left. All linear models
speeches tend to be much longer, and a few speakers
are implemented using scikit-learn [17].
n-grams as features and weigh each feature using tf-idf. the tasks have diferent levels of dificulty, so we rely on
For each task, we tune the SVM margin parameter, C,
(range 0.001-5.0), the maximum word n-grams (range
04) and the character n-grams (range 0-6). Lastly, we test
their balanced multi-task learning (BMTL) framework to
mitigate this issue [22]. This requires wrapping each loss
ℓ into a function ℎ, such that ℎ(ℓ) = exp( ℓ ), where  is a
whether n-grams should be case normalized or not (word, temperature parameter that needs to be tuned.

character, both or none). We draw 10 000 configurations
uniformly from this hyper-parameter space and pick the
setting with the highest average F1-score using 10-fold</p>
      </sec>
      <sec id="sec-2-4">
        <title>1Political parties encouraging a broad spectrum of views among</title>
        <p>their members.</p>
      </sec>
      <sec id="sec-2-5">
        <title>We train the XLM-R model in three diferent settings:</title>
        <p>ifrst on the combination of the full PoliticES 2023 dataset
and the PoliticIT 2023 training set, then on the Italian
dataset alone and lastly only on the Spanish data.
Fi</p>
      </sec>
      <sec id="sec-2-6">
        <title>2https://huggingface.co/dbmdz/bert-base-italian-cased.</title>
        <p>The dataset con- [19].2
2.3. Transformer-based models
We also experiment with multi-task and cross-lingual
transfer learning by fine-tuning pretrained language
models. Specifically, we work both with a multilingual model,
XLM-RoBERTa [18], and a monolingual version of BERT</p>
      </sec>
      <sec id="sec-2-7">
        <title>We customize both models, adding a ‘common</title>
        <p>block’ followed by three linear classification heads. Both
solutions are implemented with PyTorch [20] and the</p>
      </sec>
      <sec id="sec-2-8">
        <title>Transformers library [21].</title>
        <p>The ‘common block’ is a 2-layer, bi-directional
Recurmodel’s representation of the CLS token for each tweet
in a given group. We include a ReLU activation function
one. The model leverages the hard-shared parameters
in the RNN and tries to learn the three tasks
simultaneously (i.e. in a multi-task fashion) by minimizing the
sum of their training losses: binary cross entropy loss
for the two binary tasks and categorical cross entropy
loss for the multi-class one. As observed by Liang and
Zhang, it is reasonable to assume that in such a setting,
have many more speeches than others in the corpus. To
get a balanced dataset, we randomly select at most 10
speeches from each speaker and concatenate them with
a separator token. The resulting corpus contains samples
with an average of 2575.95 space-separated tokens per
instance.</p>
        <sec id="sec-2-8-1">
          <title>PoliticES 2023 shared task data</title>
          <p>sists of 2797 instances, each containing 80 tweets written
by diferent users who share the same gender and
political views. All tweets are anonymized, mirroring the
anonymization procedure for the PoliticIT dataset. We
lingual transfer learning experiments between Italian and
Spanish. Further information about the data collection
2.2. Linear models
As a first, simple approach, we use linear support vector
machines (SVMs) with bag-of-n-grams features. For all
the results reported here, we train a separate linear SVM
classifier for each task, with both word and character
use the PoliticES 2023 shared task data to conduct cross- rent Neural Network which takes as input the language
and labeling can be obtained from García-Díaz et al. [16]. and a dropout layer after each RNN layer, except the last
nally, the BERT-based model is also trained on the Italian sole exception of the gender classification task, where the
dataset alone. monolingual model achieves a macro-averaged F1-score</p>
          <p>In an attempt to improve the multi-class ideology task of 81.17, outperforming the SVM model.
score, we include a post-processing step similar to the Even though it does not completely exceed the
perone described in 2.2. If the binary score is left and formance of the SVM-PoliticIT classifier, the
multilinthe multi-class prediction is moderate_right, we set the gual model trained on both the Italian and the Spanish
multi-class label to moderate_left and the other way data improves the scores substantially. Specifically, the
around if the two scores are right and moderate_left. model achieves a macro-averaged F1-score of 80.56 in the
Similarly, if the binary score is left and the multi-class gender task, outperforming SVMs. The efectiveness of
prediction is right, we set the latter to left and the cross-lingual transfer learning can be observed by
examother way around if the two scores are right and left. ining the results of the monolingual BERT model which,</p>
          <p>All models use the AdamW optimizer and a scheduler with the only exception of gender, are significantly lower.
with a learning rate that decreases following the values However, it should be noted that XLM-RoBERTa and
of the cosine function. As a pre-processing step, we BERT are pretrained using diferent training regimes. To
remove all punctuation marks and masking tokens origi- get a clearer understanding of the potential benefits of
nally used to anonymize the tweets (e.g., [POLITICIAN], cross-lingual transfer, we trained XLM-RoBERTa on the
[POLITICAL_PARTY]), expand all hashtags and convert Italian dataset alone. The scores we report for the latter
emojis into their text descriptions. model are marginally lower than those obtained with the</p>
          <p>All models are trained for 15 epochs and their hyper- model trained on both the Italian and the Spanish data,
parameters are optimized with Ray Tune [23], relying on suggesting a beneficial efect of cross-lingual transfer
a Bayesian optimization strategy. learning. This is further corroborated by the results in a
zero-shot setting, where the XLM-RoBERTa model was
trained on Spanish data only. Although the scores are
3. Results well below those of the other models, they are clearly
better than random, indicating signal in the cross-lingual
data.</p>
          <p>Our best performing system (based on linear SVMs)
ranked first among 7 participating teams on both the
binary and the multi-class ideology prediction tasks,
achieving a macro-averaged F1-score of 92.82 and 75.15, respec- 4. Discussion and Concluding
tively. With respect to the gender prediction task, our Remarks
system ranked second with a score of 79.25.</p>
          <p>Besides these simple models, we experimented with
parliamentary speeches as an out-of-domain dataset and
cross-lingual transfer, using pretrained language models.</p>
          <p>We summarize our results in Table 1.</p>
          <p>We presented our experiments on how to identify
political orientation and gender from social media posts
as part of the EVALITA 2023 PoliticIT shared task. Our
systems based on linear SVMs achieved the best overall
scores. We also experimented with classifiers based on
Out-of-domain data To investigate the efectiveness deep pretrained language models, focusing particularly
of training on out-of-domain data, we train an SVM on cross-lingual transfer.
model on parliamentary speeches from the ParlaMint Our first finding is that ‘traditional’ linear models show
dataset and test it on the PoliticIT test-set. We present a better overall performance than the deep learning ones.
the results in Table 1 (indicated as SVM-ParlaMint). Given their simplicity and lack of information other than</p>
          <p>ParlaMint includes both political orientation and gen- immediate surface features in the training data, their
sucder data for the speakers in the Italian parliament. How- cess may come as a surprise. However, this is not the first
ever, as noted in Section 2.1, it does not include speakers time that such models have been found to perform
simibelonging to any (non-center) left parties for the multi- larly or better than deep networks [24, 25, 26, 27, just to
class classification. As a result, the scores for multi-class name a few]. Similarly, the logistic regression classifiers
ideology detection are rather poor. Nonetheless, both the of Mosquera [28] obtained the third place, only slightly
binary ideology and gender scores are not too far behind behind the first two systems, in the PoliticES 2022 shared
the models trained on the in-domain data. task. This, of course, may be due to a lack of proper
tuning of the pretrained models. However, the fact that the
Cross-lingual transfer The results obtained with linear SVMs outperformed all other participating teams
monolingual and cross-lingual pretrained models, Italian (presumably the majority of them using pretrained
modBERT and XLM-RoBERTa, are also presented in Table 1. els) shows that these simpler models have some strong
Contrary to our expectations, the BERT-based model per- merits in this task.
forms rather poorly compared to SVM-PoliticIT, with the Our preliminary experiments on out-of-domain data
SVM-PoliticIT
SVM-ParlaMint
BERT-based model
XLM-R-based model (it)
XLM-R-based model (es)
XLM-R-based model (es-it)</p>
          <p>Task
ideology binary
ideology multi
gender
ideology binary
ideology multi
gender
ideology binary
ideology multi
gender
ideology binary
ideology multi
gender
ideology binary
ideology multi
gender
ideology binary
ideology multi
gender</p>
        </sec>
        <sec id="sec-2-8-2">
          <title>Precision</title>
        </sec>
        <sec id="sec-2-8-3">
          <title>Recall F1-score</title>
          <p>show that training an SVM model on parliamentary also picked up unmasked tokens left in the data that could
speeches results in lower, but comparable performance to suggest a party or politician’s name. This is the case with
training on in-domain-data. The results clearly indicate words like Giorgia, FI and Democratico, which allude to
that there is a considerable cross-domain signal for both the current right-wing Prime Minister Giorgia Meloni,
ideology detection and gender prediction. Although we a well-known Italian right-wing party Forza Italia, and
did not experiment with the use of both in- and out-of- a prominent Italian left-wing party Partito Democratico,
domain data together in this study, the results highlight respectively. Further and better approaches to explain
the potential for using out-of-domain data to improve model behavior is another direction for future research
ideology prediction in social media. that may shed light in narratives of diferent ideological</p>
          <p>Our experiments also show promising results on cross- groups.
lingual transfer. The best results we achieve with
pretrained language models come from fine-tuning a
crosslingual model using both Spanish and Italian data. Part of References
this success is probably due to the similar methodologies
followed by the organizers of both shared tasks. Further
investigation of cross-lingual and cross-country transfer
on ideology detection may shed light into universal and
culture-specific aspects of ‘ideology’.</p>
          <p>Our best-performing model is a bag-of-words SVM
classifier, relying only on token unigrams as features.</p>
          <p>Compared to the transformer-based ones, this model
allows for better interpretability3 with its unigram
features revealing some linguistic indications of political
orientation. Words such as immigrazione ‘immigration’,
clandestini ‘illegal immigrants’ and confini ‘borders’
appear constantly in right-wing discourse, while terms such
as diritti ‘rights’, democrazia ‘democracy’ and solidarietà
‘solidarity’ seem to point to more leftist views. The model
3Although one should be cautious because tokens are not
independent features, the weights assigned to individual tokens are,
nevertheless, easier to interpret than weights of a neural network or
weights in a model with overlapping features.
//www.aclweb.org/anthology/2020.emnlp-demos.6.
[22] S. Liang, Y. Zhang, A simple general approach to
balance task dificulty in multi-task learning, ArXiv
abs/2002.04792 (2020).
[23] R. Liaw, E. Liang, R. Nishihara, P. Moritz, J. E.
Gonzalez, I. Stoica, Tune: A research platform for
distributed model selection and training, arXiv
preprint arXiv:1807.05118 (2018).
[24] M. Medvedeva, M. Kroon, B. Plank, When sparse
traditional models outperform dense neural
networks: the curious case of discriminating between
similar languages, in: Proceedings of the Fourth
Workshop on NLP for Similar Languages,
Varieties and Dialects (VarDial), Association for
Computational Linguistics, Valencia, Spain, 2017, pp.
156–163. URL: https://aclanthology.org/W17-1219.</p>
          <p>doi:10.18653/v1/W17-1219.
[25] Ç. Çöltekin, T. Rama, Tübingen-Oslo at
SemEval2018 task 2: SVMs perform better than RNNs in
emoji prediction, in: Proceedings of the 12th
International Workshop on Semantic Evaluation,
Association for Computational Linguistics, New
Orleans, Louisiana, 2018, pp. 34–38. URL: https://
aclanthology.org/S18-1004.
doi:10.18653/v1/S181004.
[26] A. Caines, P. Buttery, REPROLANG 2020:
Automatic proficiency scoring of Czech, English,
German, Italian, and Spanish learner essays, in:
Proceedings of the Twelfth Language Resources and
Evaluation Conference, European Language
Resources Association, Marseille, France, 2020, pp.
5614–5623. URL:
https://aclanthology.org/2020.lrec1.689.
[27] K. Amponsah-Kaakyire, D. Pylypenko, J. Genabith,</p>
          <p>C. España-Bonet, Explaining translationese: why
are neural classifiers better and what do they learn?,
in: Proceedings of the Fifth BlackboxNLP
Workshop on Analyzing and Interpreting Neural
Networks for NLP, Association for Computational
Linguistics, Abu Dhabi, United Arab Emirates (Hybrid),
2022, pp. 281–296. URL: https://aclanthology.org/
2022.blackboxnlp-1.23.
[28] A. Mosquera, Alejandro mosquera at PoliticEs
2022: Towards robust Spanish author profiling and
lessons learned from adversarial attacks, in:
Proceedings of the Iberian Languages Evaluation
Forum (IberLEF 2022). CEUR Workshop Proceedings,
CEUR-WS, A Coruna, Spain. D. Moctezuma, and V.</p>
          <p>Muniz-Sánchez, 2022.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Jost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Federico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Napier</surname>
          </string-name>
          , Political ideology:
          <article-title>Its structure, functions, and elective afinities</article-title>
          ,
          <source>Annual review of psychology 60</source>
          (
          <year>2009</year>
          )
          <fpage>307</fpage>
          -
          <lpage>337</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Erikson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. L.</given-names>
            <surname>Tedin</surname>
          </string-name>
          , American public opinion:
          <article-title>Its origins, content and impact</article-title>
          , Routledge,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Stachl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pargent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hilbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Harari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schoedel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vaid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Gosling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bühner</surname>
          </string-name>
          ,
          <article-title>Personality research and assessment in the era of machine learning</article-title>
          ,
          <source>European Journal of Personality</source>
          <volume>34</volume>
          (
          <year>2020</year>
          )
          <fpage>613</fpage>
          -
          <lpage>631</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Christian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Suhartono</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Z.</given-names>
            <surname>Zamli</surname>
          </string-name>
          ,
          <article-title>Text based personality prediction from multiple social media data sources using pre-trained language model and model averaging</article-title>
          ,
          <source>Journal of Big Data</source>
          <volume>8</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes-y Gómez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Overview of the 6th author profiling task at PAN 2018: multimodal gender identification</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>