<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>R. Pan);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>at PoliticES-IberLEF2023: Evaluating Transformers for Detecting Political Ideology in Spanish Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ronghao Pan</string-name>
          <email>ronghao.pan@um.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Camilo Caparrós-Laiz</string-name>
          <email>camilo.caparrosl@um.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ángela Almela</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Informática, Universidad de Murcia, Campus de Espinardo</institution>
          ,
          <addr-line>30100</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Facultad de Letras, Universidad de Murcia, Campus de La Merced</institution>
          ,
          <addr-line>30001, Murcia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Natural Language Processing</institution>
          ,
          <addr-line>Transformers, Politic ideology detection, Large Language Model, Multiclass</addr-line>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper describes the participation of the UMUTeam in the PoliticES shared task organized at IberLEF 2023 within the SEPLN conference. It is an automatic document classification task on clusters of texts, which consists of extracting the self-assigned gender and profession as demographic traits, and political ideology as a psychographic trait from a set of texts written in Spanish by several authors that share those traits. For this task, we have fine-tuned the MarIA, a transformer model proposed for Spanish, to create classification models for each feature. After several submissions for these tasks, our team ranked seventh out of 12 participants, with an average F1 score of 69.225% of all classification models.</p>
      </abstract>
      <kwd-group>
        <kwd>classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In general, individuals tend to be reluctant to heed the advice and instructions of politicians
who do not coincide with their ideology. In extreme cases, individuals may show a strong bias
towards a particular political party while roundly rejecting others’ ideology. In this line, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
analyzed the relationship between personality traits and political ideology by collecting data
from 21 countries. The author observed a correlation between political ideology and the big
ifve personality traits. Therefore, political ideology is a psychographic characteristic that helps
to understand individual and social behavior, as it encompasses moral and ethical values and
inherent attitudes, evaluations, biases, and prejudices [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Furthermore, it can enhance
microtargeting eforts, enabling public authorities and local governments to adopt more efective
communication strategies during crises [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        With the aim of detecting political ideology information from Spanish texts, the shared-task
PoliticES [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] has been organized as part of the IberLEF 2023 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] workshop within the framework
of the SEPLN 2023 conference. The organizers proposed an automatic document classification
task on clusters of texts, consisting of the extraction of self-assigned gender and profession as
demographic traits and political ideology as a psychographic trait from a set of texts written in
Spanish by several authors that shared those traits.
      </p>
      <p>This work presents the participation of the UMUTeam in this shared task, which is based
on exploring the fine-tuning of diferent pre-trained and Transformers-based Large Language
Models (LLMs) to detect gender, profession, and political ideology. The rest of the paper is
organized as follows. Section 2 presents the task and dataset provided. Section 3 describes
the methodology of our proposed system for addressing the task. Section 4 shows the results
obtained. Finally, Section 5 concludes the paper with some findings and possible future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description</title>
      <p>The shared task PoliticES 2023, organized at IberLEF workshop, aims to extract political ideology
information from a set of texts written in Spanish by several authors. Specifically, the organizers
propose a task of identifying demographic traits (gender and profession) and political ideology
as psychographic traits from a given cluster of texts (document classification). In the case of the
identification of political ideology, the organizers have defined it as a binary problem and as a
multiclass problem.</p>
      <p>
        The dataset provided is an extension of the PoliCorpus 2020 dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the corpus used
for the PoliticES 2022 shared task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It was gathered between 2020 and 2022 from Twitter
accounts of politicians, political journalists, and celebrities in Spain using UMUCorpusClassifier
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The users of the dataset are labeled with their gender (male, female), profession (politician,
journalist, celebrity), and political spectrum on two axes: binary (left, right) and multiclass
(left, moderate left, moderate right, right). Regarding the tweets collected from each user, the
organizers removed discarded retweets and tweets that contain headlines from news sites and
removed tweets written in languages other than Spanish. Moreover, they anonymized them by
replacing all mentions with @user, except for real users. Furthermore, other entities, such as
political party references, are also replaced with @political_party token. The final dataset is
composed of approximately 2,800 diferent clusters. The tweets that belong to each cluster are
selected favoring diversity, including texts from diferent dates and topics. For this shared task,
the dataset is divided into training and test sets (80%-20%).
      </p>
      <p>The training set consists of 180,000 tweets from 2,250 clusters and the test set consists of 547
clusters with a total of 43,760 tweets. The distribution of demographic and psychographic traits
of each cluster is shown in Table 1.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>This task involves identifying gender and profession (demographic traits) and political ideology
(psychographic traits) of users in a given set of texts. The pipeline used to participate in this
task can be described as follows. First, the dataset has been processed, and emoji features have
also been added. Second, the training dataset was split into training and validation at the tweet
level. Third, a classification model is created for each of the features using the fine-tuning
approach. Fourth, having the classification models at the sentence level, two strategies have
been evaluated to identify the demographic and psychographic traits of the users (at document
level): (1) mode, which consists of predicting each user’s tweet individually and selecting the
most repeated label among the results obtained with the classifier, and (2) highest probability,
which selects the label with the highest probability. The system architecture is depicted in
Figure 1</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset preprocessing</title>
        <p>As mentioned above, tweets from the same user will have the same demographic and
psychographic traits, Therefore, to carry out the problem of identifying these traits, we have created a
sentence-level classification model for each of them with all the tweets and the distribution of
demographic and psychographic traits of the dataset as shown in Table 2. In this case, we have
divided the training set into two subsets (80%-20%): training and validation. The customized
validation split is created using stratified sampling, in order to maintain a balance between the
labels.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Fine-tuning approach</title>
        <p>
          We utilized the fine-tuning approach of a transformer-based masked language model for Spanish
called MarIA [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] to carry out the identification of diferent features. MarIA is based on the
RoBERTa base model and has been pre-trained using the largest Spanish corpus known to date,
with a total of 570 GB of clean and deduplicated text processed. The fine-tuning process involves
adapting and adding a classification layer to the model to perform the training of the complete
model. In this way, the model takes advantage of MarIA’s pre-trained linguistic knowledge
and adapts it specifically for a particular classification task, which can significantly improve
performance on that task. The model has been fine-tuned with a training batch size of 16, 6
epochs, a learning rate of 2e-5, and a decay of weights of 0.01.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>This section describes the systems submitted by our team in each run and the overall results
obtained in this shared task. It should be noted that each participating team was allowed to
submit ten runs.</p>
      <p>We sent two runs for this task. The results, and a brief description of each, are depicted in
Table 3. The first run is based on the mode strategy, which consists of selecting the most frequent
label obtained in the user’s text set for each feature, achieving an average F-score of 69.225%.
The second run, which entailed the highest probability strategy, achieved an average F-score of
55.889%. It can be observed that the macro F1-score obtained with the mode strategy applied in
diferent classification models for both demographic and psychographic features performed
better than the highest probability strategy. In addition, the binary classification models (gender
and binary ideology) outperformed the multiclass models (profession and ideology multiclass).</p>
      <p>The oficial leaderboard for this task is depicted in Table 4. We achieved the seventh position
in the ranking with an average F-score of 69.225%. The teams ELiRF VRAIN and HiTZ-IXA
achieved the best position, outperformed our best run with an average F-score of 81.131% and
79.348%, respectively.</p>
      <p>In order to perform the error analysis and check what kind of wrong predictions our system
makes, a normalized confusion matrix with truth labels have been used, which consists of a
table showing the distribution of the predictions of a model with respect to the truth label of
the data. The confusion matrix of the system using the mode strategy is shown in Figure 2. It
can be noticed that our model tends to confuse the female gender with the male gender, with a
percentage of 37.95%. As for the prediction of users’ professions, the model usually confuses
celebrities with journalists with a percentage of 92.86% because there are few celebrity texts
in the training set (see Table 2). Regarding the identification of political ideology, it can be
observed that the model has a good accuracy in the identification of left-wing with a percentage
of 99.39%. However, it tends to confuse right-wing and left-wing political ideology, with a
percentage of 30.45%. In the case of the identification of political ideology at diferent levels
(multiclass classification), it is observed that by having more moderate left cases in the training
set, the model has an accuracy of 90% in the identification of this class. However, the model
tends to confuse left-wing ideology with moderate left (52.99%) and right-wing with moderate
right ideology (58.21%).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>These working notes summarize the participation of the UMUTeam in the PoliticES shared task
(IberLEF 2023). We achieved a 7/12 on the mean of all F1-score (69.225%) for the demographic
and psychographic feature identification models. For this, we used the MarIA fine-tuning
approach with the processed dataset and emoji features.</p>
      <p>As future work, we are planning to improve our pipeline using an expanded LLMs model with
political speech, i.e., fine-tuning a Masked Language Model (MLM) model with political text
(a) Gender
(b) Profession
(c) Binary ideology
(d) Multiclass ideology
and later fine-tuning this model for detecting political ideology. In addition, we are planning to
ifne-tune other pre-trained Spanish models to see if they improve MarIA’s performance. The
source code is available via GitHub.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work is part of the research projects AIInFunds (PDC2021-121112-I00) and
LT-SWM (TED2021-131167B-I00) funded by MCIN/AEI/10.13039/501100011033 and by
the European Union NextGenerationEU/PRTR. This work is also part of the
research project LaTe4PSP (PID2019-107652RB-I00/AEI/ 10.13039/501100011033) funded by
MCIN/AEI/10.13039/501100011033.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fatke</surname>
          </string-name>
          ,
          <article-title>Personality traits and political ideology: A first global assessment</article-title>
          ,
          <source>Political Psychology</source>
          <volume>38</volume>
          (
          <year>2017</year>
          )
          <fpage>881</fpage>
          -
          <lpage>899</lpage>
          . URL: https://onlinelibrary.wiley. com/doi/abs/10.1111/pops.12347. doi:https://doi.org/10.1111/pops.12347. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/pops.12347.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Verhulst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Eaves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Hatemi</surname>
          </string-name>
          ,
          <article-title>Correlation not causation: The relationship between personality traits and political ideologies</article-title>
          ,
          <source>American Journal of Political Science</source>
          <volume>56</volume>
          (
          <year>2012</year>
          )
          <fpage>34</fpage>
          -
          <lpage>51</lpage>
          . URL: https://onlinelibrary.wiley.com/doi/abs/10. 1111/j.1540-
          <fpage>5907</fpage>
          .
          <year>2011</year>
          .
          <volume>00568</volume>
          .x. doi:https://doi.org/10.1111/j.1540-
          <fpage>5907</fpage>
          .
          <year>2011</year>
          .
          <volume>00568</volume>
          .x. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1540-
          <fpage>5907</fpage>
          .
          <year>2011</year>
          .
          <volume>00568</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Colomo-Palacios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Psychographic traits identiifcation based on political ideology: An author analysis study on spanish politicians' tweets posted in 2020, Future Generation Computer Systems 130 (</article-title>
          <year>2022</year>
          )
          <fpage>59</fpage>
          -
          <lpage>74</lpage>
          . URL: https://www.sciencedirect.com/science/article/pii/S0167739X21004921. doi:https://doi. org/10.1016/j.future.
          <year>2021</year>
          .
          <volume>12</volume>
          .011.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>García-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Ureña-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Overview of PoliticES at IberLEF 2023:
          <article-title>Political ideology detection in Spanish texts</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Montes-y Gómez, Overview of IberLEF 2023: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2023), co-located with the 39th Conference of the Spanish Society for Natural Language Processing (SEPLN 2023), CEUR-WS</article-title>
          .org,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T. Martín</given-names>
            <surname>Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>García-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Ureña</surname>
          </string-name>
          <string-name>
            <surname>López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Valencia</given-names>
            <surname>García</surname>
          </string-name>
          , Overview of PoliticEs 2022:
          <article-title>Spanish Author Profiling for Political Ideology</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          ,
          <fpage>265</fpage>
          -
          <lpage>272</lpage>
          (
          <year>2022</year>
          -
          <fpage>09</fpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>J. A. G.-D. y Ángela</surname>
          </string-name>
          <article-title>Almela y Gema Alcaraz-Mármol y Rafael Valencia-García, Umucorpusclassifier: Compilation and evaluation of linguistic corpus for natural language processing tasks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>65</volume>
          (
          <year>2020</year>
          )
          <fpage>139</fpage>
          -
          <lpage>142</lpage>
          . URL: http: //journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6292.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Palao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Oller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Penagos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Maria: Spanish language models</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          ). URL: https://upcommons.upc.edu/handle/2117/367156# .YyMTB4X9A-0.mendeley. doi:
          <volume>10</volume>
          .26342/2022-68-3.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>