<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>JU_NLP_DID at Touché: An Attempt to Aspects of Power from Parliamentary Debates Identify</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adnan Khurshid</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dipankar Das</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rajdeep Khaskel</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Suchanda Datta</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Jadavpur University</institution>
          ,
          <addr-line>Kolkata, 700032</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>0</volume>
      <fpage>9</fpage>
      <lpage>12</lpage>
      <abstract>
        <p>Parliamentary debates shape critical aspects of citizens' lives and often influence global policies. Analyzing these debates computationally poses unique challenges due to the indirect and complex nature of political discourse. This paper addresses two key variables in parliamentary speeches: the political ideology of the speaker and their affiliation with either the governing party or the opposition. We approach these subtasks as binary classification problems, employing a combination of Term Frequency-Inverse Document Frequency (TF-IDF) vectorization and Support Vector Machines (SVM) for our analysis. Our methodology is designed to capture the nuanced language of parliamentary debates and effectively classify speakers based on their political stance and party alignment. The results demonstrate the efficacy of TF-IDF with SVM in handling the intricacies of political speech, providing a robust framework for further research in computational political analysis.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;TF-IDF</kwd>
        <kwd>SVM</kwd>
        <kwd>Binary Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Parliamentary debates are pivotal in shaping not only the national policies of a country but also
influencing global political landscapes. These debates, characterized by their indirect and
nuanced discourse, pose significant challenges for computational analysis. Understanding the
ideological stance and power alignment of speakers within these debates can provide valuable
insights into political dynamics and decision-making processes.</p>
      <sec id="sec-1-1">
        <title>1.1. Objective</title>
        <p>
          This paper addresses the task[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] of classifying two critical variables associated with speakers in
parlia- mentary debates: the political ideology of the speaker’s party and whether the speaker’s
party is in the governing coalition or in opposition. These tasks are formulated as binary
classification problems, necessitating sophisticated natural language processing (NLP) techniques
to handle the complexity and variability of political speech.
        </p>
        <p>The data for this study is derived from the ParlaMint corpus, a multilingual and comparable
dataset of parliamentary debates across various countries. The corpus has been curated to
minimize confounding variables, ensuring that the analysis focuses on the content and context of
the speeches rather than extraneous factors such as speaker identity. The provided data includes
both the original speeches and their English translations, facilitating the development of
multilingual models and cross-linguistic analyses.</p>
        <p>To tackle the classification tasks, we employ a combination of Term Frequency-Inverse
Document Fre- quency (TF-IDF) vectorization and Support Vector Machines (SVM). TF-IDF is
utilized to convert textual data into numerical representations that capture the importance of
words within the speeches, while SVM is used for its effectiveness in handling high-dimensional
feature spaces and binary classification problems.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Contribution</title>
        <p>This study contributes to the field of computational political analysis by offering a robust
framework for identifying political ideology and power alignment in parliamentary debates. By
leveraging advanced NLP techniques, we aim to enhance the understanding of political discourse
and provide a foundation for further research in this area. The results of our analysis demonstrate
the potential of TF-IDF and SVM in addressing the challenges posed by the indirect nature of
parliamentary speech, paving the way for more accurate and insightful political analysis tools.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>In today’s digital era, parliamentary debates have transcended the confines of legislative
chambers to include online platforms, fundamentally reshaping discourse. Within these digital
spaces, social networks wield significant influence over opinions and provide valuable data for
sentiment analysis and power identification. Understanding the intricate power dynamics at play
is essential for decoding how influence is disseminated and policies are formulated. By harnessing
the capabilities of sentiment analysis and computational techniques, we can shed light on the
underlying power structures and sentiment trends, ultimately enhancing decision-making
processes. This multifaceted approach involves analyzing various factors such as speaking
patterns, party contributions, responses to arguments, social network connections, and media
coverage to unveil influential actors and dominant dynamics within parliamentary debates.</p>
    </sec>
    <sec id="sec-3">
      <title>3. System Overview</title>
      <p>The system developed for parliamentary power identification involves several steps, including
data preprocessing, feature extraction, and classification. The primary goal is to classify
parliamentary text data using machine learning techniques. Below is a detailed overview of the
system components and processes.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Preprocessing</title>
        <p>We employed automated English translations for our experiments. The raw textual data
underwent rigorous preprocessing to facilitate feature extraction and classification. This
preprocessing pipeline encompassed the following steps:</p>
        <p>Lowercasing: All text is converted to lowercase to ensure uniformity. HTML Tag Removal:
HTML tags are removed using regular expressions to clean the text. The English translated text
provided in the dataset contained HTML commands, necessitating their removal to ensure
accurate preprocessing and avoid interference with subsequent analyses. Punctuation Removal:
All punctuation marks are removed to reduce noise. Stopword Removal: Common stop words are
removed using NLTK’s stopword list, which helps in focusing on the meaningful words in the text.
Lemmatization: Words are lemmatized to their base or dictionary form using the WordNet
lemmatizer. This involves: Tokenizing the text. Tagging each word with its part of speech.
Mapping the POS tag to WordNet’s POS tag format. Lemmatizing each word based on its POS tag.
This preprocessing ensures that the text data is clean, normalized, and stripped of irrelevant parts,
making it suitable for feature extraction.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Feature Extraction</title>
        <p>After preprocessing, the text data is transformed into numerical features using the TF-IDF (Term
Frequency-Inverse Document Frequency) vectorization technique. This method converts the text
into a matrix of TF-IDF features, which reflects the importance of words in the corpus:</p>
        <p>TF-IDF Vectorization: This technique is used to convert the preprocessed text data into
numerical vectors. It captures the importance of a word in a document relative to the entire corpus. The
Tfidf Vectorizer from scikit-learn is used with default parameters.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Model Building and Training</title>
        <p>For the classification task, a Support Vector Machine (SVM) model with a linear kernel is initially
employed. The SVM classifier is chosen for its effectiveness in high-dimensional spaces and its
capability to handle large feature sets resulting from TF-IDF vectorization. The dataset is split into
training and testing sets, with 80% of the data used for training and 20% for testing. The SVM model
is trained on the TF-IDF vectors of the training set.</p>
        <p>Initially, we had tried using bi-grams and n-grams with SVM but did not observe relevant
improvements in performance, hence we focused solely on uni-gram TF-IDF representations. To
optimize the SVM model, hyper-parameter tuning is performed using RandomizedSearchCV. This
approach is selected over GridSearchCV due to its ability to efficiently explore a wide range of
parameter combinations with fewer iterations, thus reducing computational burden while still
providing robust parameter estimates. Given our system's limited computational power,
RandomizedSearchCV is configured with 5 iterations and 2-fold cross-validation.
Hyper-parameter tuning is performed using RandomizedSearchCV with the following parameter
distribution:
● C: [0.1, 1, 10, 100, 1000]
● kernel: ['linear', 'rbf', 'sigmoid']
● probability: [True]
● gamma: ['scale', 'auto']
● coef0: [0.0, 0.1, 0.5, 1.0]
After tuning, the best parameters for the SVM model are found to be:
● C: 10
● kernel: 'rbf'
● probability: True
● gamma: 'scale'
● coef0: 0.1
These parameters enhance the SVM model's performance significantly for our classification task.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Model Evaluation</title>
        <p>The trained SVM model is evaluated using the test set. Several evaluation metrics are computed
to assess the model’s performance:</p>
        <p>Classification Report: This includes precision, recall, and F1-score for each class, providing a
detailed performance analysis. Confusion Matrix: The confusion matrix shows the true positive,
true negative, false positive, and false negative counts, offering insights into the model’s
prediction errors. ROC Curve and AUC: The Receiver Operating Characteristic (ROC) curve and
the Area Under the Curve (AUC) score are computed to evaluate the model’s ability to discriminate
between classes. The AUC score provides a single metric to summarize the model’s performance.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>Average F1
Max Precision
Max Recall
Max F1
Power</p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>This study demonstrates the application of advanced natural language processing techniques to
the analysis of parliamentary debates, focusing on identifying the political ideology of speakers
and their party’s power status. By leveraging the ParlaMint corpus, which provides a rich and
multilingual dataset of parliamentary speeches, we have developed a robust framework for
addressing these binary classification tasks.</p>
      <p>Our approach involved experimenting with various methodologies. We initially utilized Term
Frequency-Inverse Document Frequency (TF-IDF) vectorization combined with Support Vector
Machines (SVM), which proved effective in handling the complexity and nuance of political
discourse. The results highlight the capability of TF-IDF and SVM to capture significant features of
parliamentary speeches. Each method demonstrated unique strengths, contributing to a
comprehensive understanding of the political dynamics within parliamentary debates. This work
not only provides valuable insights into the political dynamics within parliamentary debates but
also sets the stage for further research in computational political analysis. Future studies can build
on this foundation by refining these models, exploring ensemble methods, and expanding the
scope to include additional political variables and more diverse datasets.</p>
      <p>In conclusion, our study underscores the importance of computational approaches in
understanding political discourse and offers a promising methodology for analyzing parliamentary
debates. The techniques and findings presented here contribute to the broader field of political text
analysis, enhancing our ability to decipher and interpret the intricate language of politics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Erjavec</surname>
          </string-name>
          , Tomaž, Maciej Ogrodniczuk, Petya Osenova, Nikola Ljubešić, Kiril Simov, Andrej Pančur,
          <source>Michał Rudolf et al. "The ParlaMint corpora of parliamentary proceedings." Language resources and evaluation 57</source>
          , no.
          <issue>1</issue>
          (
          <year>2023</year>
          ):
          <fpage>415</fpage>
          -
          <lpage>448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Pedregosa</surname>
          </string-name>
          , Fabian, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel et al.
          <article-title>"Scikit-learn: Machine learning in</article-title>
          <source>Python." the Journal of machine Learning research 12</source>
          (
          <year>2011</year>
          ):
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Çöltekin</surname>
            , Çağrı, Matyáš Kopp, Katja Meden, Vaidas Morkevicius, Nikola Ljubešić, and
            <given-names>Tomaž</given-names>
          </string-name>
          <string-name>
            <surname>Erjavec</surname>
          </string-name>
          .
          <article-title>"Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset</article-title>
          and
          <string-name>
            <given-names>Simple</given-names>
            <surname>Baselines</surname>
          </string-name>
          .
          <source>" arXiv preprint arXiv:2405.07363</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Russo, Daniel, Salud María Jiménez-Zafra, José Antonio García-Díaz, Tommaso Caselli,
          <string-name>
            <given-names>Marco</given-names>
            <surname>Guerini</surname>
          </string-name>
          , L. Alfonso
          <string-name>
            <surname>Ureña-López</surname>
          </string-name>
          , and
          <string-name>
            <surname>Rafael</surname>
          </string-name>
          Valencia-García.
          <article-title>"PoliticIT at EVALITA 2023: Overview of the Political Ideology Detection in Italian Texts Task</article-title>
          .
          <source>"</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Tarkka</surname>
          </string-name>
          , Otto, Jaakko Koljonen, Markus Korhonen, Juuso Laine, Kristian Martiskainen,
          <source>Kimmo Elo, and Veronika Laippala. "Automated Emotion Annotation of Finnish Parliamentary Speeches Using GPT-4." In Proceedings of the IV Workshop on Creating, Analysing, and Increasing Accessibility of Parliamentary Corpora (ParlaCLARIN)@ LRECCOLING</source>
          <year>2024</year>
          , pp.
          <fpage>70</fpage>
          -
          <lpage>76</lpage>
          .
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mochtak</surname>
            , Michal,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Rupnik</surname>
            , and
            <given-names>Nikola</given-names>
          </string-name>
          <string-name>
            <surname>Ljubešić</surname>
          </string-name>
          .
          <article-title>"The ParlaSent multilingual training dataset for sentiment identification in parliamentary proceedings</article-title>
          .
          <source>" arXiv preprint arXiv:2309.09783</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Eskişar</surname>
          </string-name>
          ,
          <string-name>
            <surname>Gül M. Kurtoğlu</surname>
            , and
            <given-names>Çağrı</given-names>
          </string-name>
          <string-name>
            <surname>Çöltekin</surname>
          </string-name>
          .
          <article-title>"Emotions running high? a synopsis of the state of turkish politics through the parlamint corpus."</article-title>
          <source>In Proceedings of the Workshop ParlaCLARIN III within the 13th Language Resources and Evaluation Conference</source>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>70</lpage>
          .
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Kiesel</surname>
          </string-name>
          , Ç. Çöltekin,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heinrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alshomary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Longueville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Erjavec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Handke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kopp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ljubešić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Meden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mirzakhmedova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Morkevičius</surname>
          </string-name>
          , T. ReitisMunstermann, M. Scharfbillig,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stefanovitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wachsmuth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , Overview of Touché 2024:
          <article-title>Argumentation Systems</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. M. D. Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Confer- ence of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>