<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>at FinancES 2023: Financial Targeted Sentiment Analysis in Spanish Combining Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miguel Ángel Rodríguez-García</string-name>
          <email>miguel.rodriguez@urjc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrián Riaño-Martínez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Roldán-Álvarez</string-name>
          <email>david.roldan@urjc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Soto Montalvo-Herranz</string-name>
          <email>soto.montalvo@urjc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Deep Learning</institution>
          ,
          <addr-line>Transformers, Natural Language Processing, Name Entity Recognition, Targeted Senti-</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad Rey Juan Carlos</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Financial and economic news is continuously monitored due to their impact on future stock prices. Thus, the polarity extraction from the news is a very relevant task for investment decision-making by traders. In this sense, Sentiment Analysis models can provide accurate methods to extract signals that influence this decision-making. In this work, we describe the contribution to IberLEF 2023 Challenge - FinancES, where we proposed a hybrid approach that addresses the targeted sentiment analysis by creating a pipeline with diferent phases. First, a phase for cleaning texts, followed by an entity recognition phase and, finally, the polarity extraction. The hybrid approach combines diferent models to proceed with each phase. Thus, RoBERTa transformer architecture is employed as a NER, BETO transformer model is employed for polarity analysis, and, finally, a Spanish Spacy model is used for the part-of-speech tagging process. Although the proposed approach has still scope for improvement, since it reached mid-table positions in the leaderboard, it put forwards a diferent method to carry out the proposed classification Financial markets are changeable and sensitive to global events and phenomena [1]. Factors like political news to users' opinions can produce an immediate efect directly on the market, making it grow positively if the news is good or, conversely, pull downwards if it is bad [2]. Therefore, understanding the emotions embedded in these resources can assist financial professionals and economics in predicting stock market fluctuations [ 3].</p>
      </abstract>
      <kwd-group>
        <kwd>Combining</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>(S. Montalvo-Herranz)</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        information by recognising the emotion, opinion or polarity in human language [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This type
of analysis is becoming an essential tool in various domains to transform emotions and attitudes
into actionable knowledge that assists in decision-making processes [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. In this work, the
proposed challenge is focused on the financial domain [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The IberLEF FinancES shared task
go beyond extracting the polarity of text and require recognising the financial target entity
afected by this opinion [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. To address these tasks, we propose a Deep Learning based system
that combines three diferent models to carry out polarity analysis, entity recognition, and
grammar tagging. After analysing various language models, we utilised RoBERTa
(https://huggingface.co/MMG/xlm-roberta-large-ner-spanish) for named entity recognition, a Spacy model
for Spanish is employed for identifying the grammatical category and, finally, we selected BETO
(https://huggingface.co/finiteautomata/beto-sentiment-analysis) as a polarity extractor.
      </p>
      <p>The rest of the paper is organised as follows. Section 2 presents related work, where various
approaches that address Opinion Mining in the financial domain are analysed. Section 3 details
the distribution of the dataset delivered for each task and describes the systems’ architecture
proposed. Section 4 analyses the results achieved in the challenge. Finally, Section 5 summarises
the findings harvested facing the challenge, and point out various future research line to explore.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related work</title>
      <p>
        Sentiment Analysis is growing and taking an important role in better understanding users’
opinions in several domains [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This growing importance is being applied to finance, since
several studies have demonstrated to find strong correlations between textual sentiment and
other financial measures, such as stock returns and volatilities [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Following this line of work,
Xiang et al., in [11], address the sentiment analysis in the financial domain trying to predict the
sentiment intensity of a determined target in a text. They proposed a semantic and syntactic
enhanced neural model, called (SSENM), that constructs sentiment-aware representations
considering target information, semantic features and syntactic knowledge for modelling more
precisely the correlation between target mentions in the text and sentiment-relevant keywords.
Addressing the same working task, Shang et al., in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] propose LECN, a novel Lexicon Enhanced
Collaborative Network to capture associations between financial targets and sentiment signals.
The model is mainly based on three components: the shared Encoder Layer, which receives
input sentences encoded into word embeddings by using BERT and is based on a BiLSTM
architecture; the Task-specific Attention layer responsible for driving the model to focus on text
segments linked to sentiments and targets; and finally, the message selective-passing mechanism
in charge of adding features information gathered from previous interactions from the target
extraction and sentiment analysis tasks to control the information shared between both tasks
and enhance the collaborative efect. A diferent approach that faces the same classification
problem is the work proposed by Shijia et al. [12], where they proposed a neural network
architecture composed of a stack of LSTM layers and it receives as inputs word vectors encoded
by the word2vec model.
      </p>
      <p>In this analysis, we have selected a set of proposals that face the same classification problem
thrown in the IberLEF challenge. Various architectures have been analysed, from neural
networks to transformers architectures. Given the outcomes achieved by the proposals, we
addressed the challenge by using transformers, specifically, language models, since they seem
to reach more promising results.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Material and methods</title>
      <p>This section describes the datasets proposed by the organizers and the model developed to deal
with the tasks delivered in the challenge.
3.1. Data
The FinancES dataset is constituted of Spanish news headlines harvested from digital newspapers
specialized in the targeted domain, such as Expansión, El Economista, Modaes and El Financiero
[13]. The resulting dataset was about 14k headlines, manually labelled by three individuals,
selecting the target entity and the sentiment polarity on three dimensions: target, companies,
and consumers. During this selection, some headlines were discarded cause of their short length
and controversy during the labelling process. As a result, the dataset was reduced to 6k-8k.
Table 1 shows the distribution of the datasets released for the practising and evaluation phase.
target_sentiment</p>
      <p>Total
companies_sentiment
consumers_sentiment</p>
      <p>Total
Total</p>
      <p>label
positive
neutral
negative
positive
neutral
negative
positive
neutral
negative</p>
      <p>Practice</p>
      <p>As we can see in Table 1, each label is separated into the three common values assigned in
sentiment analysing tasks: positive, neutral and negative, to depict a more detailed picture of the
dataset. This organization shows some irregularities in the datasets of practice and evaluation,
where the total count of the ‘targets_sentiment’ and ‘companies_sentiment’ does not match
the ‘consumers_sentiment’, difering in two and three examples. This inconsistency is because
these samples were labelled diferently, with a label not included in these values. Apart from
this, it is noteworthy that the clear unbalance of neutral cases in ‘target_sentiment’ for practice
and evaluation datasets has a highly unequal distribution of examples.</p>
      <p>Another characteristic that draws attention is the similar distribution in
‘companies_sentiment’, where the neutral label difers to others in a large number of samples.
3.2. Method
The FinancES challenge of the IberLEF 2023 evaluation campaign consisted of two main complex
tasks detecting the main economic target and its linked sentiment and characterizing the polarity
at the document level for companies and consumers. Figure 1 shows the architecture that was
proposed to address both these two main tasks.</p>
      <p>As we can see in Figure 1, the proposed system is configured as a pipeline, where the inputs
are the news headlines. First, the cleaner module is responsible for unifying the format of
the sentences, removing links, emojis and lowercase words. Next, the target extractor aims
at identifying potential financial targets. It receives the pre-processed text and employs a
pre-trained model to identify named entities specifically, it operates the transformer-based
language model RoBERTa. If the model does not identify entities, we came up with a second
extractor method, which was based on the premise that the sentence structure of news headlines
is not too complex and follows a simple composition like DET+NOUN+VERB. In this sense,
we use grammatical tagging to identify potential targets. Concretely, we employ the spaCy
pipeline for NER to mark up each word of each sentence to a particular part of speech. Then, we
created a set of 14 simple grammar rules to extract potential targets by combining in diferent
ways grammatical tags like DET, NOUN, ADJ, and VERB, among others. To decide how to
define these grammar rules, we accomplished a statistical analysis of the headlines given to
analyse what were the most used grammar structures by the writers. These rules are employed
specifically from the start of the sentence until the word tagged as a VERB is located since we
assume the target entity will appear in the first part of the sentence. These are the 14 grammar
rules designed are shown in Table 2. When a news headline follows this structure, the words
that match the rule are extracted systematically. For instance, if we have the following headline:
“Las empresas chinas piden menos burocracia para invertir en España”, the set of words that
will be extracted is: “Las empresas chinas”. On the Sentiment Analysis task, the sentiment
analyser module tackles each task diferently. The polarity of “target_sentiment” in the sentence
is obtained by employing the transformer model BETO, which, for each sentence, it returns two
values, the label corresponding to the sentiment that could be NEU or POS and NEG referencing
the three existing types of emotional tones, and its score. Thus, to assign this tone, we have
defined a threshold, which was established based on several experiments conducted. Therefore,
depending on the label and score inferred by the model, and if this score defeats the threshold
established, the polarity is assigned in this task, otherwise, it is left empty.</p>
      <p>For companies and consumers, the sentiment analyzer module needs first that entities are
classified into a person or organization. Two techniques are used in this classification: using
the labels inferred by the pre-trained Name Entity Recognition model or two dictionaries of
words that group nouns related to persons and organizations domains. When the entities are
classified, the sentiment analyzer module analyses the context where the entity is placed and
extracts the polarity of text equally to the ‘target_sentiment’ task described above.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Results and discussion</title>
      <p>In this section, we analyse the performance of the architecture proposed in the evaluation
dataset delivered. It is worth stressing that, to facilitate the experiments’ reproducibility, we
utilised the default parameters in pre-trained transformed models in the experiments carried
out. The standardised metrics selected to assess the performance were Precision, Recall and
F1-measure. Table 3 collects the results achieved on each challenge’s classification task.</p>
      <p>The complete classification report of the target extraction was not included since it would
need a large table to add all the terms extracted during the evaluation. However, the accuracy
reached by the system was 0.61 (see in Table 4), indicating that the combination of the RoBERTa
model as NER with the designed grammatical rules obtained admissible outcomes. However,
despite this reasonable result, the remaining subtasks got weak scores such as 0.37, 0.47 and
0.48 in ‘target_sentiment’, ‘companies_sentiment’ and ‘consumer_sentiment’, respectively. We
think this low score in extracting the target is due to there are specific expressions in the news
headlines that grammar rules do not cover, making the extractor loses some entities during the
evaluation. Concerning the polarity detection, the task where the proposed system accomplished
the best outcome was in ‘target_sentiment’, harvesting the highest value on precision and recall,
Task</p>
      <p>target_sentiment
companies_sentiment
consumer_sentiment</p>
      <p>Label
negative
neutral
positive
Macro AVG
negative
neutral
positive
Macro AVG
negative
neutral
positive
Macro AVG
Precision</p>
      <p>Recall F1-score
on positive and neutral labelling, respectively. The worst results coincided with the task and
label, obtaining 0.1 and 0.17 on precision and recall. In spite of the overall results displaying
balanced values, which vary between 0.5 and 0.4, if we focus on the positive labelling results,
it is easy to recognize that they have a negative impact on the system’s performance since
recall outcomes drop below 0.3. However, if we look at the dataset’s distribution, the number
of samples for the positives across the tasks does not reflect the outcomes obtained. In some
tasks, the number of positive examples is higher, but the system’s performance is quite low. For
instance, in the “target_sentiment” task, in the positive labelling the system reached a recall
of 0.17, but it has assigned the highest amount of samples in the dataset. Consequently, this
situation reflects that the Transformer model selected for addressing the sentiment analysis
classification problem does not work precisely. We think this behaviour might be reduced by
using augmentation techniques, which sampling generation could teach better the model to
diferentiate between the three types of classes, positive, negative and neutral. For a more
detailed study of these low results, Figure 2 shows the confusion matrix of each classification
task.</p>
      <p>Confusion matrices allow an analysis deeply of the results. In this case, it is easy
quantifiable the mistakes conducted by the system during the classification process. As we can see,
the noticeable mistakes were positive to neutral and neutral to negative classifications. For
instance, on the task ‘targets_sentiment’, the system misclassified 495 samples, and 259 on
the ‘consumers_sentiment’. From here, it can be inferred that the system had grave problems
distinguishing positive and neutral polarities since it hesitates and makes significant mistakes.
The mistakes in predicting neutral cases are elevated if we examine and compare the three
confusion matrices. We believe that this behaviour is due to the value assigned to the threshold
for classifying neutral sentences was not precise enough since it seems that a high percentage of
the cases are classified as neutral, but they are not. It is worth mentioning that in positive versus
(a) Target Sentiment Task.</p>
      <p>(b) Companies Sentiment Task.</p>
      <p>(c) Consumers Sentiment Task
negative classifications in the pictures’ corner, the number of errors is less inflated, reducing
misclassified cases. Thus, despite the problems in classification-neutral samples, we think the
system knows how to diferentiate acceptably between positive and negative polarity. Despite
the several experiments conducted to find the best system’s performance, we can infer from
these results that the pre-established value of the threshold does not work correctly for any
case. Thus, we think each polarity tone has to be addressed diferently, assigning an individual
limit for each type. These analysed issues have made the proposed system only achieve the 7ℎ
place in the leaderboard, depicted in Table 4.</p>
      <p>Ranking</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions</title>
      <p>In this work we describe the contribution to the FinancES challenge, allocated in the IberLEF 2023
shared evaluation campaign of Natural Language Processing systems. The proposal combines
several Deep Learning models to address the main subtasks in this challenge, target extraction
and sentiment analysis. Thus, the system combined diferent models: BETO for classifying the
polarity, RoBERTa for extracting the target, and a Spacy model for grammar labelling.</p>
      <p>As we can see in the evaluation section, the gathered results show that there is still a range
of improvement. In the study of the resulting confusion matrices, it can be easily observed that
our system can “easily” detect negative labels in any of the proposed tasks but works worse in
detecting positive labels in any or neutral sentences on the ‘targets_sentiment’ issue. To improve
this situation, we suggest the addition of new grammar rules for detecting more structures
in sentences. On the other hand, as a future work line, we would like to try cutting-edge
Deep Learning strategies that do not require extensive weight training processes, like prompt
engineering.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Acknowledgments</title>
      <p>This work has been partially supported by projects DOTT-HEALTH (PID2019-106942RB-C32,
MCI/AEI/FEDER, UE), grant “Programa para la Recualificación del Sistema Universitario Español
2021-2023”, and the project M2297 from call 2022 for impulse projects funded by Rey Juan Carlos
University.
[11] C. Xiang, J. Zhang, F. Li, H. Fei, D. Ji, A semantic and syntactic enhanced neural model for
ifnancial sentiment analysis, Information Processing &amp; Management 59 (2022) 102943.
[12] E. Shijia, L. Yang, M. Zhang, Y. Xiang, Aspect-based financial sentiment analysis with
deep neural networks., in: WWW (Companion Volume), 2018, pp. 1951–1954.
[13] P. Ronghao, J. A. García-Díaz, F. García-Sánchez, R. Valencia-García, Evaluation of
transformer models for financial targeted sentiment analysis in spanish, PeerJ Computer Science
9 (2023) e1377. URL: https://doi.org/10.7717/peerj-cs.1377. doi:10.7717/peerj-cs.1377.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Almalis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kouloumpris</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Vlahavas</surname>
          </string-name>
          ,
          <article-title>Sector-level sentiment analysis with deep learning</article-title>
          ,
          <source>Knowledge-Based Systems</source>
          <volume>258</volume>
          (
          <year>2022</year>
          )
          <fpage>109954</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <article-title>Financial sentiment analysis model utilizing knowledge-base and domainspecific representation</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>82</volume>
          (
          <year>2023</year>
          )
          <fpage>8899</fpage>
          -
          <lpage>8920</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Shang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>A lexicon enhanced collaborative network for targeted financial sentiment analysis</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>60</volume>
          (
          <year>2023</year>
          )
          <fpage>103187</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Naithani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. P.</given-names>
            <surname>Raiwani</surname>
          </string-name>
          ,
          <article-title>Realization of natural language processing and machine learning approaches for text-based sentiment analysis</article-title>
          ,
          <source>Expert Systems</source>
          (
          <year>2022</year>
          )
          <article-title>e13114</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chintalapudi</surname>
          </string-name>
          , G. Battineni,
          <string-name>
            <given-names>M. Di</given-names>
            <surname>Canio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Sagaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Amenta</surname>
          </string-name>
          ,
          <article-title>Text mining with sentiment analysis on seafarers' medical documents</article-title>
          ,
          <source>International Journal of Information Management Data Insights</source>
          <volume>1</volume>
          (
          <year>2021</year>
          )
          <fpage>100005</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Žitnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Blagus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bajec</surname>
          </string-name>
          ,
          <article-title>Target-level sentiment analysis for news articles</article-title>
          ,
          <source>KnowledgeBased Systems</source>
          <volume>249</volume>
          (
          <year>2022</year>
          )
          <fpage>108939</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>García-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Alcaráz</given-names>
            <surname>Mármol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Marín-Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Overview of FinancES 2023:
          <article-title>Financial Targeted Sentiment Analysis in Spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Montes-y Gómez, Overview of IberLEF 2023: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2023), co-located with the 39th Conference of the Spanish Society for Natural Language Processing (SEPLN 2023), CEURWS</article-title>
          .org,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Minaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Azimi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdolrashidi</surname>
          </string-name>
          , Deep-sentiment:
          <article-title>Sentiment analysis using ensemble of cnn and bi-lstm models</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>04206</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Renault</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis and machine learning in finance: a comparison of methods and models on one million messages</article-title>
          ,
          <source>Digital Finance</source>
          <volume>2</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>