<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Jaén, Spain
* Corresponding author.
$ sjzafra@ujaen.es (S. M. Jiménez-Zafra); daniel.gbaena@gmail.com (D. García-Baena); magc@ujaen.es
(M. García-Cumbreras); mgarcia@ujaen.es (M. García-Vega)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>SINAI at FinancES@IberLEF2023: Evaluating Popular Tools and Transformers Models for Financial Target Detection and Sentiment Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Salud María Jiménez-Zafra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel García-Baena</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel Ángel García-Cumbreras</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel García-Vega</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, SINAI, CEATIC, Universidad de Jaén</institution>
          ,
          <addr-line>23071</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>This work presents the participation of the SINAI team at FinancES@IberLEF2023 shared task, Financial Targeted Sentiment Analysis in Spanish. We have addressed the two proposed tasks, consisting on identifying the main economic target from headlines of financial news for determining their sentiment polarity and identifying the sentiment polarity of each news headline towards both companies and consumers. For target detection, we have explored some popular tools as Stanza and spaCy, and diferent transformers models from Hugging Face and ChatGPT4. For sentiment analysis, we have evaluated some of the most popular transformers models and specific financial transformers. In total, 11 systems have participated (including the baseline provided by the organizers). The best run sent by our team have been placed in position 4th for Task1 and position 2nd for Task 2 with an F1-score of 0.7780 and 0.6349, respectively, being 0.7922 and 0.6423 the best results obtained in the competition for both tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ifnancial target detection</kwd>
        <kwd>sentiment analysis</kwd>
        <kwd>financial multi-dimensional sentiment classification</kwd>
        <kwd>machine translation</kwd>
        <kwd>transformers</kwd>
        <kwd>natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        information worldwide, including, of course, financial literacy, more people started to manifest
their interest in economics. Nowadays, it is much easier to find financial information posted
online and, therefore, it is possible to monitor public information, receive early warnings and
perform positive and negative impact analysis. The efects of emotions on financial markets have
been demonstrated in several studies [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. In any case, there are many diferent factors to take
into account when we try to evaluate the efectiveness of sentiment analysis when we work in
a financial context. In the first place, some complex vocabulary frequently populates economic
texts, underlying social, and legal context [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. On the other hand, in this domain, language
is more related to circumstances and every word or expression may have either positive or
negative connotations depending on the context and subjectivity is always present because
texts are written according to the point of view, and/or the interests, of the author.
      </p>
      <p>
        Specifically, our team has participated in both of the tasks of this shared competition. For
Task 1, and with the objective of identifying the main economic target from texts, we evaluated
some strategies based on diferent transformers, ChatGPT4 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the Python natural language
analysis package Stanza [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and the popular NLP tool spaCy [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In relation to Task 2, our
team classified texts from headlines determining their sentiment polarity with the aid of, again,
transformers-based models from Hugging Face.
      </p>
      <p>Finally, this paper is divided into six diferent sections. Immediately after this introduction, we
will talk about everything related with the task description, detailing its purpose and reviewing
the dataset. Methodology section will explain what we did for generating our results and in
Experimental setup, we discuss about the tools and how they were set up for conducting the
experiments. On the other hand, Results and discussion summarizes the results that we obtained
in each experiment and reviews them in a comparative way. The last section of this work is
Conclusions and future work, and there we discern about our outcomes and, consequently,
point to future works that would ofer better results.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description</title>
      <p>This shared task comprises two subtasks. The first one is for target detection and here,
participants have to identify the economic target in newspaper headlines. After identifying the main
economic target from headlines of financial news, teams have to classify the sentiment polarity
(positive, neutral or negative) towards such target in the processed text. For the evaluation, the
systems are ranked using the arithmetic mean of the target F1-score and sentiment classification
macro-F1. Regarding the second task, participants are expected to conduct a multi-dimensional
sentiment classification. As opposed to traditional multi-target tasks, in which multiple targets
are identified within the scope of each individual processed text, here each news headline
refers to a single target entity, but the stances of other economic agents, companies (com) and
consumers (con), are also considered. For Task 2, the systems are ranked using the arithmetic
mean of the macro-F1-com and macro-F1-con.</p>
      <p>
        In relation to the dataset [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], it is an extension of the dataset published in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. It is
composed of news headlines written in Spanish collected from digital newspapers specialized in
economic, financial and political news as: Expansión, El Economista, Modaes or El Financiero.
It is important to highlight that no all the newspapers are from the same Spanish-speaking
Sentiment target companies consumer
      </p>
      <p>Total headlines
country. The creators of the dataset reviewed all headlines removing the irrelevant ones and
manually labelled each headline with the target entity and the sentiment polarity on three
dimensions: target, companies, and consumers. Three options are available for sentiment
analysis: positive, neutral, o negative. The final dataset comprises of 7,618 news headlines.
For the shared task, at a first stage, development-training and development-test sets were
made available for participants to develop their systems. Later, training set (including the
data of the development set) and test sets were released to participate in the shared task.
In Table 1, it is presented the dataset distribution. Finally, it is worth mentioning that the
competition was organized through CodaLab and can be accessed at the following link: https:
//codalab.lisn.upsaclay.fr/competitions/10052#learn_the_details.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        We evaluated several technologies for tasks 1 and 2. For the first part of Task 1, consisting in
identifying the main target of each headline in the dataset, we tested three diferent approaches,
one based on some popular tools, such as Stanza [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and spaCy [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], in this last case using the
Spanish pipeline optimized for CPU model: es_core_news_lg, other on diferent transformer
models from Hugging Face and the last one using the popular large language model ChatGPT4
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. ChatGPT4 was tested with diferent prompts and the one that worked best for extracting
the target entities was: “Dime cuál es la entidad objetivo en la siguiente oración, sólo la entidad,
sin punto final, manteniendo su forma de aparición en el texto”/ Tell me what is the target entity
in the following sentence, just the entity, without a period, keeping its form of appearance in the
text. All the results obtained during the development phase are shown in Table 2. As can be
seen, transformers-based models from Hugging Face obtained the best results over the rest
of the options. Concretely, Babelscape/wikineural-multilingual-ner model obtained the best
F1-score with a value of 0.8005.
      </p>
      <p>In relation to sentiment analysis, we tested diferent transformer-based models with the
original Spanish dataset and, for the English models, with an English-translated version of the
original Spanish corpus using Google Translator from Python’s deep_translator library [13].
After checking more than ten alternative configurations, based on finance-related and popular
transformers models, for the second part of Task 1, related to sentiment analysis on the main
target of news headlines, we obtained the best results with mDeBERTa and RoBERTuito and
the original, non translated, Spanish dataset. On the other hand, for Task 2, on identifying
sentiments for companies and consumers, the best results were obtained using MarIA and the
original Spanish dataset. All F1 scores obtained in the development phase can be consulted in
Table 3.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental setup</title>
      <p>Regarding to the software we used to translate the dataset from Spanish to English, it was
Google Translator from Python’s deep_translator library [13]. In addition, it is important to
note that we did not perform any prior data pre-processing on the dataset to perform the
experiments.</p>
      <p>With respect to the models, all were downloaded from their public profiles in Hugging
Face. During the finetuning process we always used Google Colab for coding under a Pro
configuration for being able to use their GPU based hardware options.</p>
      <p>Finally, concerning the hyperparameters, Table 4, Table 5 and Table 6 show the configurations
that provided the best results for each of the models in the tasks sentiment analysis for main
targets, sentiment analysis for companies and sentiment analysis for consumers, respectively.
For target detection task, we used default parameters.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results and discussion</title>
      <p>
        This section presents the results obtained in the evaluation phase of the shared task FinancES [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
Financial Targeted Sentiment Analysis in Spanish, at IberLEF 2023. The organizers selected the
arithmetic mean of the target F1-score and the target sentiment F1-score for ranking the systems
in Task 1: Financial targeted sentiment analysis, and the arithmetic mean of the macro-F1-com
and macro-F1-con for ranking the systems in Task 2: Financial Sentiment Analysis at document
level for companies and consumers. Each participating team could submit a maximum of 10 runs
through CodaLab, from which each team had to select the best one for ranking. We selected our
10 runs based on the experiments carried out on the training phase. The results for each of the
runs are shown in Table 7 and the models used for target detection and sentiment classification
in each of them are displayed in Table 8. The best results for each of the measures are marked
in bold and the run that provided the best performance is highlighted with a gray background.
      </p>
      <p>For the first part of Task 1, concerning the identification of the main economic target from
ifnancial news headlines, the transformer-based models
(Babelscape/wikineural-multilingualner and mrm8488/bert-spanish-cased-finetuned-ner) outperformed ChatGPT4, Stanza and the
es_core_news_lg model of spaCy. It is in this task where we appreciate the biggest diferences
between our systems and the best results are always for NER specific models that were developed
run
using some transformer approach.</p>
      <p>Continuing with the first task, but now in relation to the second part consisting of determining
the sentiment (positive, neutral or negative) towards the main target in the news headlines, we
can see that the models using the original Spanish dataset performed better than those using
the translated corpus, with MarIA being the best performing model. It should be noted that the
English model ROBERTA works slightly better than the Spanish model BETO.</p>
      <p>Regarding the second task, on determining the sentiment polarity of each news headline
towards both companies and consumers, again the Spanish models overall obtain better results
than the English models. However, it is worth noting that the English models ROBERTA and
distilrobertafinancial work better than the Spanish models RoBERTuito and BETO. On this
occasion, MarIA and mDeBERTa are the transformers that provide the best results for consumers
and companies sentiment classification, respectively.</p>
      <p>Finally, highlight that the finance-specific transformers have performed worse than the
general transformers in the subtasks related to sentiment analysis.</p>
      <p>For the competition, we selected run 7, which uses the transformers
mrm8488/bert-spanishcased-finetuned-ner and mDeBERTa for target detection and sentiment classification, respectively.
With this approach we reached 4th position for Task 1: Financial targeted sentiment analysis
and 2nd position for Task 2: Financial Sentiment Analysis at document level for companies and
consumers. The oficial leaderboards for both tasks can be consulted in Table 9 and Table 10.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and future work</title>
      <p>In this paper we have presented the participation of the SINAI team in the shared task FinancES,
Financial Targeted Sentiment Analysis in Spanish, at IberLEF 2023. The objective of our
experiments, for the target detection task, was to test the performance of the most popular
NER tools against transformer-based models. The main conclusion is that transformers-based
solutions outperformed others as Stanza or spaCy related approaches. On the other hand, for
the sentiment analysis tasks, the aim of our experiments was to test how some of the most
popular transformers models behave compared to specific financial transformers models. In
abc111
LLI-UAM
ABCD Team
SINAI
AnkitSinghRaikuni
UTB-NLP
NLP_URJC
BASELINE
mario.pv
UNAM Text Mining
fanchuyi
this case we conclude that generic transformer models perform better than existing financial
transformers models.</p>
      <p>In the future, we plan to evaluate why financial transformers perform worse. In addition, we
want to continue evaluating external resources to further improve the training phase of the
system by analyzing the contribution of each model, testing diferent transfer learning systems
as well as models trained on general topics, and using diferent machine translation systems to
generate new datasets and/or augment existing ones.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by Project CONSENSO (PID2021-122263OB-C21),
Project MODERATES (TED2021-130145B-I00) and Project SocialTox (PDC2022-133146-C21)
funded by MCIN/AEI/10.13039/501100011033 and by the European Union
NextGenerationEU/PRTR, Project PRECOM (SUBV-00016) funded by the Ministry of Consumer Afairs
of the Spanish Government, Project FedDAP (PID2020-116118GA-I00) supported by
MICINN/AEI/10.13039/501100011033 and WeLee project (1380939, FEDER Andalucía 2014-2020) funded
by the Andalusian Regional Government. Salud María Jiménez-Zafra has been partially
supported by a grant from Fondo Social Europeo and the Administration of the Junta de Andalucía
(DOC_01073).
D. Zhu, X. Li, N. Qiang, D. Shen, T. Liu, B. Ge, Summary of ChatGPT/GPT-4 Research and
Perspective Towards the Future of Large Language Models, 2023. arXiv:2304.01852.
[13] Baccouri, Nidhal, A flexible free and unlimited python tool to translate
between diferent languages in a simple way using multiple translators
https://deep-translator.readthedocs.io/, 2020.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Montes-y Gómez, Overview of IberLEF 2023: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2023), co-located with the 39th Conference of the Spanish Society for Natural Language Processing (SEPLN 2023), CEURWS</article-title>
          .org,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          , Almela,
          <string-name>
            <given-names>F.</given-names>
            <surname>García-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Alcaráz</given-names>
            <surname>Mármol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Marín-Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Overview of FinancES 2023:
          <article-title>Financial Targeted Sentiment Analysis in Spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M. M. Hasan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Popp</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Oláh</surname>
          </string-name>
          ,
          <article-title>Current landscape and influence of big data on finance</article-title>
          ,
          <source>Journal of Big Data</source>
          <volume>7</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Goodell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <article-title>Emotions and stock market anomalies: A systematic review</article-title>
          ,
          <source>Journal of Behavioral and Experimental Finance</source>
          <volume>37</volume>
          (
          <year>2023</year>
          ). URL: https: //ideas.repec.org/a/eee/beexfi/v37y2023ics2214635022000557.html. doi:
          <volume>10</volume>
          .1016/j.jbef.
          <year>2022</year>
          .
          <volume>10072</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Nemes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kiss</surname>
          </string-name>
          ,
          <article-title>Prediction of stock values changes using sentiment analysis of stock news headlines</article-title>
          ,
          <source>Journal of Information and Telecommunication</source>
          <volume>5</volume>
          (
          <year>2021</year>
          )
          <fpage>375</fpage>
          -
          <lpage>394</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Milne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chisholm</surname>
          </string-name>
          ,
          <article-title>The Prospects for Common Financial Language in Wholesale Financial Services</article-title>
          , SWIFT Institute Working Paper,
          <string-name>
            <surname>SSRN</surname>
          </string-name>
          ,
          <year>2013</year>
          . URL: https://books.google. es/books?id=ZZEhzwEACAAJ.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>OpenAI.</surname>
          </string-name>
          (
          <year>2023</year>
          ). ChatGPT (May version) [
          <source>Large language model]</source>
          , https://chat.openai.com,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Stanza:
          <string-name>
            <given-names>A Python</given-names>
            <surname>Natural Language Processing Toolkit for Many Human Languages</surname>
          </string-name>
          ,
          <year>2020</year>
          . arXiv:
          <year>2003</year>
          .07082.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Montani</surname>
          </string-name>
          , spaCy 2:
          <article-title>Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing, 2017</article-title>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ronghao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>García-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Evaluation of transformer models for financial targeted sentiment analysis in Spanish</article-title>
          ,
          <source>PeerJ Computer Science</source>
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <article-title>e1377</article-title>
          . URL: https://doi.org/10.7717/peerj-cs.
          <volume>1377</volume>
          . doi:
          <volume>10</volume>
          .7717/peerj-cs.
          <volume>1377</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Salas</given-names>
            <surname>Zarate</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hernández-Alcaraz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Gómez</given-names>
            <surname>Berbis</surname>
          </string-name>
          ,
          <source>Machine Learning Based Sentiment Analysis on Spanish Financial Tweets</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>305</fpage>
          -
          <lpage>311</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -77703-0_
          <fpage>31</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          , T. Han,
          <string-name>
            <surname>S</surname>
          </string-name>
          . Ma, J. Zhang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>