<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>From Human to Scalable Annotation: Teaching LLMs to mimic experts labeling on Economic News Sentiment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Petrocelli</string-name>
          <email>michele.petrocelli@mef.gov.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Rollin</string-name>
          <email>andrea.rollin@mef.gov.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Berta</string-name>
          <email>matteo.berta@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Zafonte</string-name>
          <email>francesca.zafonte@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Monaco</string-name>
          <email>simone.monaco@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salvatore Lo Sardo</string-name>
          <email>salvatore.losardo@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Apiletti</string-name>
          <email>daniele.apiletti@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daria Scacciatelli</string-name>
          <email>dscacciatelli@sogei.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tania Cerquitelli</string-name>
          <email>tania.cerquitelli@polito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Guglielmo Marconi University</institution>
          ,
          <addr-line>Roma</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ministry of the Economy and Finance</institution>
          ,
          <addr-line>Roma</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Politecnico di Torino, Department of Control and Computer Engineering (DAUIN)</institution>
          ,
          <addr-line>Corso Castelfidardo, 34/d, 10138 Torino TO</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>6</volume>
      <fpage>165</fpage>
      <lpage>190</lpage>
      <abstract>
        <p>Economic news sentiment ofers a timely window into how the public perceives current economic conditions. These perceptions shape expectations, economic behavior, financial markets, and policy decisions, making news-based sentiment a valuable proxy for tracking and anticipating broader economic trends. Most existing approaches to economic news sentiment analysis rely on automated or weakly supervised labeling strategies to ensure scalability. However, the reliance on such training data, together with a strong dependence on Englishcentric sentiment lexicons, limits their ability to capture the heterogeneous expression of sentiment across economic subdomains in the Italian language. To address these limitations, the proposed work introduces a human-annotated corpus of Italian economic news, provides a systematic comparison between long-context encoder-decoder models and instruction-tuned language models, and develops an instruction-tuned model adapted via LoRA that achieves strong agreement with expert annotations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Economic Sentiment</kwd>
        <kwd>Text Representation</kwd>
        <kwd>News Annotation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Sentiment analysis of economic texts has become an increasingly important tool for understanding
macroeconomic dynamics and market behavior. Media coverage of economic events captures
realtime changes in narratives, expectations, and perceptions, ofering valuable qualitative information
that complements traditional quantitative indicators [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. By analyzing how economic conditions are
discussed and framed in news articles, text-based approaches can provide timely insights into evolving
economic trends [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Traditional economic indicators remain fundamental, but they are typically released with substantial
delays [
        <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
        ]. As a result, they may fail to capture rapid changes and are often slow to reflect the impact
of exogenous shocks and unforeseen events. In this context, economic news sentiment provides a
timely and complementary source of information, enabling the detection of emerging trends as they
unfold. In particular, sentiment analysis can capture immediate shifts in expectations and reactions
to unexpected developments, ofering early signals of turning points that are not yet observable in
conventional macroeconomic indicators.
      </p>
      <p>
        Despite its relevance, measuring economic sentiment from text remains challenging. Conventional
approaches to sentiment analysis in economics often rely on lexicon-based methods or shallow models
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which struggle to capture the complexity, ambiguity, and context-dependence of economic language.
These limitations are particularly pronounced in specialized domains and for languages other than
English, where domain-specific lexical resources and high-quality annotated datasets remain scarce [ 5].
      </p>
      <p>To scale sentiment analysis to large corpora, many existing studies adopt automated or weakly
supervised labeling strategies. While eficient, these approaches typically rely on generic or noisy
labels and impose strong assumptions on sentiment expression, limiting their ability to reflect nuanced
sentiment variations across economic subtopics. As a consequence, the quality and reliability of
sentiment annotations remain a key bottleneck for downstream economic analysis.</p>
      <p>Recent advances in transformer-based architectures have substantially improved the modeling of
long and complex texts, while also increasing computational and memory requirements [6, 7]. Large
language models benefit from extended contextual representations and strong performance across many
language understanding tasks, but their practical adoption is often constrained in domain-specific and
low-resource settings.</p>
      <p>An alternative strategy is to adapt smaller transformer models to specific tasks using
parametereficient fine-tuning methods such as Low-Rank Adaptation (LoRA) [ 8]. By updating only a limited
number of parameters while keeping the backbone model fixed, these techniques reduce training
costs and memory usage, enabling the specialization of open-source models for economic text analysis
without full retraining.</p>
      <p>Beyond scalability, an important challenge is whether language models can reproduce expert human
annotation of economic news.</p>
      <p>For tasks involving long input sequences, architectures explicitly designed for long documents
provide an additional advantage. Longformer-based encoder–decoder models [9] enable the processing
of entire news articles without truncation, preserving access to long-range contextual information.
When combined with parameter-eficient fine-tuning, such models remain usable on standard hardware.</p>
      <p>In this work, we study whether annotations in Italian economic news can be replicated using
transformer-based models. We consider two complementary modeling strategies: (i) long-context
encoder–decoder architectures and (ii) instruction-based language models adapted through
parametereficient fine-tuning. The evaluation focuses on agreement with human annotators, the impact of
diferent label aggregation strategies, and generalization across heterogeneous types of economic news
(Figure 1).</p>
      <p>Our analysis is conducted on a human-annotated corpus of approximately 800 Italian news articles
published between 2007 and 2024. Articles are annotated for sentiment on a five-point ordinal scale
and assigned to one of six economic topics. The results show that instruction-tuned language models
achieve strong agreement with both individual and aggregated human annotations. In contrast, the
Longformer Encoder–Decoder (LED) model [9] exhibits a systematic tendency toward regression to the
mean, limiting its ability to capture variations. These findings indicate that, in low-data settings, while
long-context models are efective at processing extended documents, smaller instruction-based models
adapted through parameter-eficient techniques ofer more robust performance for replicating expert
economic annotation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>This section outlines the role of textual sentiment in economic analysis (2.1), discusses the main
methodological approaches to sentiment extraction (2.2), and motivates the use of large language
models as tools for replicating expert sentiment and topic annotations in economic news (2.3).</p>
      <sec id="sec-2-1">
        <title>2.1. Economic Sentiment Analysis</title>
        <p>Traditional indicators provide objective measures of economic activity, but they present two significant
limitations: publication lags and limited forward-looking information. This temporal gap has motivated
researchers to explore the use of textual data, which ofers real-time availability and forward-looking
information. Previous studies have shown that economic sentiment extracted from textual sources
can serve as an efective proxy for anticipating key macroeconomic indicators, including GDP (Gross
Domestic Product) growth, unemployment, and inflation, thereby providing valuable information
for economic forecasting and policy analysis [5, 10, 11]. Text-based measures of sentiment capture
expectations, uncertainty, and perceived economic conditions that are often not immediately reflected
in traditional quantitative indicators [12, 13]. Recent econometric studies further support the predictive
value of textual information: text-enhanced factor models incorporating news content improve GDP
forecasting accuracy [14], while text-augmented VAR (Vector AutoRegression) and dynamic factor
models highlight the role of central bank communication in predicting macroeconomic outcomes [15].</p>
        <p>A parallel research efort has been devoted to define sentiment. Prior work has typically approached
sentiment from two complementary perspectives: (i) the evaluation of sentiment expressed in the
semantic content of the text, and (ii) the analysis of tone and emotional nuances conveyed through
linguistic framing and stylistic choices. While both perspectives have proven informative in diferent
contexts, this study adopts the second approach and focuses on sentiment as tone and emotional nuance
in economic news articles. This choice reflects the central role of media framing in shaping economic
narratives and public perceptions, independently of the underlying economic facts reported.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Lexicon-Based Approaches versus Deep Learning Methods</title>
        <p>
          Lexicon-based approaches represent one of the most widely used methods for measuring economic
sentiment. These approaches rely on the construction of domain-specific dictionaries designed to
capture recurrent linguistic patterns in economic and financial texts. Prominent examples include
the news-based sentiment measures proposed by Shapiro et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the financial sentiment dictionary
developed by Loughran and McDonald [16], and related dictionary-based methods applied to news and
market analysis [12, 17].
        </p>
        <p>Lexicon-based methods are attractive due to their simplicity, transparency, long-standing use in
the literature, and low computational cost [18]. However, these approaches present well-documented
limitations, including strong dependence on domain and language specific vocabularies, limited ability
to account for context, negation, and semantic composition, and reduced efectiveness when applied to
complex or ambiguous texts [19, 20, 21]. These limitations are particularly pronounced for specialized
domains and for languages diferent from English, where domain-specific lexical resources and annotated
datasets remain scarce [5].</p>
        <p>In response to these limitations, deep learning methods based on transformer architectures have been
increasingly adopted. The literature continues to debate whether traditional lexicon-based approaches
or more recent neural models provide superior performance [22, 23]. While lexicon-based models
remain competitive in some settings, transformer-based architectures are generally better suited to
capture complex linguistic patterns and contextual dependencies when trained on high-quality data.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Large Language Models as Annotators</title>
        <p>Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in sentiment
analysis, achieving robust zero-shot and few-shot performance across diverse domains. Their ability
to follow natural language instructions and leverage broad linguistic and contextual knowledge has
made them promising candidates for tasks traditionally requiring expert human judgment, such as text
annotation. However, these strengths do not extend uniformly to more complex sentiment analysis
settings: as shown in [24] and [25], LLMs underperform smaller, task-specific supervised models on
tasks involving ambiguity or fine-grained sentiment distinctions, underscoring the need for careful
evaluation of LLM-generated annotations.</p>
        <p>
          In parallel, a growing literature has explored the use of LLMs as annotators or evaluators for tasks
traditionally performed by humans. Several studies show that, with clear guidelines and appropriate
prompting, LLMs can approximate or match crowdworker-level performance in classification and span
annotation tasks, ofering substantial gains in scalability and cost eficiency [ 26], [
          <xref ref-type="bibr" rid="ref5">27</xref>
          ]. Nonetheless,
alignment with human judgments remains imperfect, with LLMs typically achieving only moderate
agreement and exhibiting sensitivity to prompt design, task structure, and annotation granularity.
        </p>
        <p>
          Moreover, recent work highlights that LLMs inherit and may amplify systematic biases, including
verbosity, authority, and aesthetic biases, raising concerns about their reliability as judges [
          <xref ref-type="bibr" rid="ref6">28</xref>
          ], [
          <xref ref-type="bibr" rid="ref7">29</xref>
          ].
While LLMs can correlate well with average human judgments, they struggle to capture annotator
disagreement and demographic heterogeneity in subjective tasks where disagreement is informative rather
than noise [
          <xref ref-type="bibr" rid="ref8">30</xref>
          ]. Accordingly, several studies caution against treating LLMs as drop-in replacements for
expert annotators and advocate agreement-based, statistically grounded evaluation frameworks rather
than accuracy against majority labels [
          <xref ref-type="bibr" rid="ref9">31</xref>
          ], [25].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>This section describes the construction of the annotated dataset, the human annotation protocol and
agreement assessment (3.1), and the training and evaluation procedures adopted to compare
longdocument encoder–decoder models and instruction-tuned large language models in replicating humans
sentiment annotations (3.2).</p>
      <sec id="sec-3-1">
        <title>3.1. Human Annotation of Economic News</title>
        <p>The study is based on a corpus of Italian economic news articles collected from major national
newspapers through the Factiva database. A total of 800 articles were retrieved using a predefined set of
economy-related keywords that has been defined with a group of economic experts and are related
to several thematic areas (e.g, inflation, economic forecasting, social impact, politics and government
measures, central banks and geopolitics) by avoiding narrow or sentiment-driven selection.</p>
        <p>The corpus spans two distinct periods, 2007-2014 and 2017-2024, allowing the analysis to capture
diferent economic cycles and institutional contexts. To reduce potential temporal biases, articles were
balanced across publication years.</p>
        <p>Each article is treated as an independent observation and annotated at the document level, without
sentence-level or paragraph-level labeling. This choice reflects the objective of capturing the overall
tone and the dominant economic theme conveyed by each article.</p>
        <p>Sentiment is annotated on a five-point ordinal scale, where 1 indicates very negative sentiment and 5
indicates very positive sentiment. Annotators are instructed to focus on the tone and emotional nuances
of each article rather than its factual economic content.</p>
        <p>To promote consistency across annotators, detailed written guidelines are provided prior to the
annotation task. These guidelines define sentiment levels and include examples illustrating diferent
cases. In instances of ambiguity, annotators are instructed to select the closest sentiment based on the
overall emphasis and framing of the article or to assign a neutral label when sentiment is mixed.</p>
        <p>
          Inter-annotator agreement is computed on the subset of articles annotated by all seven raters using
Fleiss’  [
          <xref ref-type="bibr" rid="ref10">32</xref>
          ],
 =
        </p>
        <p>1¯ − − ¯¯ ,
.</p>
        <p>The statistic  measures the degree of agreement among annotators beyond what would be expected
by chance. Here, ¯  denotes the average observed agreement across items, while ¯  represents the
expected agreement under random assignment of categories. The ratio normalizes observed agreement
by its maximum possible value above chance, yielding  = 1 for perfect agreement,  = 0 for
chancelevel agreement and  &lt; 0 for disagreement. For this reason it estimates provide an indicator of the
consistency and reliability of the annotation process and serve as a reference point for interpreting
subsequent model performance.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model Training and Evaluation</title>
        <p>We selected three transformer-based architectures for sentiment annotation of Italian economic news:
one long-context encoder–decoder model and two large language models. Their difering architectures,
context windows, and training strategies enable a comparative assessment of suitability.</p>
        <p>The first architecture is a Longformer Encoder–Decoder (LED) model [ 9], included as a
representative long-context encoder–decoder baseline to assess whether access to full document context alone
is suficient to replicate expert annotation in a limited-data setting. LED model follows a supervised
ifne-tuning paradigm and is trained directly on the annotated dataset. Its architecture enables the
processing of entire news articles without truncation, which is particularly relevant for economic texts
where contextual information may be distributed across long documents.</p>
        <p>
          The final two models considered are Gemma-2-9B-IT 1 [
          <xref ref-type="bibr" rid="ref11">33</xref>
          ] and Phi-3.5 Mini Instruct2 [
          <xref ref-type="bibr" rid="ref12">34</xref>
          ], both
instruction-tuned large language models. While both models adopt a decoder-only Transformer
architecture with grouped-query attention, they difer in parameter scale and context capacity, with
Gemma-2-9B-IT featuring a larger parameter count and Phi-3.5 Mini Instruct designed for higher
eficiency and extended context handling. These models were selected based on their strong
performance in Italian language understanding tasks, as documented by the Evalita-LLM benchmark [
          <xref ref-type="bibr" rid="ref13">35</xref>
          ]3.
Evalita-LLM provides a systematic evaluation framework for Italian NLP, comparing a wide range of
open-source LLMs across ten tasks spanning both textual and multimodal settings. The benchmark
results support the selection of these models as representative state-of-the-art systems for Italian text
processing.
        </p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Longform Encoder-Decoder Training</title>
          <p>The chosen backbone is allenai/led-base-16384 because it supports sparse attention and is widely
used for long-document modeling [9]. The maximum input length is set to 4096 tokens to process entire
article. Global attention is restricted to the first token, which provides a document-level aggregation
mechanism in Longformer-based models and preserves computational eficiency for long sequences.</p>
          <p>We evaluate multiple loss functions because the sentiment labels are ordinal and diferent objectives
impose diferent inductive biases. In particular, we consider: MSE as a regression baseline, modeling
sentiment as a continuous variable bounded to the annotation scale; Cross-Entropy, which treats
sentiment prediction as a multi-class classification task; Soft Cross-Entropy which extends cross-entropy
to soft targets to capture annotation uncertainty; Coral, which formulates sentiment prediction as an
ordinal regression problem, explicitly modeling the ordered structure of labels and Distribution, which
1https://huggingface.co/google/Gemma-2-9B-IT
2https://huggingface.co/microsoft/Phi-3.5-mini-instruct
3https://huggingface.co/spaces/evalitahf/evalita_llm_leaderboard
models sentiment as a discrete probability distribution and penalizes prediction errors according to
their ordinal distance.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Instruction-Tuned LLM Training</title>
          <p>The two large language models, Gemma-2-9B-IT and Phi-3.5 Mini Instruct, are fine-tuned using a
supervised instruction-following setup. Training examples are formatted as instruction-completion
pairs, where the model is prompted to annotate a full Italian economic news article with a sentiment
score. The instruction explicitly defines sentiment as tone rather than factual content and constrains
the output format to a valid JSON object, ensuring structured and machine-readable predictions. (see
Appendix A).</p>
          <p>Fine-tuning is performed using Low-Rank Adaptation (LoRA) [8], which enables eficient
specialization of large models by updating a limited set of low-rank matrices while keeping the backbone
parameters frozen. LoRA adapters are applied to the attention and feed-forward projection layers,
automatically inferred from the model architecture. This strategy substantially reduces memory usage
and training cost, allowing fine-tuning on standard hardware while preserving the representational
capacity of the base models.</p>
          <p>The subset of articles annotated by all human annotators is reserved exclusively for evaluation,
serving as a common benchmark for assessing agreement between human annotators themselves and
between models and human annotators. These articles are not used during training or validation.</p>
          <p>The remaining articles, each annotated by a single annotator, are used for model training. For
instruction-tuned large language models fine-tuned with LoRA, all available training articles (750
documents) are used for supervised fine-tuning, without an explicit validation split. Model selection
and early stopping are based on loss computed on the held-out evaluation set.</p>
          <p>For the Longformer Encoder-Decoder model, the 750 training articles are further split into training
and validation sets using an 85/15 split. The validation set is used for early stopping and hyperparameter
selection, while the final evaluation is conducted on the 50 multi-annotated articles. This split reflects
the diferent training paradigms of the two approaches and ensures a fair comparison on a shared
evaluation set.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Preliminary Evaluation Results</title>
      <p>
        Results are primarily reported using agreement-based metrics, which are well suited to capturing
alignment with human judgments, while traditional error and accuracy measures are reported as
complementary indicators of model performance [
        <xref ref-type="bibr" rid="ref9">31</xref>
        ].
      </p>
      <sec id="sec-4-1">
        <title>4.1. Dataset</title>
        <p>Human annotation was conducted by seven annotators. To assess the reliability of the annotation
scheme, a subset of 50 articles was annotated by all annotators, enabling the computation of
interannotator agreement. The remaining articles were annotated by a single annotator and subsequently
used for model training, balancing the need for agreement assessment with dataset scalability. For
modeling purposes, sentiment labels are treated using label encoding that preserves the ordinal ordering
of the scale.</p>
        <p>The results, visible in the first line of Table 3, highlight the intrinsic dificulty of the sentiment
annotation task. Agreement across all annotators is relatively low for sentiment, with  = 0.24 ,
reflecting the subjective and nuanced nature of tone-based sentiment assessment. When considering
pairwise agreement, the highest concordance is observed for selected annotator pairs, with  = 0.41.</p>
        <p>
          Following Fleiss [
          <xref ref-type="bibr" rid="ref10">32</xref>
          ],  values are interpreted relative to chance agreement, with larger values
indicating stronger agreement. For qualitative interpretation, we additionally refer to commonly used
benchmarks proposed by Landis and Koch [
          <xref ref-type="bibr" rid="ref14">36</xref>
          ], according to which values between 0.21 and 0.40
correspond to fair agreement, values between 0.41 and 0.60 to moderate agreement, and values above
0.60 to substantial agreement or perfect agreement.
        </p>
        <p>Each article in our dataset is also annotated with exactly one topic label corresponding to the most
salient economic theme discussed. The predefined topic categories, Monetary Policy and Central
Banks, Financial Markets, Inflation and Prices, Fiscal Policies and Taxes, Labor Market, and Geopolitics
and Social Impact, were defined in consultation with economic experts to reflect standard thematic
distinctions in economic analysis.</p>
        <p>It is interesting to see in Figure 2 and Figure 3 that expert sentiment annotations exhibit substantial
dispersion across both economic topics and information sources, despite being measured on a discrete
ifve-point scale. While most annotations cluster around negative, neutral, and positive categories,
the presence of variation within each topic and source highlights the inherently subjective nature
of sentiment assessment in economic text. Moreover, the similarity of the sentiment distributions
across topics and sources suggests that heterogeneity in judgments is not driven by a single thematic
dimension or outlet, but rather reflects systematic diferences in interpretation among annotators.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Performance evaluation</title>
        <sec id="sec-4-2-1">
          <title>4.2.1. Longformer Encoder–Decoder</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Instruction-Tuned Large Language Models</title>
          <p>reduction in MSE. These results demonstrate that parameter-eficient fine-tuning efectively aligns
instruction-tuned large language models with the annotation guidelines and the linguistic characteristics
of Italian economic news.</p>
          <p>To better characterize the error patterns of the models, we also report confusion matrices (Figure 4).
The labels predicted by the model are compared with labels derived from those assigned by humans:
the sentiment target is the average human score rounded to the nearest class.</p>
          <p>
            To further assess alignment with human judgments, we compute pairwise Cohen’s  [
            <xref ref-type="bibr" rid="ref10">32</xref>
            ] between
model predictions and human annotations for both Gemma-2-9B-IT and Phi-3.5 Mini Instruct
(Table 3). For Gemma, LoRA fine-tuning substantially improves agreement with annotators for sentiment
classification.
          </p>
          <p>In contrast, Phi-3.5 Mini Instruct exhibits limited or negative efects from LoRA fine-tuning:
sentiment agreement improves only marginally.</p>
          <p>Human–human agreement provides important context for interpreting these results. Topic annotation
shows high consistency across annotators ( ≈ 0.49 ), whereas sentiment agreement is substantially
lower ( ≈ 0.24 ), reflecting the inherently subjective nature of tone-based sentiment annotation. Within
this context, sentiment performance for all models remains bounded by annotator disagreement.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussions</title>
      <p>This study investigates whether sentiment annotations in Italian economic news can be reliably
replicated by transformer-based models, with a particular focus on comparing long-context encoder-decoder
architectures and instruction-tuned large language models adapted through parameter-eficient
finetuning.</p>
      <p>The Longformer Encoder-Decoder model demonstrates the ability to process entire news articles
without truncation and achieves competitive aggregate error metrics for sentiment prediction under
several loss functions. However, a detailed analysis reveals a systematic regression-to-the-mean behavior
across all training configurations. This indicates that optimizing for aggregate error alone is insuficient
to capture the nuanced and subjective nature of economic sentiment, particularly when inter-annotator
agreement is limited and the training data are relatively small.</p>
      <p>In contrast, instruction-tuned large language models exhibit greater flexibility in replicating human
annotation behavior. LoRA fine-tuning substantially improves sentiment accuracy for
Gemma-2-9BIT. These results suggest that instruction-following pretraining, combined with parameter-eficient
adaptation, provides a strong inductive bias for modeling evaluative tone in economic texts. At the
same time, the more heterogeneous behavior observed for Phi-3.5 Mini Instruct highlights that the
efectiveness of LoRA depends on model capacity and alignment with the task-specific taxonomy.</p>
      <p>A central aspect of this work is the use of human–human agreement as a benchmark for evaluating
model performance. Sentiment annotation is inherently subjective and model outputs should therefore
be interpreted relative to this level of human consistency rather than in absolute terms. The results also
highlight the importance of high-quality human-annotated data and well-defined annotation guidelines:
in their absence, even powerful model architectures may converge to trivial or weakly informative
solutions.</p>
      <p>Future work includes applying the proposed framework to larger corpora and to additional languages
and economic contexts in order to provide a more efective test of generalizability, exploring LoRA
finetuning across a broader range of instruction-tuned models with diferent architectures and parameter
scales, and investigating uncertainty-aware annotation schemes, such as soft labels, distributions, or
confidence scores, that could better capture the subjective nature of sentiment annotation. Further
research should also evaluate commercial large language models and, considering both zero-shot and
few-shot settings, compare proprietary systems with open-source fine-tuned models to clarify trade-ofs
in terms of transparency, cost, and annotation quality. Finally, integrating these methodologies into
forecasting models of macroeconomic target variables may enhance their predictive performance,
particularly by incorporating timely sentiment-based signals alongside traditional indicators.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We would like to sincerely thank all the expert annotators who contributed to the creation of the dataset.
Their domain expertise, careful and consistent annotations, and the numerous fruitful discussions
throughout the annotation process were essential to ensuring the quality, reliability, and relevance of
the data.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used GPT 5.2 in order to: grammar, spelling check
and help in the correction of sentence structure. After using these tool, the author reviewed and edited
the content as needed and takes full responsibility for the publication’s content.
[5] S. R. Baker, N. Bloom, S. J. Davis, Measuring economic policy uncertainty, The Quarterly Journal
of Economics 131 (2016) 1593–1636.
[6] R. Thoppilan, et Al, Lamda: Language models for dialog applications, CoRR abs/2201.08239 (2022).</p>
      <p>URL: https://arxiv.org/abs/2201.08239. arXiv:2201.08239.
[7] J. W. Rae, et Al, Scaling language models: Methods, analysis &amp; insights from training gopher,</p>
      <p>CoRR abs/2112.11446 (2021). URL: https://arxiv.org/abs/2112.11446. arXiv:2112.11446.
[8] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, Lora: Low-rank
adaptation of large language models, in: Proceedings of the International Conference on Learning
Representations (ICLR), 2022.
[9] I. Beltagy, M. E. Peters, A. Cohan, Longformer: The long-document transformer, CoRR
abs/2004.05150 (2020). URL: https://arxiv.org/abs/2004.05150. arXiv:2004.05150.
[10] S. A. Sharpe, N. R. Sinha, C. A. Hollrah, The power of narrative sentiment in economic
forecasts, International Journal of Forecasting 39 (2023) 1097–1121. URL: https://www.sciencedirect.
com/science/article/pii/S0169207022000590. doi:https://doi.org/10.1016/j.ijforecast.
2022.04.008.
[11] M. Gentzkow, B. Kelly, M. Taddy, Text as data, Journal of Economic Literature 57 (2019) 535–74.</p>
      <p>URL: https://www.aeaweb.org/articles?id=10.1257/jel.20181020. doi:10.1257/jel.20181020.
[12] P. C. Tetlock, Giving content to investor sentiment: The role of media in the stock market, The</p>
      <p>Journal of Finance 62 (2007) 1139–1168.
[13] P. C. Tetlock, M. Saar-Tsechansky, S. Macskassy, More than words: Quantifying language to
measure firms’ fundamentals, The Journal of Finance 63 (2008) 1437–1467.
[14] B. Seo, Econometric forecasting using ubiquitous news text: Text-enhanced factor model,
International Journal of Forecasting 41 (2025) 1055–1072. doi:10.1016/j.ijforecast.2024.11.001.
[15] L. N. Ferreira, Forecasting with VAR-teXt and DFM-teXt models: Exploring the predictive power
of central bank communication, Technical Report 559, Banco Central do Brasil, Working Paper
Series, 2021.
[16] T. Loughran, B. McDonald, When is a liability not a liability? textual analysis, dictionaries, and
10-ks, The Journal of Finance 66 (2011) 35–65.
[17] D. Garcia, Sentiment during recessions, The Journal of Finance 68 (2013) 1267–1300.
[18] P. J. Stone, D. C. Dunphy, M. S. Smith, The General Inquirer: A Computer Approach to Content</p>
      <p>Analysis, MIT Press, 1966.
[19] B. Pang, L. Lee, Opinion mining and sentiment analysis, Foundations and Trends in Information</p>
      <p>Retrieval 2 (2008) 1–135.
[20] M. Taboada, J. Brooke, M. Tofiloski, K. Voll, M. Stede, Lexicon-based methods for sentiment
analysis, Computational Linguistics 37 (2011) 267–307.
[21] L. Young, S. Soroka, Afective news: The automated coding of sentiment in political texts, Political</p>
      <p>Communication 29 (2012) 205–231.
[22] R. Catelli, S. Pelosi, M. Esposito, Lexicon-based vs. bert-based sentiment analysis: A comparative
study in italian, Electronics 11 (2022). URL: https://www.mdpi.com/2079-9292/11/3/374. doi:10.
3390/electronics11030374.
[23] E. Öhman, The validity of lexicon-based sentiment analysis in interdisciplinary research, in:
M. Hämäläinen, K. Alnajjar, N. Partanen, J. Rueter (Eds.), Proceedings of the Workshop on Natural
Language Processing for Digital Humanities, NLP Association of India (NLPAI), NIT Silchar, India,
2021, pp. 7–12. URL: https://aclanthology.org/2021.nlp4dh-1.2/.
[24] W. Zhang, Y. Deng, B. Liu, S. J. Pan, L. Bing, Sentiment analysis in the era of large language models:</p>
      <p>A reality check, 2023. URL: https://arxiv.org/abs/2305.15005. arXiv:2305.15005.
[25] C. Shen, L. Cheng, X.-P. Nguyen, Y. You, L. Bing, Large language models are not yet human-level
evaluators for abstractive summarization, in: H. Bouamor, J. Pino, K. Bali (Eds.), Findings of
the Association for Computational Linguistics: EMNLP 2023, Association for Computational
Linguistics, Singapore, 2023, pp. 4215–4233. URL: https://aclanthology.org/2023.findings-emnlp.
278/. doi:10.18653/v1/2023.findings-emnlp.278.
[26] X. He, et Al, AnnoLLM: Making large language models to be better crowdsourced
annotaNote: The prompt shown here is the English-translated version of the original Italian prompt used in the
experiments.</p>
      <p>Listing 1: Instruction prompt used for LLM fine-tuning
SYSTEM PROMPT
------------You are an expert annotator of Italian economic news articles.</p>
      <p>Your task is to analyze an article and produce the following annotation:</p>
      <p>Response format (mandatory):
{"Sentiment": &lt;number between 1 and 5&gt;}
No other text before or after the JSON.</p>
      <p>USER PROMPT
----------Analyze the following article and produce the required annotation.</p>
      <p>ARTICLE:
[Article text]
Respond with JSON only.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Giannone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Reichlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Small</surname>
          </string-name>
          ,
          <article-title>Nowcasting: The real-time informational content of macroeconomic data</article-title>
          ,
          <source>Journal of Monetary Economics</source>
          <volume>55</volume>
          (
          <year>2008</year>
          )
          <fpage>665</fpage>
          -
          <lpage>676</lpage>
          . URL: https://www. sciencedirect.com/science/article/pii/S0304393208000652. doi:https://doi.org/10.1016/j. jmoneco.
          <year>2008</year>
          .
          <volume>05</volume>
          .010.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Barbaglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Consoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Manzan</surname>
          </string-name>
          ,
          <article-title>Forecasting with economic news</article-title>
          ,
          <source>Journal of Business &amp; Economic Statistics</source>
          <volume>41</volume>
          (
          <year>2023</year>
          )
          <fpage>708</fpage>
          -
          <lpage>719</lpage>
          . URL: https://doi.org/10.1080/07350015.
          <year>2022</year>
          .
          <volume>2060988</volume>
          . doi:
          <volume>10</volume>
          .1080/07350015.
          <year>2022</year>
          .
          <volume>2060988</volume>
          (
          <issue>online</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Stock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. W.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <article-title>Macroeconomic forecasting using difusion indexes</article-title>
          ,
          <source>Journal of Business &amp; Economic Statistics</source>
          <volume>20</volume>
          (
          <year>2002</year>
          )
          <fpage>147</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Shapiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sudhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <article-title>Measuring news sentiment</article-title>
          ,
          <source>Journal of Econometrics</source>
          <volume>228</volume>
          (
          <year>2022</year>
          )
          <fpage>221</fpage>
          -
          <lpage>243</lpage>
          . URL: https://www.sciencedirect.com/science/article/pii/S0304407620303535. doi:https://doi.org/10.1016/j.jeconom.
          <year>2020</year>
          .
          <volume>07</volume>
          .053. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .naacl-industry.
          <volume>15</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kasner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zouhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schmidtová</surname>
          </string-name>
          , I. Kartáč,
          <string-name>
            <given-names>K.</given-names>
            <surname>Onderková</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Plátek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gkatzia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mahamood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dušek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Balloccu</surname>
          </string-name>
          ,
          <article-title>Llms as span annotators: A comparative study of llms and humans, 2025</article-title>
          . URL: https://arxiv.org/abs/2504.08697. arXiv:
          <volume>2504</volume>
          .
          <fpage>08697</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Humans or LLMs as the judge? a study on judgement bias</article-title>
          , in: Y.
          <string-name>
            <surname>Al-Onaizan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bansal</surname>
            ,
            <given-names>Y.-N.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Miami, Florida, USA,
          <year>2024</year>
          , pp.
          <fpage>8301</fpage>
          -
          <lpage>8327</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .emnlp-main.
          <volume>474</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .emnlp-main.
          <volume>474</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elangovan</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Bodapati</surname>
          </string-name>
          , D. Roth,
          <article-title>ConSiDERS-the-human evaluation framework: Rethinking human evaluation for generative large language models</article-title>
          , in: L.
          <string-name>
            <surname>-W. Ku</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
          </string-name>
          , V. Srikumar (Eds.),
          <source>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Bangkok, Thailand,
          <year>2024</year>
          , pp.
          <fpage>1137</fpage>
          -
          <lpage>1160</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>63</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>63</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zouhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rooein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hoyle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sachan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leippold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          , E. Ash,
          <article-title>Can reasoning help large language models capture human annotator disagreement?</article-title>
          ,
          <year>2026</year>
          . URL: https: //arxiv.org/abs/2506.19467. arXiv:
          <volume>2506</volume>
          .
          <fpage>19467</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>N.</given-names>
            <surname>Calderon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reichart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dror</surname>
          </string-name>
          ,
          <article-title>The alternative annotator test for LLM-as-a-judge: How to statistically justify replacing human annotators with LLMs</article-title>
          , in: W. Che,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nabende</surname>
          </string-name>
          , E. Shutova, M. T. Pilehvar (Eds.),
          <source>Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Vienna, Austria,
          <year>2025</year>
          , pp.
          <fpage>16051</fpage>
          -
          <lpage>16081</lpage>
          . URL: https://aclanthology.org/
          <year>2025</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>782</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2025</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>782</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Fleiss</surname>
          </string-name>
          ,
          <article-title>Measuring nominal scale agreement among many raters</article-title>
          ,
          <source>Psychological Bulletin</source>
          <volume>76</volume>
          (
          <year>1971</year>
          )
          <fpage>378</fpage>
          -
          <lpage>382</lpage>
          . doi:
          <volume>10</volume>
          .1037/h0031619.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>G.</given-names>
            <surname>Team</surname>
          </string-name>
          , Gemma:
          <article-title>Open models based on gemini research and technology</article-title>
          ,
          <source>arXiv preprint arXiv:2403.08295</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdin</surname>
          </string-name>
          , et al.,
          <article-title>Phi-3 technical report: A highly capable language model locally on your phone</article-title>
          ,
          <source>arXiv preprint arXiv:2404.14219</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zanoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Resta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cimmino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Madeddu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          , Evalita-llm:
          <article-title>Benchmarking large language models on italian</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2502.02289. arXiv:
          <volume>2502</volume>
          .
          <fpage>02289</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Landis</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. G. Koch,</surname>
          </string-name>
          <article-title>The measurement of observer agreement for categorical data</article-title>
          ,
          <source>Biometrics</source>
          <volume>33</volume>
          (
          <year>1977</year>
          )
          <fpage>159</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>