<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Fine-Tuning a Compact Multimodal Model on Consumer-Grade Hardware at PastReader 2025</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yanco Amor Torterolo-Orta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marina Miguez-Lamanuzzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Nacional de Educación a Distancia (UNED)</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>This paper presents the experiments conducted as part of the IberLEF 2025 shared task PastReader: Transcribing Texts from the Past. A compact, open-source multimodal model was fine-tuned on the dataset provided by the organising team. The model was run on consumer-grade hardware, as one of our main goals is to democratise the use of AI while contributing to the preservation and digitisation of historical documents. The results were notable given the hardware limitations. Future work will focus on evaluating other models of similar size and exploring new techniques, with a special emphasis on leveraging the shared task dataset in collaboration with the Biblioteca Nacional de España (BNE).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;OCR</kwd>
        <kwd>Historical Documents</kwd>
        <kwd>Digital Humanities</kwd>
        <kwd>Multimodal OCR</kwd>
        <kwd>Fine-tuning</kwd>
        <kwd>Granite3</kwd>
        <kwd>2-vision</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The rest of this work is structured as follows: section 2 contextualises the shared task and introduces
our approach; while section 3 focuses on an analysis of the dataset, highlighting certain aspects that
hinder model performance. Section 4 describes our Granite-based approach in detail, and section 5
provides an in-depth analysis of the results. Some conclusions are drawn at the end in section 6, ofering
insights into the task and directions for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description and approach</title>
      <p>PastReader 2025 consists of applying OCR on PDF files with limited quality. Two diferent, yet related,
tasks were proposed. On the one hand, OCR outputs from said PDF files are provided. These outputs are
in plain text and purposely exhibit a suboptimal quality. Hence, the task 1 is to provide clean versions
of these texts. The clean texts are meant to align with the ground truth. On the second hand, the task 2
requires developing an end-to-end OCR system. PDF files are the intended input, ignoring the proposed
OCR text outputs from task 1. From the PDF files, new text outputs are expected to be provided directly.
The main challenge lies in the dataset quality and the variety of the sources.</p>
      <p>As anticipated, our participation consists of fine-tuning Granite3.2-vision:2b [13] on the provided
dataset. This model was selected over other open-source multimodal models for two main reasons.
The first one is that we value the democratisation of LLMs and being able to fine-tune small models
on consumer-grade hardware. This is a small model, with only 2 billion parameters, which raised the
chances of it fitting in the 16gb of VRAM of the Nvidia RTX 5080 employed for this matter. The second
reason is the high performance this model ofers despite its size.</p>
      <p>According to IBM2, this model focuses on being especially performant for enterprises. It has been
ifne-tuned on IBM’s DocFM dataset. It is a “large instruction tuning dataset” consisting of high-quality
enterprise data. Basically, it prioritises visual document understanding with both image and text, which
encompasses document characteristics such as layouts, fonts, charts, etc. Conversely, other models
are mainly trained with natural images, allegedly yielding worse results than IBM’s model. In fact,
based on IBM’s claims, this model rivals larger models in benchmarks such as DocVQA and ChartQA.
Granite stands out in document understanding and multimodal retrieval-augmented generation (RAG).
The remarkable ability of multimodal models to “understand” images and answer questions about said
images makes them a perfect choice for OCR. Therefore, it was considered a suitable choice for this
shared task.</p>
      <p>Given time constraints and the end-to-end nature of Granite’s OCR system, we decided to participate
in Task 2 only.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets analysis</title>
      <p>The dataset [14] provided consisted of historical documents sourced from various archives and
repositories. As a result, it was highly heterogeneous in nature, including variations in font styles, print
quality, background colouration, presence of handwritten marginalia, and embedded visual elements
such as stamps or illustrations. This diversity closely reflects the real-world conditions under which
digitisation eforts must operate, but it also presents a significant challenge for OCR systems that are
typically trained on more standardised inputs.</p>
      <p>These inconsistencies in the visual presentation of the documents have a direct impact on OCR
performance. Diferences in font type and noise introduced by paper degradation or scanning artefacts
can significantly increase the character error rate (CER) and word error rate (WER). For example,
documents with faded ink or coloured backgrounds may result in missed characters or misclassification,
particularly when the contrast between text and background is insuficient.</p>
      <p>The generalisation ability of a model trained on one subset of documents may decline when faced
with very diferent styles from another subset. This is especially problematic in historical corpora,
where typographic standards and orthographic conventions were far from unified.
2https://www.ibm.com/new/announcements/ibm-granite-3-2-open-source-reasoning-and-vision</p>
      <p>Granite’s broad contextual understanding and vision-language architecture ofer promise for
interpreting complex layouts and mixed media. However, like any foundation model, its efectiveness
is highly dependent on how similar the dataset is to its pre-training data. The wide variation in the
dataset can either help by diversifying test conditions or hinder performance if the model has not seen
enough similar samples during training. These results underline the importance of dataset consistency
or adaptive fine-tuning when deploying OCR systems in heterogeneous historical collections.</p>
      <sec id="sec-3-1">
        <title>3.1. Some interesting aspects of the Dataset</title>
        <p>The provided dataset is composed by isolated pages of 8 periodical publications from the 19th and 20th
centuries in Spanish language. The topics they cover are very varied, assorted from cultural and general
news magazines (Juventud, El Español), to satirical and humorous magazines (El Duende Satírico del día),
chronicles of spiritualism (La luz del porvenir), serial publications of narrative fiction and poetry ( Revista
nueva, La patria de Cervantes) or even newspapers about fashion and feminine customs (Periódico de las
damas), as well as scientific publications on medical issues ( Revista frenopática española). As observed,
these press pages difer significantly in their goals, themes, and formats.</p>
        <p>The following are examples of pages to highlight several complexities that models should be designed
to handle. A dedicated model should be trained with these and other common challenges found in
historical Spanish texts in mind, as contemporary texts are not a suitable substitute. However, due to
time constraints, our team has not yet developed such a model, as creating a dataset with a suficient
number of representative examples remains pending.</p>
        <p>First of all we can see that many of the examined documents have in common that they combine
images and text on the page, for the most part narrative (novels, short stories), where the image is simply
a complement to the narration, to provide it with greater expressive vividness. However, in other cases,
there are documents where the image is the main element, and the text is simply a complement to it. For
example, this is the case of a page from a fashion magazine (9062, Figure 1), where the text associated
with the image is the description of the outfit; in other cases, it is the portrait of a phrenologist doctor
(9171, Figure 2), where the associated text is just the doctor’s name and their speciality. Thus, it is
necessary to take into account that the hierarchy of information is reversed in these two cases, with
respect to the informative importance that the images have in relation to the text.</p>
        <p>Also, in many of the examined pages, we find that the image, and a brief text associated with it in
small capitals, interrupts the transcription of the sentences of the narrative texts (8963, Figure 3). In fact,
paratextual elements such as image captions are expected to disrupt the typical linear reading order
followed by OCR systems—from top to bottom, left to right. Another aspect that can hinder the OCR is
the presence of typographical ornamentation, like small and decorated capitals. This phenomenon can
be observed in the examples above. This can potentially lead the models to misinterpret some letters.</p>
        <p>In order to obtain a good training of AI models and reliable results, we consider that it is essential
that the dataset is clean and well organised, as well as the application of systematic transcription rules
based on coherent choices from a philological and palaeographic point of view. For this reason, we are
preparing annotation and transcription guidelines based on the issues detected in the documents of this
training dataset for future work.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Fine-tuning Granite3.2-vision:2b</title>
      <p>The remarkable ability of multimodal models to “understand” images and answer questions about
said images makes them a perfect choice for OCR. However, without fine-tuning, i.e., only through
prompting, the results are less reliable. These models are prone to analyse pictures or answering
relatively complex questions about specific aspects of the image. Therefore, fine-tuning for OCR is
highly advisable, mainly to generate the expected transcription exclusively, with no further explanations.
Providing a benchmark for Granite3.2-vision:2b using only prompting was our intention, but due to
time constraints it was eventually impossible. However, despite inconsistencies with format and several
issues, initial testing hinted promising results, especially with bigger models like llama3.2-vision:11b.
Without fine-tuning, Granite3.2-vision:2b was more prone to wrongly interpret or misspell Spanish
words, probably because its training data is in English language.</p>
      <p>We would like to give proper acknowledgments to the author of this ipynb [15], as it provided a
preliminary template to work with. In this example, Eli Schwartz fine-tunes Granite3.1-vision:2b, the
previous version of the model we used, using the Geometric Perception dataset [16] from Hugging Face,
hence requiring us to adapt it to the dataset of the shared task, among other changes.</p>
      <p>In order to fine-tune the model, the first step was to turn the PDF files into image files, in this case
PNG files. In general terms, this kind of models use image files, so any pipeline will generally include a
conversion at some point. We performed this conversion in advance for convenience, as it saves some
time during processing, although at the cost of an important amount of storage.</p>
      <p>After conversion, the dataset consists of the image of the document in PNG format, and the expected
output in TXT format. However, this model requires a specific input format, similar to an interaction
with a chat model. Figure 4 depicts the structure of the dataset after the input format conversion. As it
can bee seen, for each example, it mainly requires a system prompt, an image, a message from the
user, which is defined as user prompt, and what we expect from the model: the ocr text. The user
prompt is: “Please perform OCR on this Spanish document.” This is the structure of the dataset, which
is passed to the model with the values of the parameters according to each example.</p>
      <p>DATASET STRUCTURE
chat = [
"role": "system", "content": ["type": "text", "text": SYSTEM_PROMPT],
"role": "user", "content": [
"type": "image", "mage": image,
"type": "ext", "ext": USER_PROMPT
],
"role": "assistant", "content": ["type": "text", "text": ocr_text],
]
examples.append(chat)</p>
      <p>As mentioned, it also requires a system prompt. It is more detailed and contains precise instructions
of what is expected from the model. The system prompt used can be seen in Figure 5 below. It instructs
the model not to invent information or modify the text, while trying to obtain the raw text with no
further additions.</p>
      <sec id="sec-4-1">
        <title>SYSTEM_PROMPT</title>
        <p>INSTRUCTIONS:
You are an OCR expert specialised in Spanish documents.</p>
        <p>You are analysing an old book scan with potentially low quality.</p>
        <p>Extract ALL text exactly as it appears.</p>
        <p>Do not correct, interpret or modify the text in any way.</p>
        <p>Return ONLY the raw text, without any additional comments or formatting.</p>
        <p>Do not invent content not present in the image.</p>
        <p>The output must be EXACTLY the recognised text, without adding anything else.</p>
        <p>It is worth noting that a challenging aspect was the use of RAM memory for loading the dataset image
ifles as pixel values. The employed gaming pc features 32gb of RAM memory. Despite the fact that this
is not a low amount of RAM for consumer-grade hardware, it falls short when attempting to process
all the images at full resolution from the dataset, running out of memory. Said resolution could vary
depending on the file, within the ranges of 1000–1500 pixels by 2000–3000 pixels. Even at a resolution
close to 827x1169 pixels, the same error persisted. Considering that Granite3.2-vision:2b’s processor
scans images by cropping them and analysing areas of 384x384 pixels, we reduced the resolution of
the dataset images to a similar size: 414x585 pixels. Using multiples of the crop area dimensions might
have been more eficient. It is worth mentioning that, since the images in the dataset have slightly
diferent aspect ratios, some additional measures had to be taken to ensure all images had the same
dimensions for fine-tuning. This posed the challenge of potential information loss when resizing them
to a fixed resolution given that this process usually crops the images. To address this, each image was
resized while preserving its original aspect ratio, ensuring it fitted within the 414 ×585 pixel target. The
remaining space was then filled with padding to reach the exact required dimensions. There are less
hardware-demanding alternatives to loading images as pixel values, such as including image paths in
the dataset instead of the actual image data and loading them during fine-tuning, or converting the
pixel values to tensors ahead of time for later use. However, these options were not explored in this
paper.</p>
        <p>With respect to quantisation to reduce memory consumption, it is a cornerstone for fine-tuning on
consumer-grade hardware. The model was fine-tuned using QLoRA (Quantized Low-Rank Adapter)
[17], an approach that combines 4-bit quantisation with parameter-eficient fine-tuning based on LoRA
(Low-Rank Adaptation of Large Language Models) [18] adapters. The base model was loaded in 4-bit
precision using the NF4 (Normalised Float 4) quantisation scheme, along with double quantisation
and bfloat16 computation to optimise numerical stability and performance. Specific modules, such as
the vision tower and output head, were excluded from quantisation to preserve their full precision.
Lightweight LoRA adapters were injected into the projection layers of the language model, and only
these adapters were updated during training. This configuration was fundamental for fitting the model
within the RTX 5080’s 16gb of VRAM, although it trades of some performance.</p>
        <p>Regarding hyperparameters, this aspect remained largely unexplored, as time limitations prevented
any meaningful experimentation. The configuration was mostly based on default or initial guesses,
with minimal testing. Despite the use of quantisation strategies, special care was still required. Since
the task involves images, VRAM usage was inherently higher. The hyperparameters used are shown
in Table 1. As observed, the per-device batch size is set to a minimal value of 1, which is partially
mitigated by using 8 gradient accumulation steps. This efectively simulates a larger batch size without
exceeding memory limits. In addition, bfloat16 precision is used alongside gradient checkpointing to
further reduce memory usage during training. The fused adamw_torch_fused optimiser was selected
to maximise computational eficiency. A conservative learning rate of 1e-4 and a weight decay of 0.01
were also applied to ensure stable fine-tuning. Finally, checkpoint saving was limited to the most recent
state to save disk space and simplify checkpoint management.</p>
        <p>Hyperparameter</p>
      </sec>
      <sec id="sec-4-2">
        <title>Number of training epochs</title>
      </sec>
      <sec id="sec-4-3">
        <title>Per-device batch size</title>
      </sec>
      <sec id="sec-4-4">
        <title>Gradient accumulation steps</title>
      </sec>
      <sec id="sec-4-5">
        <title>Warmup steps</title>
      </sec>
      <sec id="sec-4-6">
        <title>Learning rate</title>
      </sec>
      <sec id="sec-4-7">
        <title>Weight decay</title>
      </sec>
      <sec id="sec-4-8">
        <title>Logging steps</title>
      </sec>
      <sec id="sec-4-9">
        <title>Save strategy</title>
      </sec>
      <sec id="sec-4-10">
        <title>Save steps</title>
      </sec>
      <sec id="sec-4-11">
        <title>Save total limit</title>
      </sec>
      <sec id="sec-4-12">
        <title>Optimizer</title>
        <p>bfloat16 precision</p>
      </sec>
      <sec id="sec-4-13">
        <title>Remove unused dataset columns</title>
      </sec>
      <sec id="sec-4-14">
        <title>Gradient checkpointing</title>
      </sec>
      <sec id="sec-4-15">
        <title>Dataset text field</title>
      </sec>
      <sec id="sec-4-16">
        <title>Skip dataset preparation</title>
        <p>Even with all the measures and strategies taken to fine-tune Granite3.2-vision:2b, it narrowly fitted in
the VRAM. The training dataset, comprising both the development and training sets, was randomly split
into two parts: 90% for training and 10% for testing. After successfully fine-tuning on the training dataset,
the model was used for inference on the final test dataset, made available later by the organisation. Apart
from requiring the same dataset format adaptation into a chat format, there was no other significant
step worth mentioning. This Granite approach resulted in GRESEL2_run1.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Analysis of results</title>
      <sec id="sec-5-1">
        <title>5.1. Quantitative analysis</title>
        <p>Table 2 ofers a chart with the results of the baseline, the winning run and our run. Originally, six
diferent metrics were planned to be implemented in the shared task: Word Error Rate (WER), Sentence
Error Rate (SER), Levenshtein Distance, Normalised Edit Distance (NED), BLEU (Bilingual Evaluation
Understudy) and ROUGE (Recall-Oriented Understudy for Gisting Evaluation). However, SER was
apparently disregarded, whereas ROUGE was extended to include four variants—ROUGE1, ROUGE2,
ROUGEL and ROUGELSUM—, and both WER and Levenshtein Distance were preserved, totalling eight
metrics. Consequently, the OCR evaluation relies on a diverse set of complementary metrics that capture
diferent dimensions of performance.</p>
        <p>Moreover, literal accuracy is assessed through character-level metrics such as Levenshtein Distance
and NED, which count the number of required edit operations; and WER, which quantifies lexical
mismatches. Therefore, these metrics favour exact textual reproduction, where lower values indicate
better performance. In contrast, semantic quality is measured using BLEU, which evaluates n-gram
overlap; and the ROUGE family of metrics: ROUGE-1 focuses on term-level recall, ROUGE-2 captures
OCRTITS_run1
GRESEL2_run1</p>
        <p>BASELINE</p>
        <p>ROUGE1 ROUGE2 ROUGEL ROUGELSUM
short-range contextual relationships, and ROUGE-L/ROUGE-LSum assess overall discourse coherence.
In these metrics, higher values indicate better preservation of linguistic meaning, even when
characterlevel diferences are present. Together, these metrics provide a holistic view of system performance,
from raw text fidelity to higher-level comprehension. Figure 6 provides a more visual insight into the
results.</p>
        <p>Table 2 reveals a decent performance in our run. Our fine-tuned Granite3.2-vision:2b (GRESEL2_run1)
achieved a solid second place in five of the eight evaluation metrics, including WER and all ROUGE
variants. In fact, the diferences between Granite and the top-performing system in these metrics are
marginal: ROUGE1 (0.8841 vs. 0.8849), ROUGE2 (0.8049 vs. 0.8065), ROUGE-L (0.8806 vs. 0.8834),
ROUGE-LSUM (0.8837 vs. 0.8843), and WER (0.2643 vs. 0.2344). It also ranked third in BLEU. However,
Granite underperformed in character-level precision, ranking fourth in both Levenshtein Distance
and NED. These results suggest that the model demonstrates excellent semantic understanding and
word-level alignment with the ground truth, but has room for improvement in literal character accuracy.
Given the compact size of the model and the consumer-grade hardware constraints noted earlier, the
performance achieved by Granite3.2-vision:2b is highly notable.</p>
        <p>Regarding emissions, Figure 7 presents a comparative analysis of computational costs. The left chart
uses a logarithmic scale to visualize Granite’s fine-tuning and inference metrics during a simulated
10hour execution window, enabling direct comparison of energy consumption (kWh) and CO2 emissions
(kg) despite their orders-of-magnitude diferences. The right chart similarly uses logarithmic scaling
to display per-example eficiency metrics across the 3,000-sample dataset. These charts highlight a
dramatic diference between the fine-tuning process and the inference process in Granite’s run.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Qualitative insights</title>
        <p>Upon inspecting the output generated by our fine-tuned model, Granite, it was quite satisfactory overall,
but some patterns were identified. For instance, when a word is split across lines by a hyphen at the
end of a line, the model joins the parts, removes the hyphen, and places the line break after the end of
the reconstructed word. The model also misinterprets certain letters and words, particularly those with
diacritical marks. This might be due to the fact that Granite3.2-vision:2b was not specifically pre-trained
on Spanish texts. In general, when the model is unable to recognise a dificult area, it tends to hallucinate
content. It also tends to “correct” misspelled words found in the original, even when those words reflect
historical variants of Spanish used at the time. These tendencies contribute to lower character-level
accuracy, which aligns with the poor results observed in the Levenshtein Distance and NED metrics. On
a positive note, the model consistently refrains from generating output when the original file is blank.
Figure 8 (file 2466), presents two particularly interesting elements: the stamp featuring the words
EL</p>
        <p>CENSO¿
niSCURSO</p>
        <p>LXVIL
V:.Jrribus Anticyris caput instíñabile. .vJ
- • ' ' '
• , - . .</p>
        <p>• • • . . • !</p>
        <p>.• •A
• •</p>
        <p>Horat. in Art. Poét. y."3bdí
" • P a t e t a , «jue sanarla no podrfk
n"'El-hielébofo todo</p>
        <p>.
:j.;jgne en (tres Islas Antíciras se c^ia.
-"•••ií'.". . " • • ' . '
' •
n</p>
        <p>\</p>
        <p>BIBLIOTECA NACIONAL (Spanish National Library), and the drop cap in the letter “S.” Regarding the
words in the stamp, they were ignored by the model, whereas the drop cap was correctly recognised.
A word misspelled in historical Spanish, “Mathe-máticas” (Maths), was merged and corrected to
Matemáticas. In general, the model is more prone to confuse letters or words when the area is blurry or
stained.</p>
        <p>Regarding the second example, shown in Figure 9, the image might initially appear extremely
challenging, as the original physical page seems to have been stained by ink from the facing page on its
right. As a result, it contains horizontally mirrored lines superimposed between the actual lines of text.
Surprisingly, despite this potential challenge, the model did not struggle with these artefacts. What
hindered the recognition of certain words was the presence of blurred letters or smudged areas, yet the
overall transcription is quite acceptable given the circumstances.</p>
        <p>The last example, Figure 10 displays a two-column layout. This should not pose a major challenge,
as OCR models are increasingly better at handling multi-column formats. However, the transcription
of this page was particularly inaccurate. The second (right-hand) column was apparently ignored,
while the first (left-hand) column was heavily hallucinated. The page contains a menu, and the model
invented several dishes for no apparent reason, even adding fabricated information about portion sizes
in some cases. This failure to correctly render the column structure might be due to a lack of suficient
multi-column examples in the training dataset—although this remains a hypothesis.</p>
        <p>Finally, a last observation about the dataset. While a more thorough analysis would be required to
confirm this, the dev/train sets used for fine-tuning appear to be more diverse in terms of sources than
the test set. The latter lacks documents containing images, while the former includes materials with
diferent colour palettes in the paper background. Figures 1, 2, and 3, from the dev/train sets, clearly
illustrate this. The test set displays less variety in page types and tends to feature whiter backgrounds,
in contrast to the yellowish tones found in the dev/train documents. For example, El Censor is one of the
predominant sources in the test set (Figure 8). This mismatch between the dev/train and test sets might
have negatively afected the model’s learning and generalisation, although this remains a hypothesis.
Ensuring both representativeness and diversity across dataset splits is crucial for efective fine-tuning.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>This paper has explored our participation in this challenging OCR shared task. Despite the dificulties
imposed by the nature and diversity of the dataset, our approach has achieved satisfactory results
compared to the baseline. Naturally, some errors commonly associated with AI models were observed,
such as the tendency to hallucinate words when the original text is illegible, or to “correct” misspelled
words found in the source. We have provided a thorough analysis of the dataset from the competition
and ofered some ideas about how to fine-tune small multimodal models for OCR tasks and what the
expected results are. It is worth highlighting the notable performance of Granite3.2-vision:2b, mainly in
semantic and word-selection metrics. Although consumer-grade hardware constraints forced the use of
a smaller variant and several measures that downgraded fine-tuning quality, it excelled in performance.
Overall, this experience has enabled us to learn from our experiments and share our insights into this
OCR task.</p>
      <p>In future work, we will explore the implementation of a dataset variant from this shared task, but
with transcription guidelines focusing on philological and palaeographic criteria, as anticipated. Our
hypothesis is that this dataset variant will improve the model’s learning. Furthermore, other similarly
sized multimodal models will be employed to compare performance between models. Additionally,
larger models will be explored, and full-quality fine-tuning will be applied without downgrading the
quality of the data through the use of more capable hardware and cloud computing at our disposal. We
expect to further enhance our results in OCR tasks by exploring new approaches and ideas.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>
        This work is framed under the Spanish National Project GRESEL: UAM
        <xref ref-type="bibr" rid="ref5">(PID2023-151280OB-C21)</xref>
        . It
has also been partially funded by the PTA2023-023812-I grant, awarded to Yanco Amor Torterolo Orta
by MICIU/AEI/10.13039/501100011033 and the European Social Fund Plus (ESF+).
      </p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT-4o and Deepseek-V3 in order to:
Grammar and spelling check, paraphrase and reword. Besides, Microsoft’s Copilot was also used in
order to: Formatting assistance (latex commands, image labelling and table creation). Further, the
authors used the first two models for figures 6 and 7 in order to: Generate charts based on data.
After using these tools/services, the authors reviewed and edited the content as needed and take full
responsibility for the publication’s content.
M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, J. Lin, Qwen2-vl: Enhancing vision-language
model’s perception of the world at any resolution, arXiv preprint arXiv:2409.12191 (2024).
[10] E. Garcia-Arias, A. Garcia-Serrano, Creación de un modelo de descripciones de imágenes
especializado en arqueología griega, Procesamiento del Lenguaje Natural 75 (2025). In press.
[11] A. Garcia Serrano, A. Menta Garuz, La inteligencia artificial en las humanidades digitales: dos
experiencias con corpus digitales, Revista de Humanidades Digitales 7 (2022) 19–39. URL: https:
//revistas.uned.es/index.php/RHD/article/view/30928. doi:10.5944/rhd.vol.7.2022.30928.
[12] M. Mishra, M. Stallone, G. Zhang, Y. Shen, A. Prasad, A. M. Soria, M. Merler, P. Selvam, S. Surendran,
S. Singh, M. Sethi, X.-H. Dang, P. Li, K.-L. Wu, S. Zawad, A. Coleman, M. White, M. Lewis,
R. Pavuluri, (...), R. Panda, Granite code models: A family of open foundation models for code
intelligence, 2024. URL: https://arxiv.org/abs/2405.04324. arXiv:2405.04324.
[13] G. V. Team, L. Karlinsky, A. Arbelle, A. Daniels, A. Nassar, A. Alfassi, B. Wu, E. Schwartz, D. Joshi,
J. Kondic, N. Shabtay, P. Li, R. Herzig, S. Abedin, S. Perek, S. Harary, U. Barzelay, A. R. Goldfarb,
A. Oliva, B. Wieles, (...), R. Feris, Granite vision: a lightweight, open-source multimodal model for
enterprise intelligence, 2025. URL: https://arxiv.org/abs/2502.09927. arXiv:2502.09927.
[14] A. Montejo-Ráez, E. Sánchez Nogales, G. Expósito Álvarez, A. Ureña López, M. T. Martín-Valdivia,
J. Collado-Montañez, I. Cabrera de Castro, M. V. Cantero Romero, A. García Serrano, R.
Ortuño Casanova, Y. A. Torterolo Orta, Pastreader 2025, https://doi.org/10.5281/zenodo.15084265,
2025. [Data set].
[15] E. Schwartz, Fine-tuning granite vision with trl and peft (lora), https://colab.research.google.
com/github/huggingface/cookbook/blob/main/notebooks/en/fine_tuning_granite_vision_sft_trl.
ipynb, 2024. URL: https://colab.research.google.com/github/huggingface/cookbook/blob/main/
notebooks/en/fine_tuning_granite_vision_sft_trl.ipynb, accessed: 2025-05-08.
[16] J. Zhang, O. Liu, T. Yu, J. Hu, W. Neiswanger, Euclid: Supercharging multimodal llms with synthetic
high-fidelity visual descriptions, arXiv preprint arXiv:2412.08737 (2024).
[17] T. Dettmers, A. Pagnoni, A. Holtzman, L. Zettlemoyer, Qlora: Eficient finetuning of quantized
llms, 2023. URL: https://arxiv.org/abs/2305.14314. arXiv:2305.14314.
[18] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, Lora: Low-rank
adaptation of large language models, 2021. URL: https://arxiv.org/abs/2106.09685. arXiv:2106.09685.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Montejo-Ráez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sánchez-Nogales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Expósito-Álvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Ureña-López</surname>
          </string-name>
          , M. T. MartínValdivia, J.
          <string-name>
            <surname>Collado-Montañez</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Cabrera-de Castro</surname>
            ,
            <given-names>M. V.</given-names>
          </string-name>
          <string-name>
            <surname>Cantero-Romero</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Ortuño-Casanova</surname>
          </string-name>
          ,
          <article-title>Overview of pastreader shared task in iberlef 2025: Transcribing texts from the past</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>75</volume>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>González-Barba</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <article-title>Overview of IberLEF 2025: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Society for Natural Language Processing (SEPLN 2025), CEUR-WS</article-title>
          . org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreno-Sandoval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Porta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Carbajo-Coronado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Torterolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Samy</surname>
          </string-name>
          ,
          <article-title>The financial document causality detection shared task (FinCausal 2025)</article-title>
          , in: C.
          <string-name>
            <surname>-C. Chen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moreno-Sandoval</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ananiadou</surname>
          </string-name>
          , H.
          <string-name>
            <surname>-H. Chen</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Joint Workshop of the 9th Financial Technology and Natural Language Processing (FinNLP)</source>
          ,
          <source>the 6th Financial Narrative Processing (FNP), and the 1st Workshop on Large Language Models for Finance and Legal (LLMFinLegal)</source>
          ,
          <article-title>Association for Computational Linguistics, Abu Dhabi</article-title>
          ,
          <string-name>
            <surname>UAE</surname>
          </string-name>
          ,
          <year>2025</year>
          , pp.
          <fpage>214</fpage>
          -
          <lpage>221</lpage>
          . URL: https:// aclanthology.org/
          <year>2025</year>
          .finnlp-
          <volume>1</volume>
          .21/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Porta-Zamorano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Torterolo-Orta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreno-Sandoval</surname>
          </string-name>
          ,
          <article-title>LLI-UAM Team at FinancES 2023: Noise, Data Augmentation and Hallucinations</article-title>
          ,
          <source>in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF</source>
          <year>2023</year>
          )
          <article-title>co-located with the Conference of the Spanish Society for Natural Language Processing (SEPLN</article-title>
          <year>2023</year>
          ), volume
          <volume>3496</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR-WS, Jaén, Spain,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3496</volume>
          /finances-paper1.
          <fpage>pdf</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Lee</surname>
          </string-name>
          , Visual instruction tuning,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2304.08485. arXiv:
          <volume>2304</volume>
          .
          <fpage>08485</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Grattafiori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jauhri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kadian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Dahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Letman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mathur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schelten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaughan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hartshorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sravankumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Korenev</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Hinsvark,</surname>
          </string-name>
          (...),
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <source>The llama 3 herd of models</source>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/ 2407.21783. arXiv:
          <volume>2407</volume>
          .
          <fpage>21783</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kamath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ferret</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pathak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Vieillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Merhej</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Perrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Matejovicova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rivière</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rouillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mesnard</surname>
          </string-name>
          , G. Cideron, J. bastien Grill,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramos</surname>
          </string-name>
          , E. Yvinec,
          <string-name>
            <given-names>M.</given-names>
            <surname>Casbon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pot</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Penchev</surname>
          </string-name>
          , (...), L. Hussenot,
          <source>Gemma 3 technical report</source>
          ,
          <year>2025</year>
          . URL: https: //arxiv.org/abs/2503.19786. arXiv:
          <volume>2503</volume>
          .
          <fpage>19786</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Poznanski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Borchardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunkelberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Huf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rangapur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wilhelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          , L. Soldaini,
          <article-title>olmocr: Unlocking trillions of tokens in pdfs with vision language models</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2502.18443. arXiv:
          <volume>2502</volume>
          .
          <fpage>18443</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fan</surname>
          </string-name>
          , K. Dang,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>