<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of ICT Research and Applications 11 (2017) 199-211. doi:10.5614/itbj.
ict.res.appl.2017.11.2.6.
[23] M. H. Yahia</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/ICDAR.2009.155</article-id>
      <title-group>
        <article-title>Proposing a Comprehensive Dataset for Arabic Script OCR in the context of Digital Libraries and Religious Archives (Extended Abstract)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Riccardo Amerigo Vigliermo</string-name>
          <email>vigliermo@fscire.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Sullutrone</string-name>
          <email>giovanni.sullutrone@unimore.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Bergamaschi</string-name>
          <email>sonia.bergamaschi@unimore.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Sala</string-name>
          <email>luca.sala@unimore.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Università di Modena e Reggio Emilia (UNIMORE)</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>3365</volume>
      <issue>3</issue>
      <fpage>24185</fpage>
      <lpage>24198</lpage>
      <abstract>
        <p>Optical Character Recognition (OCR) technology is integral to digitizing and accessing historical documents within digital libraries. However, OCR systems often struggle to accurately recognize and classify complex document structures, especially historical texts with diverse layouts and languages. This preliminary study addresses this challenge by proposing the building of a comprehensive and community accessible dataset of Arabic title pages using advanced Vision Language Models (VLMs) and OCR tools. In this context, by extracting the first pages of each document at high resolution, we focused on accurately classifying frontispieces and distinguishing them from the main text to improve metadata quality and document retrieval in digital libraries. The Qwen-2vl-72B model was utilized to classify each page as either a 'frontispiece'or 'non-frontispiece'using a specially designed prompt. The identified frontispieces will be processed using Google Vision AI to automatically extract a Ground Truth to be evaluated by linguistic experts before completing the dataset. Further steps will envisage training of an open source solution such as Kraken OCR to also evaluate the efectiveness of the dataset. The innovative approach introduced here not only addresses the current lack of comprehensive datasets but also advances the efectiveness and precision of digital library initiatives such as the Digital Maktaba project.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Arabic OCR</kwd>
        <kwd>Datasets</kwd>
        <kwd>Title pages</kwd>
        <kwd>Digital Libraries</kwd>
        <kwd>Religious Archives</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        As digital mediums become central to information consumption, the demand for accessible and
searchable digital texts has grown exponentially. Optical Character Recognition (OCR) technology has played a
pivotal role in this transformation by converting various document formats into editable and searchable
data. This technology is crucial for digital archiving, information retrieval, and data analysis. OCR
technology still faces notable challenges, particularly with complex and historical documents. These
challenges are amplified when processing Arabic script, where unique linguistic, typographic and
calligraphic characteristics afect OCR accuracy. One of the primary barriers in OCR research and
application is the lack of comprehensive, high-quality datasets designed for library usage. Moreover,
in comparison to other languages and scripts, existing datasets often lack the breadth and specificity
required to address these complex features, hindering the development and benchmarking of advanced
OCR algorithms. This limitation is especially critical for elements like frontispieces (i.e., title pages)
which contain unique artistic and typographical components that demand specialized OCR handling.
Fig. 1 shows an example of title pages. In developing a cataloging tool, which is the main aim of the
Digital Maktaba project 1, we present here a possible pipeline for the creation of a title pages dataset to
efectively train an Open Source OCR model such as Kraken [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], through the escriptorium VRE [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
to extract cataloging metadata from Arabic printed frontispieces. The presented work also considers
recent advances in Vision Language Models [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] that could contribute significantly to data extraction
from images by integrating visual, as well as textual understanding. By leveraging VLMs and a closed
source OCR such as Google Vision AI this study addresses these challenges by developing an extensive,
community-accessible dataset of title pages in the Arabic script. We focus specifically on accurately
classifying frontispieces and distinguishing them from the main text within the initial pages of each
document. Additionally, we aim to generate enriched metadata for improved organization and retrieval
in digital libraries. Currently, we are in the process of constructing a dataset that captures the diverse
typographic and structural challenges of Arabic texts. Our immediate goal is to finalize this dataset
and utilize the best available OCR tools to automatically extract text, or portions of it, from these
documents. This extracted text will then be meticulously reviewed and corrected by experts to create a
gold-standard reference, ensuring accuracy for future OCR advancements. In the sections that follow,
we provide background on OCR technology, elaborate on the unique challenges posed by Arabic script,
outline our objectives and methodology, and discuss the preliminary workflow and the anticipated
outcomes and implications of our study.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Works</title>
      <sec id="sec-2-1">
        <title>2.1. Arabic Optical Character Recognition in Digital Libraries: the Title page as FPG and VrD</title>
        <p>
          OCR technology has played a key role in library and archive digitization eforts across the globe.
By converting printed text into machine-readable formats, OCR facilitates the storage, retrieval, and
analysis of vast document collections. Traditional OCR systems rely on pattern recognition and machine
learning techniques to interpret character shapes and word patterns in scanned images [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. While
efective for documents with conventional fonts and layouts, these systems struggle with documents
that deviate from these standards. The Arabic script poses several specific challenges due to its unique
linguistic and typographic features. Arabic is written in a cursive, right-to-left (RtL) script, complicating
1The Digital Maktaba (DM), defined as Work Package (WP) 5 in the ITSERR (Italian Strengthening of the ESFRI RI RESILIENCE)
project, functions as the primary source and ultimate goal of the study presented here. This WP is dedicated to crafting
a digital library that can analyze and extract information from multi-lingual documents, particularly from Arabic scripts
(Arabic, Persian and Azerbaijani), ofering a state-of-the-art cataloging methodology designed specifically for religious
studies libraries that need to manage multi-lingual and multi-alphabetic cultural resources
character segmentation. Some Arabic graphemes have diacritical dots (above or below the baseline) that
can shift in diferent calligraphic-typographic styles leading to recognition errors. The script primarily
represents consonants, with vowels indicated by optional diacritical marks (that are often not presented).
Arabic characters change shape depending on their position within a word (initial, medial, final, or
isolated). Some graphemes do not connect (bind) with others creating words with multiple disconnected
components. Lastly, in some calligraphic styles the characters may overlap, touch, or appear in slanted
orientations. In the present preliminary investigation, the title pages could be conceived as part of
a series of pages that we refer to as Frontispiece Pages Group (FPG). This definition is motivated by
the context of frontispiece OCR analysis and character extraction for the development of a system
able to support the librarian’s work. FPGs are groups of pages where most of the metadata useful for
cataloging is present. It should be also considered as such since we often have the recurrence of the same
information in diferent scripts on diferent pages and with diferent layouts (e.g., title is represented
in the title page, as well as in other pages, sometimes in other fonts). The title page in many cases is
a black-on-white re-proposition of the cover page where the mere binarization of the text does not
solve other graphical issues (decorations, vocalization, etc.). Subsequent pages of the FPG usually report
useful data (even in a fragmented manner across several pages) in more ‘normalized’scripts allowing for
easier text extraction and cataloging. In some cases, the information is placed inside special boxes on
one of the pages following the title page. Moreover, the title page in the FPG, from a Document Analysis
perspective, could be considered in many cases the same as a Visually Rich Documents (VrD) [
          <xref ref-type="bibr" rid="ref6 ref7 ref8 ref9">6, 7, 8, 9</xref>
          ],
especially considering that scanned PDFs from the physical realm could also bear noise elements such
as library stamps or marks and several other issues related to the state of the paper as support.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Peculiar challenges posed by Arabic script FPGs</title>
        <p>Frontispieces present unique challenges for OCR extraction that are not typically encountered with
internal pages. Frontispieces often exhibit a high degree of variability and complexity. Several factors
contribute to the increased dificulty:
Variety in Layouts and Designs. Frontispieces may include ornate designs, decorative elements, and
unconventional layouts that intertwine text and images. This visual richness can confuse OCR systems,
which are primarily trained on text-centric pages.</p>
        <p>Diverse Backgrounds and Noise. The presence of backgrounds with various colors, textures, or
deteriorated conditions adds noise to the images. Such backgrounds can interfere with text recognition
by obscuring characters or creating false positives.</p>
        <p>Non-Standard Fonts and Scripts. Frontispieces often feature artistic or custom typefaces, including
calligraphic styles like Naskh, Nasta’l¯ıq, or Ku¯f¯ı. These fonts have unique graphical peculiarities that
are not always well-represented in standard OCR training datasets.</p>
        <p>Multiscript Content. They may contain text in multiple scripts, such as Arabic and Latin sometimes
within the same page. Moreover, the use of numerals (for dates) and alphabetic script is by itself a
challenge since Arabic-indic numerals have a LtR orientation while Arabic script is RtL.
Presence of Vocalization and Diacritics. The inclusion or omission of vowels and diacritical marks
can vary, afecting character recognition. Decorations or artistic elements might be mistaken for
diacritics, leading to misinterpretation of the text.</p>
        <p>
          These challenges are compounded by external variables such as overall image quality, character
resolution, diferent levels of support degradation and the presence of colored fonts or backgrounds. Previous
works [
          <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13">10, 11, 12, 13</xref>
          ] highlighted that these issues requires not only advanced OCR algorithms but also
carefully curated datasets.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Related Works: Vision-Language Models in OCR</title>
        <p>
          VLMs models have applications across a range of fields, including image captioning, visual question
answering, and object recognition, where both visual and textual data are combined [14]. Leveraging
large-scale neural networks and extensive datasets, VLMs are designed to interpret complex images
that may include text, graphics, and other visual elements [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Although VLMs are still emerging in
document analysis, their potential for handling mixed-media layouts and intricate document structures
is promising. Unlike traditional OCR systems, which are focused solely on recognizing text, VLMs take
a holistic approach by analyzing both visual and textual components of an image. For instance, the
Qwen-2vl-72B [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] model is designed to analyze multimodal data and can perform tasks such as image
captioning and visual context interpretation, which could theoretically aid in recognizing text within
complex visual contexts. At the time of writing, it is the best overall performing open-source VLM
according to benchmark evaluations [15].
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Related Works: Arabic Printed Characters Datasets</title>
        <p>In the last two decades Arabic script OCR studies have made significant steps forward. An example of
dataset useful for both handwritten and printed Arabic text is ARABASE [16]. The Dataset is a collection
of documents images and Part of Arabic Words (PAWs). In 2009, the APTI dataset (Arabic Printed
Text Image) was composed of 45,313,600 word images covering 250 million characters representing
one of the most extensive [17]. In 2010 PATDB (Printed Arabic Text Data Base) [18] was published
as a corpus consisting of 6954 scanned pages images with diferent dpi resolutions. Few years later
the multi-font dataset APTID/MF (Arabic Printed Text Image Dataset/Multi-Font) emerged as another
solution for segmentation and automatic font identification researches with 387 pages of documents
scanned in grayscale, from which 1845 text-blocks and a large dataset of 27,402 samples were extracted
[19]. Similarly, the ALTID dataset was composed by 1,845 text blocks in Arabic and 2,328 in Latin
alphabet from 731 greyscale images [20]. Worth to mention is also the KAFD dataset [21], which consists
in 40 fonts in 10 sizes and 4 diferent styles (e.g., bold, italic, etc.), while in 2017 another PAW dataset was
created from 83,056 text images representing all the words of the Arabic language in diferent Arabic
fonts: (e.g., Thuluth, Naskh, etc.) for a combined total of 415,280 images [22]. In 2022 a bilingual dataset
with the name of BPTI (Arabic/English) was ideated to address the lack of availability of bilingual text
datasets. It consists of 97,812 text images categorized into two groups: Scanned page and digitized line
images [23]. Finally, two datasets designed and developed for the recognition of Arabic printed text
with examples and text images collected from the Qur’a¯n: QTID (Quran Text Image Dataset) consisting
of 309,720 images with a total of 2,494,428 characters from the Qur’anic text [24]; and the second [25]
containing 604 images at page level and 8,927 images at text line level from the Medina Qur’a¯n (mus. h.af
al-mad¯ına). At the best of our knowledge no specific dataset has been developed for the analysis of
Arabic printed title pages OCR handling in the context of librarian use and digital libraries.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        The primary focus of this study is to improve the OCR capabilities for digital libraries by developing a
comprehensive, community-accessible frontispiece dataset. Our methodology involves assembling an
initial set of historical documents from the FSCIRE "La Pira" digital archive, classifying their pages as
’frontispiece’ or ’non-frontispiece’, and creating an initial OCR draft of the frontispieces using Google
Vision AI, which will be evaluated with common metrics such as Carachter Error Recognition (CER) and
Word Error Recognitin (WER), analyzed and corrected by linguistic experts. The decision to implement
this tool for dataset creation is based on previous tests of open-source solutions such as EasyOCR,
Tesseract, and Google Docs [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] where the latter emerged for better performances on the languages in
our collection. Despite that, the large number of images to be processed and the accuracy and flexibility
needed for our specific goal led us to the selection of Google Vision AI. A visual representation of the
pipeline is shown in Fig.2 while further details are provided in the following sections.
Document Selection and Preparation. We collected approximately 140,000 donated documents,
selected for their diversity in content, formats and cultural significance. This collection ensures that the
dataset spans a range of subjects, periods, languages, layouts, fonts, and visual elements, aligning with
our goals of cataloging and preserving large non-latin cultural heritages. The documents were then
divided into digitized and non-digitized with the latter selected for the next processing steps.
Resource management.To manage resources efectively and maintain consistency, we extracted the
ifrst six pages of each document at high resolution, as these initial pages typically contain frontispieces
and other introductory materials. Each page is scaled proportionally to 4096x4096 pixels to reduce the
amount of converted tokens for the VLM used in the next step.
      </p>
      <p>Processing Documents. To process these pages, we selected the Qwen-2vl-72B model, one of the best
open-source VLMs currently available. Using a specialized prompt, we guided the model in classifying
pages as frontispieces or non-frontispieces. After obtaining the subset of frontispieces, we used Google
Vision AI to produce an initial text extraction. Given that the analyzed pages are particularly challenging
(as stated in Section 2.2), the results of this step will be given to human experts for correction, resulting
in a high-quality OCR data set at the end of the pipeline.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion and Future Directions</title>
      <p>In this extended abstract, we have presented our planned approach to improve OCR capabilities for
Arabic frontispieces in digital libraries. Recognizing the unique challenges of Arabic script and the
scarcity of specialized datasets, we aim to develop a comprehensive, high-quality dataset by processing
approximately 140,000 historical documents. By employing advanced Vision-Language Models like
Qwen-2vl-72B for page classification and OCR tools such as Google Vision AI for initial text extraction,
we intend to create a reliable resource for training and benchmarking OCR algorithms. Our future
work will focus on finalizing this dataset, refining the OCR pipeline, and collaborating with linguistic
experts to ensure accuracy. We believe this efort will significantly contribute to the preservation and
accessibility of Arabic texts in digital libraries, supporting advanced cataloging and research initiatives.
Moving forward, we plan to publicly release our curated frontispieces dataset, providing a valuable
resource for the research community. With this dataset, we plan to train the open source Kraken
OCR engine to develop a model specifically tailored for the librarian and cataloger usage, also aiming
to improve OCR accuracy for these complex pages and to facilitate better metadata extraction and
cataloging. Additionally, by ofering the dataset as a benchmark, we hope to support the evaluation and
advancement of OCR systems focused on frontispiece recognition.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was supported by the PNRR project Italian Strengthening of Esfri RI Resilience (ITSERR)
funded by the European Union – NextGenerationEU (CUP:B53C22001770006).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kiessling</surname>
          </string-name>
          , G. Kurin,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Smail</surname>
          </string-name>
          ,
          <article-title>Advances and Limitations in Open Source ArabicScript OCR: A Case Study</article-title>
          ,
          <source>Digital Studies / Le champ numérique 11</source>
          (
          <year>2021</year>
          ). URL: http://arxiv.org/ abs/2402.10943. doi:
          <volume>10</volume>
          .16995/dscn.8094, arXiv:
          <fpage>2402</fpage>
          .10943 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Stokes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kiessling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Stökl Ben</surname>
          </string-name>
          <string-name>
            <surname>Ezra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tissot</surname>
          </string-name>
          , E. Gargem,
          <article-title>The EScriptorium VRE for Manuscript Cultures</article-title>
          ,
          <source>Classics@ Journal, Ancient Manuscripts and Virtual Research Environments</source>
          <volume>18</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Qwen-vl: A versatile vision-language model for understanding, localization</article-title>
          , text reading, and beyond,
          <year>2023</year>
          . URL: https://api.semanticscholar.org/CorpusID:261101015.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Improved baselines with visual instruction tuning</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>26296</fpage>
          -
          <lpage>26306</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Naz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Umar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Shirazi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Razzak</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Siddiqi</surname>
          </string-name>
          ,
          <article-title>Segmentation techniques for recognition of Arabic-like scripts: A comprehensive survey</article-title>
          ,
          <source>Education and Information Technologies</source>
          <volume>21</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10639-015-9377-5.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Graph Convolution for Multimodal Information Extraction from Visually Rich Documents</article-title>
          , in: A.
          <string-name>
            <surname>Loukina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Morales</surname>
          </string-name>
          , R. Kumar (Eds.),
          <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Industry</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>32</fpage>
          -
          <lpage>39</lpage>
          . URL: https://aclanthology.org/N19-2005. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -2005.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine Reading</article-title>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/ 2310.14802. doi:
          <volume>10</volume>
          .48550/arXiv.2310.14802, arXiv:
          <fpage>2310</fpage>
          .14802 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>K.-A. L. Nguyen</surname>
          </string-name>
          ,
          <article-title>Document Understanding with Deep Learning Techniques</article-title>
          . Document and
          <string-name>
            <given-names>Text</given-names>
            <surname>Processing</surname>
          </string-name>
          .,
          <string-name>
            <surname>Ph</surname>
          </string-name>
          .D. thesis, Sorbonne Université,
          <year>2024</year>
          .,
          <year>2024</year>
          . URL: https://theses.hal.science/ tel-04626992/file/140733_NGUYEN_
          <year>2024</year>
          <article-title>_archivage</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Piwowarski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Laborde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Moyse</surname>
          </string-name>
          , Learning Reading Order via
          <article-title>Document Layout with Layout2Pos</article-title>
          , in: A.
          <string-name>
            <surname>Antonacopoulos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Hinze</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Piwowarski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Coustaty</surname>
            ,
            <given-names>G. M.</given-names>
          </string-name>
          <string-name>
            <surname>Di Nunzio</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Gelati</surname>
          </string-name>
          , N. Vanderschantz (Eds.),
          <source>Linking Theory and Practice of Digital Libraries</source>
          , Springer Nature Switzerland, Cham,
          <year>2024</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -72437-
          <issue>4</issue>
          _
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Martoglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ruozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Vigliermo</surname>
          </string-name>
          , S. De Nardis,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanzini</surname>
          </string-name>
          ,
          <article-title>Preserving and conserving culture: First steps towards a knowledge extractor and cataloguer for multilingual and multi-alphabetic heritages</article-title>
          ,
          <source>in: Proceedings of the Conference on Information Technology for Social Good</source>
          , GoodIT '21,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>301</fpage>
          -
          <lpage>304</lpage>
          . URL: https://doi.org/10.1145/3462203.3475927. doi:
          <volume>10</volume>
          .1145/3462203.3475927.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          , S. De Nardis,
          <string-name>
            <given-names>R.</given-names>
            <surname>Martoglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ruozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Vigliermo</surname>
          </string-name>
          ,
          <article-title>Novel perspectives for the management of multilingual and multialphabetic heritages through automatic knowledge extraction: The digitalmaktaba approach</article-title>
          ,
          <source>Sensors</source>
          <volume>22</volume>
          (
          <year>2022</year>
          ). URL: https://www.mdpi. com/1424-8220/22/11/3995. doi:
          <volume>10</volume>
          .3390/s22113995.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Martoglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Vigliermo</surname>
          </string-name>
          ,
          <article-title>A tool for semiautomatic cataloguing of an islamic digital library: A use case from the digital maktaba project</article-title>
          , in: A.
          <string-name>
            <surname>Paschke</surname>
            , G. Rehm,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Neudecker</surname>
          </string-name>
          , L. Pintscher (Eds.),
          <source>Proceedings of the Third Conference on Digital Curation Technologies (Qurator</source>
          <year>2022</year>
          ), Berlin, Germany, Sept.
          <fpage>19th</fpage>
          -
          <lpage>23rd</lpage>
          ,
          <year>2022</year>
          , volume
          <volume>3234</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEURWS.org,
          <year>2022</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3234</volume>
          /paper1.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Martoglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ruozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Vigliermo</surname>
          </string-name>
          ,
          <article-title>Knowledge extraction, management and long-term preservation of non-Latin cultural heritages - Digital Maktaba project presentation</article-title>
          , in: B.
          <string-name>
            <surname>Alessia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Alex</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Stefano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Stefano</surname>
          </string-name>
          , R. Domenico (Eds.),
          <source>Proceedings of the 19th Conference on Information and Research Science Connecting to Digital and Library Science</source>
          , volume
          <volume>3365</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR, Bari, Italy,
          <year>2023</year>
          , pp.
          <fpage>153</fpage>
          -
          <lpage>161</lpage>
          . URL:
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>