<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Arabic Maghribi Handwritten Text Recognition with RASAM 2: A Comprehensive ⋆ Dataset and Benchmarking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chahan Vidal-Gorène</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Clément Salah</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noëmie Lucas</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aliénor Decours-Perez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antoine Perrier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNRS, Centre Jacques Berque</institution>
          ,
          <addr-line>Maroc</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Calfa</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>École Nationale des chartes-Université PSL, Centre Jean-Mabillon</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Sorbonne Université (UMR 8167), Université de Lausanne (IHAR)</institution>
          ,
          <addr-line>France</addr-line>
          ,
          <country country="CH">Suisse</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Edinburgh</institution>
          ,
          <addr-line>Scotland</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>4</fpage>
      <lpage>6</lpage>
      <abstract>
        <p>Recent advancements in handwritten text recognition (HTR) for historical documents have demonstrated high performance on cursive Arabic scripts, achieving accuracy comparable to Latin scripts. The initial RASAM dataset, focused on three Arabic Maghribi manuscripts, facilitated rapid coverage of new documents via fine-tuning. However, HTR application for Arabic scripts remains constrained due to the vast diversity in spellings, ambiguities, and languages. To overcome these challenges, we present RASAM 2, an extended dataset with 3,750 lines from 15 manuscripts in the BULAC library, showcasing various hands, layouts, and texts in Arabic Maghribi script. RASAM 2 aims to establish a new benchmark for HTR model training for both Maghribi and Oriental scripts, covering text recognition and layout analysis. Preliminary experiments using a word-based CRNN approach indicate significant model versatility, with a nearly 40% reduction in Character Error Rate (CER) across new in-domain and out-of-domain manuscripts.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;dataset</kwd>
        <kwd>Arabic scripts</kwd>
        <kwd>handwritten text recognition</kwd>
        <kwd>historical manuscripts</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>other non-Latin scripts. The initial dataset, RASAM 1, included 300 images from three
Bibiothèque des langues et civilisations (BULAC) manuscripts copied between 1734 and 1875,
achieving promising results with an in-domain Character Error Rate (CER) of 4.8%.</p>
      <p>However, the limited scope of RASAM 1 restricted its efectiveness in recognizing
out-ofdomain manuscripts, even those with similar contemporary scripts and themes (see Table
1). To overcome these limitations, we introduce RASAM 2, an expanded dataset comprising
3,750 lines from fiteen manuscripts, encompassing a broader range of themes and handwriting
styles. RASAM 2 aims to provide a comprehensive reference for training HTR models for
Arabic scripts, enhancing their robustness and applicability across diverse Arabic Maghribi and
Oriental texts. This paper presents the technical details of RASAM 2, its composition, and the
initial results of using a new word-based Convolutional Recurrent Neural Network (CRNN)
approach, which shows significant improvement in model versatility and a substantial reduction
in CER for both in-domain and out-of-domain manuscripts.</p>
      <p>©JJ.Ë@ ©J®Ë@
Commentary: The fā is confused with a bā (in both cases, a single point is subscribed).</p>
      <p>Qå•Ag ñ“Ag
Commentary: The rā is confused with a wāw (more or less open and long final).</p>
      <p>XYªK. hYë
Commentary: The pair of letters bā and ’ayn were confused with anhā (the subscript
point of the bā was not spotted). The final dāl is confused with a ḥā, they may have
a close ending.</p>
      <p>àAÓQË@ àAÓñË@
Commentary: The rā of rummān (pomegranates) became a wāw, both often very close
realisations - a possible example of a food word unknown by the model.</p>
      <p>Që@ñm.Ì'@ Qå”¯@ñmÌ'@
Commentary: The first subscribed point is misunderstood and the ǧīm of ǧawāhir
(jewels or gems) is confused with an ḥā. The unusually wide realisation of the hā is
mistaken for a qāf (the dot on the line below is mistakenly equated with this line)
followed by a ṣād. The rā is well understood.</p>
    </sec>
    <sec id="sec-2">
      <title>2. State-of-the-art datasets for Arabic scripts</title>
      <p>
        The study of documents in Arabic constitutes a separate field within the handwritten text
recognition and document analysis questions more generally, owing to the great diversity and
variability they encompass, hence the workshops dedicated to this specific issue held at the last
ICDAR and ICFHR conferences. The latest developments in HTR for Arabic have however
demonstrated that the use of dedicated CRNN enables to overcome the issue of text
recognition for these scripts, with CER below 5%, even below 3% in specific cases, with few training
data [
        <xref ref-type="bibr" rid="ref16 ref9">16, 9</xref>
        ]. At this stage, these specialized models exceed the performance achieved by
Transformers for Arabic, the latest results on Al-Soudani Maghrebi script achieving an average of
10% CER with large dataset[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The text detection is also efective on Arabic documents, for
instance, the use of FCN [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] allows for a good text-line detection. For the semantic
classification of contents, using a non-specialized U-net 1[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] outperforms the FCN results, which is
notably facing problems in diferentiating two close text regions of the same type, unlike U-net.
Several open-ended questions remain, such as the processing of very cursive scripts, the issue
of transcription and the ambiguity of diacritics, or the reading of abbreviations.
      </p>
      <p>
        In recent years, numerous datasets have emerged in an attempt to overcome these diferent
tasks. In the instance of non-historical documents, the IFN/ENIT dataset 1[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], focused on
modern scripts and produced in a very restricted context, is an important point of reference, notably
used for the automatic generation of handwritten lines 5[]. Not designed for HTR purposes,
the KHATT dataset ofers a dataset in modern scripts with 1,000 diferent copyists[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], mainly
intended for writer identification, as well as the QUWI and LAMIS-MSHD datasets[
        <xref ref-type="bibr" rid="ref10 ref4">10, 4</xref>
        ].
      </p>
      <p>
        In the instance of historical documents, very specialized datasets exist, such as WAHD1[],
dedicated to writer identification, or KERTAS [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], dedicated to manuscript dating. There exist
datasets non-specialized on a specific Arabic script, such as HADARA80P [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and VML-HD [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
notably for RASM2018 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] comprised of scientific manuscripts from the Qatar Digital Library,
or BADAM [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] focused on line detection in Arabic documents, particularly complex ones. More
recently, the RASAM 1 dataset [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] targets Arabic Maghribi scripts, in contrast to RASM and
BADAM, which focus on oriental scripts. It ofers typical layouts and hands as representative of
the common Maghribi production, selected for the purpose of quickly developing HTR models
operable for both research and production. The dataset has since been extended within the
scope of the TARIMA project, with 120 pages manually transcribed from 28 various Arabic
Maghribi sources, including lithographs1. The dataset has been designed for fine-tuning tasks
from RASAM 1. For the oriental scripts, we can also mention the Iskandar dataset from the
Alexander Hackathon, focusing on 5 manuscripts of the Alexander romance in Middle Arabi2c.
      </p>
      <p>Together, these datasets are already covering a vast part of the production of documents
in Arabic scripts (subject to their compatibility, see Table2). Although the proof of concept
is successful for text recognition, the challenge today is to increase the versatility of existing
models by providing a greater variety of fully annotated and transcribed documents.</p>
      <sec id="sec-2-1">
        <title>1https://github.com/calfa-co/tarima 2https://gitlab.huma-num.fr/lipa/iskandar</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset composition</title>
      <sec id="sec-3-1">
        <title>3.1. Quantitative description</title>
        <p>Summary: RASAM 2 dataset comprises 250 images from 15 diferent manuscripts. 3,750 lines
in total have been transcribed, 250 lines by manuscript on average, regardless of the type (main
text or marginal notes). It entails 5,702 annotated lines in total and focuses on Arabic Maghribi
manuscripts (see Table 5 in appendix for the complete list of manuscripts). Its purpose is to
extend the variety of cases encountered in RASAM 1, in order to provide a robust training basis
for documents in Arabic scripts.</p>
        <p>
          • Dataset availability (v.1.0): https://github.com/calfa-co/rasam-dataset.
• License: Apache2.0
• Data format: pageXML with Text regions and lines
• Annotation tool: Calfa Vision3 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
• Ontology for annotation: SegmOnto [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
• Transcription guidelines: Same as RASAM 1 (no missing hamza or diacritics added)
Methodology for data creation: The images have been randomly selected in the
manuscripts to constitute a representative sample of the production, of the states of
conservation, and of the handwriting quality. The images have been pre-annotated with the baseline
and text region detection models trained on RASAM 1 and available within the project type
”Arabic Manuscript (default)” on the annotation platform. Afterwards, the predictions have
been manually checked by the participants during the hackathons. Transcription guidelines
follow RASAM 1 recommendations [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>The dataset holds 522,371 characters (divided in 54 classes) for a total of 93,855 words (divided
in 22,027 classes). The ḍammatan and @ classes in particular are under-represented and are
likely to be less encountered, and so less recognized in a character-based approach (see below
Section 4). The words waw ( ð ), min/man ( á Ó ) and fī ( ú¯ ) are the most represented in the
dataset, with 4,398; 2,246 and 2,189 occurrences respectively, a contrario the wordsal-akhdūd
(XðYgB@), qaṭām (ÐA¢¯) and la‘ād (XAªË) are among the least represented (a single occurrence).</p>
        <p>We retained four text regions and two annex regions for the semantic classification of
contents:
• MainZone: the main text region of the document. This region can appear several times
within a single page, when the text is segmented or in case of a multiple column layout;
• MainZone:title: text region located at the same level as the main text, for headings and
stylized titles;
• MarginTextZone: marginal text region regardless of its location in the page;
• MarginTextZone:catchword: marginal text region corresponding to the catchwords,
systematically under the main text region;
• StampZone: stamps present on the page;
• TableZone: region corresponding to a table.</p>
        <p>A summary of the text regions distribution is given in Table3.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Qualitative description</title>
        <p>
          As outlined in the introduction, the aim of this new dataset is to enhance the versatility and
robustness of RASAM 1 by training it on a wider variety of manuscripts in order to expand the
base of its (1.) vocabulary, (2.) layouts and (3.) scripts. As a result, 15 manuscripts make up
this new dataset.
(1.) Of the fiteen new manuscripts, five (1/3 of the corpus) have themes and/or vocabulary
related to the first dataset. Like MS.ARA.1977 (RASAM 1), MS.ARA.1944 (RASAM 2) belongs
to the historical genre; and like MS.ARA.609 (RASAM 1), the manuscripts MS.ARA.1936, 1943,
1960 and 1983 deal with Islamic law – with the diference that, on the one hand, the legal issues
are not identical, which means that a new vocabulary has to be learned, and that, on the other
hand, MS.ARA.1936 also includes Berber written in Arabic. The other ten manuscripts of the
new dataset (the remaining 2/3 of the corpus) cover new themes, not yet dealt with by RASAM
1. In detail, MS.ARA.1947 is a classical Arabic literature text, MS.ARA.1926 a collection of
litanies, MS.ARA.23, 24, 45b and 1982 cover vocabulary related to Arabic grammar and linguistics.
MS.ARA.6, 9, 65 and 1946 consist of collections on various topics ranging from Islamic
jurisprudence to Arabic grammar, including private correspondence and exegesis of Qur’anic verses.
In addition, several manuscripts show significant variations in handwriting, particularly for
the latter collections.
(2.) From the layout perspective, the RASAM 1 dataset already covered complex layouts:
MS.ARA.609 integrated many tables within the body of the text and MS.ARA.1977 recorded
many lines of poetry which traditionally are ofset from the main text [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The RASAM 2
dataset intends to enhance the capabilities of the model in handling complex layouts. In detail
(see Figure 1), the RASAM 2 dataset reinforces its capabilities in the treatment of poetry verses
(MS.ARA.6), tables (MS.ARA.65) and marginal comments, whether they are aligned with the
main text as in MS.ARA.1943, or rounded, or even inverted as in MS.ARA.1936. Moreover, the
RASAM 2 dataset develops new skills, in particular in the identification of interlinear comments
(MS.ARA.1947) or particularly stylised titles (MS.ARA.1926) as well as in the processing of more
complex page layouts, notably with the presence of gap texts (MS.ARA.1960).
        </p>
        <p>
          (3.) From a strictly palaeographic point of view, the RASAM 2 dataset intends to deal with
a broader variety of hands. The emphasis has been placed on three points in particular. (a.)
Firstly, particular interest has been given to the use of colors within these diferent manuscripts.
Some recent experiments conducted on the basis of RASAM 1 show that the use of colors largely
hinders the models’ good recognition of characters9[]. Therefore, many manuscripts in the
RASAM 2 corpus aim at providing the model with many color realizations (see MS.ARA.1926
and MS.ARA.6 supra, where blue, green, red and yellow are used in particular). (b.) Secondly,
RASAM 2 intends to be able to handle diferent text densities. RASAM 1 was indeed based
on only 3 manuscripts which, although diferent from the density aspect [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], did not cover
the multiple realizations of Arabic manuscripts in Arabic Maghribi scripts. In order to fill this
gap, RASAM 2 is built on a broad continuum in terms of density from very airy manuscripts
— such as MS.ARA.1926 with less than ten lines per page and less than ten words per line —
to extremely dense manuscripts — such as MS.ARA.1982 with more than forty lines per page
and slightly less than twenty words per line, or MS.ARA.1943 with thirty-five lines per page
and more than twenty words per line. (c.) Finally, RASAM 2 covers a wider range of Arabic
Maghribi scripts. The model is thus built from very careful and stylized, almost calligraphic
hands following the example of MS.ARA.1926 (see below6) or hands that are characterized
by a wide amplitude of their final tails — see in particular the realization of the final lām in
the word qāla of MS.ARA.6, 1926, 1946, 1947 (see Table 6 in appendix). Conversely, RASAM
2 also includes very cursive and crowded scripts, as is the case for MS.ARA.1943, 1982. In
sum, and as schematically represented in Figure2, RASAM 2 covers a wider reality of Arabic
Maghribi hands. It leads to a pre-generic model for the treatment of Arabic Maghribi scripts,
far exceeding the possibilities ofered by RASAM 1, which was still only a proof of concept
until then.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. HTR of Arabic versatility experiments</title>
      <sec id="sec-4-1">
        <title>4.1. Methodology</title>
        <p>
          The latest developments in HTR for handwritten documents in Arabic scripts have shown that
operating a word-based CRNN (where every word is considered as a diferent class to identify)
outperforms a basic character-based CRNN (where each character is considered as a diferent
class to identify) on documents with a steady lexicon (both in learning time and CER)9[]. This
approach, despite being dependent on the targeted lexicon, relies on recognizing a word in
context, which appears a more robust approach for cursive Arabic scripts) 9[]. We hold onto
this approach, which is a variation of the one implemented for RASAM 1[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Some
underrepresented word classes are in a few-shot learning situation. In this case, the word-based
approach is based on context for predictions, and failing that relies on character recognition.
        </p>
        <p>
          Lucas et al. have notably demonstrated that a fine-tuning strategy limited to 10 images (160
transcribed lines on average) for the Arabic Maghribi scripts, on the basis of a RASAM-trained
model is sufÏcient to reach a CER below 10% and to shorten the transcription work [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>
          We are taking this fine-tuning approach from the RASAM model and testing it on two
samples: one in-domain sample, derived from RASAM 1 and RASAM 2, and one out-of-domain
sample derived from manuscripts from Lucas et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] (see Figure 3). The latter dataset is
twice out-of-domain, with new scripts and new lexicon. We compare this new model with the
one strictly trained on RASAM 1 (see Figure4 and Table 4).
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Results</title>
        <p>In out-of-domain documents but belonging to the same family of scripts as RASAM 1 and 2,
such as the Arabic Maghribi scripts, RASAM 2 demonstrates notable efÏciency, as evidenced
in its application to TARIMA. Particularly noteworthy is its performance on Oriental scripts
(RASM and Iskandar), where RASAM 2 not only outperforms RASAM 1 but also achieves
significantly lower average CER scores (20.34 for RASM and 16.73 for Iskandar). These improved
results not only enhance accuracy but also facilitate faster processing with minimal data
requirements.</p>
        <p>Besides the versatility of RASAM 2 model, Figure4 also shows its robustness with a very
consistent CER per page and very little dispersion as in the case of RASAM 1. It is
particularly visible on RASAM 2 dataset for which RASAM 1 model (out-of-domain test) reaches a
CER between 11.67% (on the manuscript BULAC.MS.ARA.1982) and 48.80% (on the manuscript
BULAC.MS.ARA.9).</p>
        <p>A contrario, the CER of RASAM 2 model ranges between 1.71% and 28.47% in an in-domain
instance, and between 7.26% and 26.88% in an out-of-domain instance. The extreme values are
therefore practically twice as small as those for RASAM 1. Thus, there remain pages for which
our new model does not immediately succeed in producing workable outcome, for these pages,
it will then be necessary to adopt a fine-tuning strategy, which should be fast. 4 The median
observed in Figure 5 is 27.97% for RASAM 1 for out-of-domain documents, and is reduced to
15.83% for RASAM 2, hence a 42% decrease in the error rate.</p>
        <p>Figure 5 presents the average CER for each manuscript. In the in-domain instance,
several manuscripts have a CER of less than 5%: this is the case for the manuscripts
BULAC.MS.ARA.1943 (3.43%), BULAC MS ARA 1977 (4.91%), BULAC. MS.ARA.1982 (3.26%),
BULAC.MS.ARA.1983 (3.58%), and BULAC MS ARA 45b (3.20%). The BULAC.MS.ARA.1936 and
BULAC.MS.ARA.1947 manuscripts, even if they largely benefit from the new model, retain a
high CER, higher than 15% and up to 16.25% for the BULAC.MS.ARA .1936 (compared with the
46.47% CER achieved with RASAM 1, but which is out-of-domain).
4In Lucas et al., a CER of 3.23% was reached with a diferent split and a slightly redesigned architecture, based on a
meta-word-based approach (in the context of a specialized in-domain model). It also shows in particular that for
the manuscript BULAC.MS.ARA.1957, the initial CER of 30.46% (RASAM 1) is reduced to 21.8% after a fine-tuning
of only 20 lines. Applied to the same manuscript (see Figure5), RASAM 2 model obtains an initial CER of 25.5%9[].</p>
        <p>In the out-of-domain instance, the gap between the results of RASAM 1 and RASAM 2 is
narrower. If the manuscripts BULAC.MS.ARA.1922 (31.44% vs 26.38%) and BULAC.MS.ARA.1957
(35.95% vs 26.33%) retain a very high CER, the manuscripts BULAC.MS.ARA.1944 and BULAC.
MS.ARA.1929 achieve a CER of 7.67% and 10.16%, better than the CER obtained in-domain for
the manuscripts previously cited.</p>
        <p>Despite the diversity of the TARIMA corpus, with both manuscripts and lithographs, the
results remain very good. This is due to the proximity between the RASAM 1 &amp; 2 dataset and
the palaeographic characteristics of the TARIMA corpus, all of which are in Maghribi script.
4.2.2. Out-of-domain results (Oriental scripts)
Out-of-domain results (Oriental scripts) RASAM 2 also demonstrates significantly enhanced
efÏciency when applied to Oriental manuscripts, as illustrated by its performance with RASM
and Iskandar. Its versatility is particularly evident in Iskandar, where the CER remains
below 30%, with an average CER ranging between 8% and 20% (Fig. 4 and 5). Except for one
manuscript (MS_Orient_A_02393), all the CER remain below 20% with RASAM 2. While
RASM results exhibit some dispersion (albeit less than with RASAM 1), RASAM 2’s
performance varies across the four manuscripts comprising the RASM dataset. Its highest result is
observed in Dehli_Arabic_1901 (slightly above 16%), but none exceed 25%. The disparity in
out-of-domain results between RASM and Iskandar likely arises from the diference in dataset
adherence to RASAM guidelines. While Iskandar follows the RASAM guidelines, the RASM
dataset diverges from them, which may explain the observed gap in CER results. For
example, when the scribe omitted expected diacritics on certain letters, the transcriber left the letter
without them, whereas the RASAM guidelines would have added the diacritics where
necessary. This suggests that with minimal fine-tuning, RASAM 2 could readily adapt to various
manuscripts, regardless of their script families.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Qualitative interpretation</title>
        <p>RASAM 2 sets a new standard for the recognition of Arabic Maghribi scripts. Figure5 shows
that it nevertheless produces many more errors than the average on four in-domain and
outof-domain manuscripts, leading to an increase in the CER. Observation of the manuscripts (see
Figure 6) reveals several situations where the CER decreases naturally.</p>
        <p>
          Manuscript with vowel signs and numerous interlinear notes: This is the case of the
manuscripts BULAC.MS.ARA.1936 and BULAC.MS.ARA.1957 for which we observe an
important vocalization which is rarely present in these manuscripts. It leads, at this stage, to a greater
ambiguity of the forms to be recognized, but is however not insurmountable: a specialized
approach from RASAM shows for example that 160 lines are enough with a word-based approach
to reach a CER of 10.41% for the manuscript BULAC.MS.ARA.1957 [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>
          Variation in line color: This is a phenomenon already observed in RASAM 1 1[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], with
an over-representation of colored lines among lines with high CER. The MS.ARA.1947, which
alternates blue and red lines (marginally present in training) is therefore penalized. Its CER
drops to 6.56% without these lines.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In conclusion, the RASAM 2 dataset ofers a high representativeness of Arabic Maghribi scripts.
The word-based model trained on this dataset obtains very high in-domain and out-of-domain
accuracies, achieving a 40-point CER reduction in all scenarios, which ensures an important
coverage of Arabic Maghribi manuscript traditions. The dataset also demonstrates its
versatility and can be easily fine-tuned on a new target, including Oriental scripts and new varieties of
Arabic (Middle Arabic, Berber written in Arabic). In the future, we will study this transfer of
RASAM models to other types of Arabic scripts, in particular Oriental ones. Additionally, we
plan to conduct experiments using transformer-based models, as the critical mass of data for
Arabic has now been reached, thanks to the RASAM team and all datasets produced within this
scope. More generally, the datasets created in recent years around the RASAM team (TARIMA,
Iskandar) have made it possible to create a set of open data decisive for the HTR of Arabic
scripts.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was carried out within the framework of cooperation between the Research
Consortium Middle-East and Muslim Worlds (GIS MOMM), the BULAC, and Calfa. It aligns with the
scientific focus defined by the GIS MOMM, which prioritizes North African studies and digital
humanities.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Data availability</title>
      <p>• RASAM 1 and 2 datasets: https://github.com/calfa-co/rasam-dataset
• TARIMA dataset: https://github.com/calfa-co/tarima
• Iskandar dataset: https://gitlab.huma-num.fr/lipa/iskandar</p>
    </sec>
    <sec id="sec-8">
      <title>B. Paleographical features of RASAM 2 dataset</title>
      <p>èYë
½ËX
ÈA¯</p>
      <sec id="sec-8-1">
        <title>RASAM 1 (Magribi script) MS.ARA.417 MS.ARA.609 MS.ARA.1977</title>
        <p>MS.ARA.6
MS.ARA.9
MS.ARA.23
MS.ARA.24
MS.ARA.45b
MS.ARA.65
MS.ARA.1926
MS.ARA.1936
MS.ARA.1943
MS.ARA.1944</p>
      </sec>
      <sec id="sec-8-2">
        <title>RASAM 2 (Magribi script)</title>
        <p>Add.MS.7474
Add.MS.23494
Dehli.Arabic.1901
OR.3366</p>
      </sec>
      <sec id="sec-8-3">
        <title>Orient.A.0238X</title>
        <p>Btv1b10031486f
Btv1b11003097s
Btv1b110032696
Leeds.ms153</p>
      </sec>
      <sec id="sec-8-4">
        <title>RASM (Oriental script) Iskandar (Oriental script)</title>
        <p>TOTAL
250
3,750
5,702
simple
simple
simple
simple
medium
simple
simple
simple
simple
simple
complex
high
complex
high
low
low
low
low
low
low
low
low
high
high
simple
very low
damaged
Litany
medium
low
damaged
complex
high
damaged
good
good
good
good
good
good
good
good
good
good
good
good</p>
        <p>Law
Grammar</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdelhaleem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Droby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Asi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kassis</surname>
          </string-name>
          , R. Al Asam,
          <article-title>and</article-title>
          <string-name>
            <surname>J. El-sanaa.</surname>
          </string-name>
          “
          <article-title>Wahd: a database for writer identification of arabic historical documents”</article-title>
          .
          <source>In: 2017 1st International workshop on arabic script analysis and recognition (ASAR)</source>
          .
          <source>Ieee</source>
          .
          <year>2017</year>
          , pp.
          <fpage>64</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Adam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Al-Maadeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bouridane</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>El-Menshawy</surname>
          </string-name>
          .
          <article-title>“KERTAS: dataset for automatic dating of ancient Arabic manuscripts”</article-title>
          .
          <source>In:International Journal on Document Analysis and Recognition (IJDAR) 21</source>
          (
          <year>2018</year>
          ), pp.
          <fpage>283</fpage>
          -
          <lpage>290</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Clausner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Antonacopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mcgregor</surname>
          </string-name>
          , and
          <string-name>
            <surname>D.</surname>
          </string-name>
          Wilson-Nunn. “
          <article-title>Icfhr 2018 competition on recognition of historical arabic scientific manuscripts-rasm2018”</article-title>
          .
          <source>In: 2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR)</source>
          .
          <source>Ieee</source>
          .
          <year>2018</year>
          , pp.
          <fpage>471</fpage>
          -
          <lpage>476</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Djeddi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gattal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Souici-Meslati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Siddiqi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chibani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. El</given-names>
            <surname>Abed</surname>
          </string-name>
          . “
          <article-title>LAMISMSHD: A Multi-script OfÒine Handwriting Database”</article-title>
          .
          <source>In: 2014 14th International Conference on Frontiers in Handwriting Recognition</source>
          .
          <year>2014</year>
          , pp.
          <fpage>93</fpage>
          -
          <lpage>97</lpage>
          . doi:
          <volume>10</volume>
          .1109/icfhr.
          <year>2014</year>
          .
          <volume>23</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eltay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zidouri</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Ahmad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Elarian</surname>
          </string-name>
          . “
          <article-title>Generative adversarial network based adaptive data augmentation for handwritten Arabic text recognition”</article-title>
          .
          <source>InP:eerJ Computer Science</source>
          <volume>8</volume>
          (
          <year>2022</year>
          ),
          <year>e861</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-B. Camps</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Pinche</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Jahan</surname>
          </string-name>
          . “
          <article-title>SegmOnto: common vocabulary and practices for analysing the layout of manuscripts (and more)”</article-title>
          .
          <source>In:1st International Workshop on Computational Paleography (IWCP ICDAR</source>
          <year>2021</year>
          ).
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kassis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdalhaleem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Droby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alaasam</surname>
          </string-name>
          , and J.
          <string-name>
            <surname>El-Sana</surname>
          </string-name>
          .
          <article-title>“Vml-hd: The historical arabic documents dataset for recognition systems”</article-title>
          .
          <source>In2:017 1st international workshop on Arabic script analysis and recognition (ASAR)</source>
          .
          <source>Ieee</source>
          .
          <year>2017</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kiessling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S. B.</given-names>
            <surname>Ezra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Miller</surname>
          </string-name>
          . “
          <article-title>BADAM: a public dataset for baseline detection in Arabic-script manuscripts”</article-title>
          .
          <source>In:Proceedings of the 5th International Workshop on Historical Document Imaging and Processing</source>
          .
          <year>2019</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Lucas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Salah</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Vidal-Gorène</surname>
          </string-name>
          .
          <article-title>“New Results for the Text Recognition of Arabic Maghribi Manuscripts - Managing an Under-resourced Script”</article-title>
          .
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Maadeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ayouby</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Hassaın̈e, and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Aljaam</surname>
          </string-name>
          . “
          <article-title>QUWI: An Arabic and English Handwriting Dataset for OfÒine Writer Identification”</article-title>
          .
          <source>In: 2012 International Conference on Frontiers in Handwriting Recognition</source>
          .
          <year>2012</year>
          , pp.
          <fpage>746</fpage>
          -
          <lpage>751</lpage>
          . doi:
          <volume>10</volume>
          .1109/icfhr.
          <year>2012</year>
          .
          <volume>256</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Mahmoud</surname>
          </string-name>
          , I. Ahmad,
          <string-name>
            <given-names>W. G.</given-names>
            <surname>Al-Khatib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alshayeb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Tanvir</given-names>
            <surname>Parvez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Märgner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Fink</surname>
          </string-name>
          . “KHATT:
          <article-title>An open Arabic ofÒine handwritten text database”</article-title>
          .
          <source>In: Pattern Recognition 47.3</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>1096</fpage>
          -
          <lpage>1112</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.patcog.
          <year>2013</year>
          .
          <volume>08</volume>
          .009.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Maouloud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. O. M.</given-names>
            <surname>Dyla</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Ba</surname>
          </string-name>
          . “
          <article-title>Transformer-based Model For Handwritten Recognition Arabic Words Al-soudani Maghrebi Script”</article-title>
          .
          <source>InJ:ournal of Theoretical and Applied Information Technology 101.24</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13] [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Pantke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dennhardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Märgner</surname>
          </string-name>
          , and
          <string-name>
            <surname>T. Fingscheidt. “</surname>
          </string-name>
          <article-title>An historical handwritten arabic dataset for segmentation-free word spotting-hadara80p”</article-title>
          .
          <source>In2:014 14th International Conference on Frontiers in Handwriting Recognition. Ieee</source>
          .
          <year>2014</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Pechwitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Maddouri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Märgner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ellouze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Amiri</surname>
          </string-name>
          , et al. “
          <article-title>IFN/ENIT-database of handwritten Arabic words”</article-title>
          .
          <source>In:Proc. of CIFED</source>
          . Vol.
          <volume>2</volume>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          .
          <year>2002</year>
          , pp.
          <fpage>127</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Vidal-Gorène</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dupin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Decours-Perez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Riccioli</surname>
          </string-name>
          .
          <article-title>“A Modular and Automated Annotation Platform for Handwritings: Evaluation on Under-Resourced Languages”</article-title>
          . In:
          <article-title>Document Analysis</article-title>
          and
          <string-name>
            <surname>Recognition - ICDAR</surname>
          </string-name>
          <year>2021</year>
          . Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Lladós</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lopresti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Uchida</surname>
          </string-name>
          . Cham: Springer International Publishing,
          <year>2021</year>
          , pp.
          <fpage>507</fpage>
          -
          <lpage>522</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Vidal-Gorène</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lucas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Salah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Decours-Perez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Dupin</surname>
          </string-name>
          . “
          <article-title>RASAM - A Dataset for the Recognition and Analysis of Scripts in Arabic Maghrebi”</article-title>
          .
          <source>InD:ocument Analysis and Recognition - ICDAR</source>
          <year>2021</year>
          Workshops. Ed. by
          <string-name>
            <given-names>E. H. Barney</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>U.</given-names>
            <surname>Pal</surname>
          </string-name>
          . Cham: Springer International Publishing,
          <year>2021</year>
          , pp.
          <fpage>265</fpage>
          -
          <lpage>281</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -86
          <fpage>198</fpage>
          -
          <lpage>8</lpage>
          \_
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          BULAC.MS.
          <source>ARA.6</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/b6q5p6 BULAC.MS.
          <source>ARA.9</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bqnm44 BULAC.MS.ARA.
          <volume>23</volume>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bnvxrc BULAC.
          <source>MS.ARA.24</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>https://bina.bulac.fr/s/bina/ark:/73193/bsn0x6 BULAC.MS.ARA.45b</mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/brv21m BULAC.MS.ARA.
          <volume>65</volume>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bnckkp BULAC.MS.ARA.
          <year>1926</year>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/b7d8bx BULAC.MS.ARA.
          <year>1936</year>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bcc306 BULAC.MS.ARA.
          <year>1943</year>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/b5x6qh BULAC.MS.ARA.
          <year>1944</year>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bj10cm BULAC.MS.ARA.
          <year>1946</year>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/b8pkg9 BULAC.MS.ARA.
          <year>1947</year>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bdfnnt BULAC.MS.ARA.
          <year>1960</year>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bstrdn BULAC.MS.ARA.
          <year>1982</year>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/73193/bvmdrp BULAC.MS.ARA.
          <year>1983</year>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          https://bina.bulac.fr/s/bina/ark:/
          <source>73193/bz8x88 14 14 14 14 12 12 34 16 20 34 14 13 13 14 12</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>