<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards the Corpus of Latvian Romani Texts: Deciphering the Manuscripts in Jānis Leimanis' Archive</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Natalia Perkova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kirill Kozhanov</string-name>
          <email>kozhanov@uni-potsdam.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Handwritten Text Recognition</institution>
          ,
          <addr-line>Low-resource Languages, Digitization, Latvian Romani</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Helsinki University</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Potsdam University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Uppsala University</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>1</volume>
      <fpage>5</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Latvian Romani is a Northeastern Romani dialect with a limited number of publicly available sources. Two large archival collections of texts in Latvian Romani, compiled primarily in the 1930s in Latvia and Estonia, have been recently digitized as images and made available online for a wider public. In our study, we focus on one of these collections, the Latvian Romani folklore texts collected by Jānis Leimanis in interwar Latvia. In this paper, we describe how initial manual transcriptions, most of which have been created with the help of a special crowdsourcing platform, were integrated in the handwritten text recognition (HTR) workflow in Transkribus. We present two HTR models trained on the basis of Leimanis' collection and discuss various issues related to the work on these texts. Folklore Texts (http://garamantas.lv/en/collection/886320/Romani-folklore-collection-of-Janis-Leimanis), (https://fennougrica.kansalliskirjasto.fi/handle/10024/87064).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        This study focuses on the development of a corpus for Latvian Romani (Lotfitka), one of the
understudied Romani dialects. It belongs to the group of Northeastern Romani dialects [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] and is
spoken in Latvia, Estonia, and northern Lithuania. There exist only a handful of texts published in
Latvian Romani (a translation of the Gospel of John by Jānis Leimanis (1933) and several fairy tales
and legends in [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ]). The data in the Romani-Latvian-English dictionary [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the Romani
Morpho-Syntax Database (https://romani.humanities.manchester.ac.uk/rms/, see [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) contain only
words and separate unrelated phrases. Our attempt at creating a corpus of Latvian Romani texts has
been encouraged by recent digitization initiatives in several countries (Estonia, Latvia, and Finland)
which resulted in providing open access to the two important archives of Latvian Romani texts, both
compiled before the World War II: the collection by Jānis Leimanis, a prominent Latvian Romani
personality
the
interwar
      </p>
      <sec id="sec-1-1">
        <title>Latvia, at the</title>
      </sec>
      <sec id="sec-1-2">
        <title>Archive of</title>
      </sec>
      <sec id="sec-1-3">
        <title>Latvian folklore (http://garamantas.lv/en/collection/886320/Romani-folklore-collection-of-Janis-Leimanis), in</title>
        <p>and</p>
        <p>Riga
the
collection by Paul Ariste, a brilliant Estonian linguist, archived in the Estonian literary museum and
available
online
at
the</p>
      </sec>
      <sec id="sec-1-4">
        <title>National library of</title>
      </sec>
      <sec id="sec-1-5">
        <title>Finland</title>
        <p>
          Digitization of manuscripts with field notes in an endangered, understudied and obviously
lowresource language variety using automatic handwritten recognition still represents a rather novel
direction in applying computational methods to such data. Among numerous public models available
in Transkribus, the Evenki-Russian bilingual model trained on Konstantin Rychkov’s manuscripts from
the 1910s seems to be the only example of similar research, see [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for more details on this project. In
our case, the texts in focus are also bilingual and handwritten, collected by the same person about a
        </p>
        <p>
          2022 Copyright for this paper by its authors.
century ago and not previously published in full. Another important recent initiative focusing on
digitizing old manuscripts with texts in minor languages is Manuscripta Castreniana
(https://www.sgr.fi/manuscripta/), an impressive collection of texts transcribed, translated and
commented by the linguist Matthias Alexander Castrén in the middle of the 19th century, see [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for an
overview.
2. Jānis Leimanis and his Latvian Romani folklore collection
        </p>
        <p>
          Jānis Leimanis (1886-1950) was a prominent Romani figure of interwar Latvia (see [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]). Founder
of the first Romani organization in Latvia, translator and activist, he was officially involved in the work
of the Archives of Latvian Folklore (Latvijas Folkloras Krātuve, further LFK) in 1933–1934. In 1939,
Leimanis’ son Juris published a fiction book titled “Gypsies in Latvia’s forests, homes and markets”,
primarily based on the materials collected by his father (it was recently reprinted with some
commentaries by Māra Vīksna [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]). This book, however, was published in Latvian and did not truly
present any text from the collection at its length.
        </p>
        <p>Leimanis’ archive in LFK comprises 75 copybooks (three of them currently unavailable), with about
500 folklore units of different genres (463 of them accessible) and 1254 manuscript pages in total. A
diversity of genres is impressive: there are about 65 longer narrative texts (fairy tales and stories), a
number of songs, short forms (often incorporated in longer texts, but still classified as separate folklore
units), description of traditional practices (e.g., Latvian pirts ‘sauna’), proverbs and parables, as well as
an appendix with a list of obsolete and disappearing words. This collection is mostly bilingual, as all
Romani texts have Latvian translations provided by Leimanis himself: overall, 884 pages contain such
bilingual texts. This gives us a unique opportunity to compile a parallel corpus for a Romani dialect.</p>
        <p>
          In 2014, LFK launched a massive crowdsourcing campaign on deciphering the handwritten
manuscripts which were scanned and uploaded to a special platform (http://garamantas.lv/en/). This
initiative became a huge success in terms of the citizens’ voluntary engagement into the preservation
of cultural heritage (see more in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]). As part of this campaign, Leimanis’ archive was digitized and
uploaded to the crowdsourcing platform as a separate collection
(http://garamantas.lv/lv/collection/886320/Jana-Leimana-ciganu-folkloras-vakums). Since then, some
files have been deciphered by volunteers, but by October 2020 the deciphered pages comprised about
25% of all files and only about 21% of files with Romani text. Although a dozen of languages spoken
by the ethnic groups of Latvia are presented in the LFK materials, most volunteers do not master these
languages and pay their attention primarily to the texts written in Latvian. This is one of the reasons
why the Latvian Romani collection remained mostly non-deciphered. Our original question was
whether it would be possible to accelerate this process, bearing in mind that the entire collection was
written by the same person, or, in other words, it has the same handwriting.
        </p>
        <p>The orthography used by Leimanis for Romani texts is based on Latvian; no extra symbols or
diacritics (for instance, to mark stress) are used. Leimanis’ handwriting is very clear, and the text is
almost always very accurately adjusted to the lines. No doubt, such handwritten texts make an ideal
case for automatic recognition.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Transkribus and the Latvian Romani HTR pipeline</title>
      <p>
        One of the well-known and freely available HTR tools is Transkribus
(https://transkribus.eu/Transkribus/). It is possible to use it as a desktop application with the access to
the server or work directly in the browser in Transkribus Lite. The platform can be used for transcribing
texts and training various text recognition models on them. Transcription can be manual or automatic,
conducted with the help of the existing models (OCR models for printed texts are also available for
users), see more details in [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. By now, 119 public models for at least a dozen of languages are
available for Transkribus users; a smaller number of 87 models are presented at the website2. Currently,
there are no models available for Latvian print or handwritten texts, not to mention Latvian Romani.
      </p>
      <p>In January 2021, we intensively worked on transferring the deciphered texts to the Transkribus
platform as ground truth data for the Leimanis collection. Preliminary transcription experiments were
conducted already in 2020. At the initial stage, the texts were taken from finished garamantas.lv
transcriptions, with some minor editing in most obvious cases. The almost perfect quality of the
transcriptions available at garamantas.lv accelerated the process of ground truth preparation. The major
difficulty was related to the distribution of transcription blocks between the lines, as Transkribus
requires a very detailed image-based reproduction of texts in transcriptions. In contrast, the guidelines
for crowdsourcers ask for normalized form of erratic words to be put in square brackets. This
information is not based on the scanned images, and such additions had to be removed from the
transcriptions.</p>
      <p>
        Even though this work was initially conducted just by one person, a trained linguist with only
theoretical knowledge of Latvian Romani, we hope that access to the resources on this and other
dialects, most importantly, the dictionary [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], allowed for the correct interpretation of most cases, and
that the number of errors is minimal. At the same time, this person speaks Latvian at an advanced level,
which allowed for more thorough checking of the Latvian part of the texts. At the later stage, two
researchers who specialize in Romani dialects, including Lotfitka, joined the process of deciphering,
which accelerated proofreading of automatically recognized pages. These pages were also lately
approved as ground truth by the more experienced transcriber after checking the proofread pages. In
this way, we managed to proofread several more folklore units by the middle of 2021. The priority was
given to bilingual texts, with monolingual ones left for the later deciphering stage.
      </p>
      <p>Due to the standard format of the copybooks used by Leimanis, the standard size of the files is about
20–34 lines (usually two-page layouts with the Latvian Romani text on the left and the Latvian
translation on the right side). It takes about 6–10 minutes to copy and paste a previously deciphered text
and distribute it correctly across recognized lines, as well as to correct minor things (e.g., add
strikethrough annotation for some words which were not relevant for the transcribers, but which are
more important for the HTR training in Transkribus). Figure 2 shows what the text from Figure 1 looks
like after having been transcribed.</p>
      <p>It is claimed that training a good Transkribus model requires about 15000 words, or 75 pages, of
ground truth material. In our case, an average file with a two-page layout has about 160 words. This
gives the following preliminary calculation: at least 94 files (15000/163) are needed to train the HTR
model. As there were more deciphered files available at garamantas.lv, it was decided to add at least
200 deciphered pages as ground truth transcriptions.</p>
      <p>The first model, Leimanis_test, was trained on 212 pages and validated on 23 pages. About 30 pages
in the training data are copybook covers, which have only several shortened lines and describe metadata
on the copybook content. The second model, Leimanis_updated, has the first model as its base model.
Only several pages were additionally transcribed and used as ground truth data. The proportion of train
and validation data was shifted to the improvement of the latter in terms of size. Still, using the base
model and a higher number of epochs (100 in the second model vs. 50 in the first model) resulted in a
considerable improvement of quality, as character error rate (CER) was already below the threshold of
5%. Both models were trained with the CITlab HTR+ method. The comparison of our models is given
in Table 1 below.</p>
      <p>A preliminary layout analysis had been previously performed with the help of the CITlab Advanced
tool with default settings. For ground truth pages, it was always manually corrected, when necessary,
but for many pages it is still left uncorrected: this is seen as a preparatory step before each automatic
page recognition procedure, so it is conducted consecutively.</p>
      <p>After obtaining the HTR model and transcribing the rest of the files automatically, one should have
a look at the quality of the obtained transcriptions. There is some chance that the initial quality of the
model would not be enough, then adding more manual transcription could probably improve the
training. In any case, in order to continue work with the texts, one needs to correct all possible errors in
automatic transcriptions. The project page gives information on character error rate about 5% for some
of the available models, which, of course, requires some postprocessing of the automatically annotated
files.</p>
      <p>For evaluation of the quality of recognition of the two available models and their comparison, we
launched HTR on a page outside the ground truth data (page 1245 in our collection, originally notebook
68, page 53). It is a fragment of the unit 426 (http://garamantas.lv/lv/unit/400700/LFK-1389-426), a
fairytale ‘How a baron married a Roma girl, and how she preferred forest and her people’ (“Kā skaisto
čigānieti apprecēja barons, un ka tai mežs un tautieši labāk patika...” / “Sir šukārune romane čha lija
baronos, te sir lake, vešs te roma fidīr kamdžapes nasir filačin te...”).</p>
      <p>Figure 3 shows the result of automatic layout analysis (CITlab Advanced with default settings). The
text in notebooks is very consistently and clearly written line-to-line, and corrections are minor, in this
case only as a superscript word in line 5 on the right page. The layout recognition works very well, and
in this case all the errors are minor and expected and belong to the regularly occurring types. First, the
very first line is split into two parts; this happens occasionally, but still not too frequently to noticeably
disturb the process of post-correction. Second, superscripts are usually assigned separate lines in
automatic layout analysis. This is not incorrect from a purely visual perspective, but for structural
reasons we want to incorporate such superscript words in the lines where they truly belong by merging
them with the next line. As Transkribus provides additional markup for superscript and subscript (also
for bold, italic, underlined, and strikethrough, among others), the post-corrected result can be easily
used for training next models and for the restoration of the original.
3Here by pages we mean double pages as scanned and uploaded to garamantas.lv; page numbers are also reflected in the file names
(1389-68-05 at the website and 1389_68_05.jpg in our collection).</p>
      <p>See the tutorial available at
https://readcoop.eu/transkribus/howto/how-to-train-a-handwritten-text-recognition-model-intranskribus/ (accessed on 15/02/2021).</p>
      <p>WER
2.70
2.36
4.73
5.41</p>
      <p>Recognition errors are often related to diacritics either in the Latvian or the Latvian Romani text;
sometimes diacritics can also be incorrectly added, as in our case with sūne (incorrect šūne). In addition,
some problematic cases are really challenging, as they are related to less standard letter combinations
occasionally made in handwritten texts (e.g., when one letter is distorted, or when borderlines between
adjacent letters are not as clear as they normally are). Sometimes, capitalization errors are made, if
Leimanis doesn’t make clear size difference between the first capital letter and the following one (see
the error Viņa / viņa in Figure 4). Finally, sometimes different letters are indeed similar in their graphic
forms, therefore automatic recognition models make mistakes like with human recognition, as we as
transcribers also sometimes struggle, trying to decipher some letters; in our case, the capitalized F in
the word Fidir got interpreted as either L or T.</p>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusions</title>
      <p>It has been shown how the opportunities provided by Transkribus can be successfully exploited with
respect to unique handwritten Latvian Romani data. Initial volunteer-based manual deciphering efforts
have been successfully integrated in the development of automatic handwritten recognition for Latvian
Romani; we are thankful for being granted access to such preliminary ground truth data.</p>
      <p>The character of handwriting and accurate original copybook layout made it possible to train
recognition models, the best of which has very satisfactory quality, so that the further deciphering
process is based on post-correction of automatically recognised pages. Errors made by the model
Leimanis_updated are not numerous and in many cases can be compared to similar difficulties faced
by ordinary transcribers.</p>
      <p>We hope that our initial effort at digitizing Leimanis’ archive will result in the development of a
bigger Latvian Romani corpus. This would, of course, imply more serious normalization of texts,
additional translations (at least in English), developing morphological annotation for Latvian Romani,
etc. The bilingual character of Leimanis’ collection makes it possible to consider the compilation of a
parallel corpus, with such options as, for instance, word alignment. It is also worth mentioning that for
several texts later translations are available in [15] and in the materials from Pasakas.net5. Various
Latvian Romani materials are already available in our repository
(https://github.com/LatvianRomani/Lotfitka).</p>
    </sec>
    <sec id="sec-4">
      <title>5. Acknowledgements</title>
      <p>We are thankful to Sanita Reinsone for giving access to the collection and kind support and to Ieva
Tihovska for her help with text transcriptions and sharing her knowledge about the history of Latvian
Roma people and Jānis Leimanis in particular. We would also like to recognize the contribution by
Anette Ross who helped us with proofreading the automatically recognized transcriptions and gave us
access to various texts in Latvian Romani.</p>
    </sec>
    <sec id="sec-5">
      <title>6. References</title>
      <p>5As the original domain was taken by other people, a duplicated archival copy is currently stored at asakas.net.
Heyer, L. Hirvonen, T. Hodel, M. Jokinen, Ph. Kahle, M. Kallio, Fr. Kaplan, Fl. Kleber, R. Labahn,
E. M. Lang, S. Laube, G. Leifert, G. Louloudis, R. McNicholl, J.-L. Meunier, J. Michael, E.
Mühlbauer, N. Philipp, I. Pratikakis, J. Puigcerver, Pérez, H. Putz, G. Retsinas, V. Romero, R.
Sablatnig, J. Andreu Sánchez, Ph. Schofield, G. Sfikas, Chr. Sieber, N. Stamatopoulos, T. Strauß,
T. Terbul, A. Héctor Toselli, B. Ulreich, M. Villegas, E. Vidal, J. Walcher, M. Weidemann, H.
Wurster, K. Zagoris. Transforming scholarship in the archives through handwritten text
recognition: Transkribus as a case study, Journal of Documentation 75, 5 (2019) 954-976. doi:
10.1108/JD-07-2018-0114
[15] S. Brice (Ed.), Laimes puķe: čigānu tautas pasakas, Sprīdītis, Rīga, 1992.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matras</surname>
          </string-name>
          ,
          <article-title>Romani: A linguistic introduction</article-title>
          , Cambridge University Press, Cambridge,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Tenser</surname>
          </string-name>
          , Northeastern Group of Romani Dialects,
          <source>Ph.D. thesis</source>
          , The University of Manchester,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ariste</surname>
          </string-name>
          , Romenge paramiši: Mustlaste muinasjutte, Tartu,
          <year>1938</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ariste</surname>
          </string-name>
          ,
          <article-title>Supplementary review concerning the Baltic Gypsies and their dialect</article-title>
          ,
          <source>Journal of the Gipsy Lore Society</source>
          , Third Series,
          <volume>43</volume>
          ,
          <issue>1</issue>
          /2 (
          <year>1964</year>
          )
          <fpage>35</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ariste</surname>
          </string-name>
          ,
          <string-name>
            <surname>Einige Märchen</surname>
          </string-name>
          Čuchný-Zigeuner,
          <source>Tartu Riikliku Ülikooli Toimetised</source>
          <volume>309</volume>
          (
          <year>1973</year>
          )
          <fpage>5</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Mānušs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Neilands</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rudevičs</surname>
          </string-name>
          ,
          <article-title>Čigānu-latviešu-angļu un latviešu-čigānu vārdnica</article-title>
          ,
          <string-name>
            <surname>Zvaigzne</surname>
            <given-names>ABC</given-names>
          </string-name>
          , Riga,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matras</surname>
          </string-name>
          , Ch. White,
          <string-name>
            <given-names>V.</given-names>
            <surname>Elšik</surname>
          </string-name>
          ,
          <article-title>The Romani morpho-syntax (RMS) database</article-title>
          , in: M.
          <string-name>
            <surname>Everaert</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Musgrave</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Dimitriadis (Eds.),
          <article-title>The use of databases in crosslinguistic studies</article-title>
          , Mouton de Gruyter, Berlin,
          <year>2009</year>
          , pp.
          <fpage>329</fpage>
          -
          <lpage>362</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Arkhipov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barinskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shtefura</surname>
          </string-name>
          ,
          <article-title>Using handwritten text recognition on bilingual EvenkiRussian manuscripts of Konstantin Rychkov</article-title>
          ,
          <source>Scripta &amp; eScripta</source>
          <volume>21</volume>
          (
          <year>2021</year>
          )
          <fpage>233</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Partanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rueter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hämäläinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Alnajjar</surname>
          </string-name>
          ,
          <string-name>
            <surname>Processing M.</surname>
          </string-name>
          <article-title>A. Castrén's materials: Multilingual typed and handwritten manuscripts</article-title>
          , in: M.
          <string-name>
            <surname>Hämäläinen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Alnajjar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Partanen</surname>
          </string-name>
          , J. Rueter (Eds.),
          <source>Proceedings of the Workshop on Natural Language Processing for Digital Humanities</source>
          ,
          <article-title>The Association for Computational Linguistics</article-title>
          , Stroudsburg,
          <year>2021</year>
          , pp.
          <fpage>47</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>I. Tihovska</surname>
          </string-name>
          ,
          <article-title>Jānis Leimanis and the Beginnings of Latvian Roma Activism</article-title>
          , in: E.
          <string-name>
            <surname>Marushiakova</surname>
          </string-name>
          , V. Popov (Eds.), Roma Portraits in History: Roma Civic Emancipation Elite in Central,
          <source>SouthEastern and Eastern Europe from the 19th Century until World War II, Brill | Schöningh, Leiden</source>
          ,
          <year>2022</year>
          . doi: https://doi.org/10.30965/9783657705191_
          <fpage>011</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Leimanis</surname>
          </string-name>
          , Čigāni Latvijas mežos, mājās un tirgos, Zinātne, Rīga,
          <year>2005</year>
          [1939].
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Reinsone</surname>
          </string-name>
          ,
          <article-title>Searching for deeper meanings in cultural heritage crowdsourcing</article-title>
          , in: P. Hetland,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pierroux</surname>
          </string-name>
          , L. Esborg (Eds.),
          <article-title>A History of Participation in Museums and Archives: Traversing Citizen Science</article-title>
          and Citizen Humanities, Routledge,
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .4324/
          <fpage>9780429197536</fpage>
          -
          <lpage>14</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Colutto</surname>
          </string-name>
          , G. Hackl, G. Mühlberger,
          <article-title>Transkribus - a platform for transcription, recognition and retrieval of document images</article-title>
          ,
          <source>IAPR International Conference on Document Analysis and Recognition (ICDAR)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Mühlberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Seaward</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Terras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ares</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Colutto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Déjean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fiel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Greinoecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grüning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hackl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Haukkovaara</surname>
          </string-name>
          , G.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>