<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>In Codice Ratio : Scalable Transcription of Historical Handwritten Documents (Extended Abstract)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Serena Ammirati</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Donatella Firmani</string-name>
          <email>donatella.firmani@uniroma3.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Maiorino</string-name>
          <email>m.maiorino@asv.va</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Merialdo</string-name>
          <email>merialdo@dia.uniroma3.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Nieddu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Rossi</string-name>
          <email>and.rossi.516@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Roma Tre University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vatican Secret Archives</institution>
          ,
          <addr-line>Archivum Secretum Apostolicum Vaticanum</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Huge amounts of handwritten historical documents are being published by digital libraries world wide. However, for these raw digital images to be really useful, they need to be annotated with informative content. State-of-the-art Handwritten Text Recognition (HTR) approaches require an impressive training e ort by expert paleographers. Our contribution is a scalable, end-to-end transcription work- ow { that we call In Codice Ratio { based on ne-grain segmentation of text elements into characters and symbols, with limited training e ort. We provide a preliminary evaluation of In Codice Ratio over a corpus of letters by pope Honorii III, stored in the Vatican Secret Archive.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Large document collections are sources of important correlations between entities
such as people, events, places, and organizations. Previous studies [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] have shown
that it is possible to detect macroscopic patterns of cultural change over
periods of centuries by analyzing large textual time series. Such automatic methods
promise to empower scholars with a quantitative and data-driven tool to study
culture and society, but their power has been limited by the amount of digitally
transcribed sources. Indeed, the World Wide Web only contains a small part of
the traditional archives. (It is evocative to think that it may only contain a few
millimeters out of the 85km of linear shelves in the Vatican Secret Archives.)
Recently, many historical archives have begun to digitize their assets, sharing
high-resolution images of the original documents. Notable examples include the
Bibliothque Nationale de France3, the Virtual Manuscript Library of
Switzerland4, and the Vatican Apostolic Library5. In this scenario, expert paleographers
3 http://gallica.bnf.fr
4 http://www.e-codices.unifr.ch/en
5 http://www.digitavaticana.org
(a)
(b)
can largely bene t from computer-assisted transcription technologies. This
includes not only full transcriptions, but also partial transcription and other kinds
of automatically produced meta-data, useful for indexing and searching.
      </p>
      <p>
        Popular automatic tools for transcribing the text content of digital images
include Optical Character Recognition (OCR) systems, which work great for
typewritten text but are not suitable for handwritten text recognition (HTR).
Since most digitized documents by historical archives are manuscripts, HTR has
recently gained more and more attention by researchers worldwide. Handwritten
text is more challenging to recognize than typewritten one because characters
have less regular shapes, and are often combined into single units known as a
ligatures. Therefore, while OCR systems are trained to recognize individual
typewritten glyphs, most state-of-the-art HTR systems use holistic approaches: all
text elements (sentences, words, and characters) of a single text line are
recognized as a whole, without any prior segmentation of the line into these elements.
So-called segmentation free models [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] can be automatically obtained using well
known training techniques, but they require the whole transcripts of a number of
these unsegmented images. In order to use these technologies, users with
experience in handwritten documents transcription are required to transcribe
manually signi cant portions of the original documents. Currently available HTR
technologies are still far from o ering scalable automated solutions.
Our contribution. The approach of our project In Codice Ratio is in the
middle of a spectrum, where on the one side there are OCR systems and on the
other segmentation free HTR technologies, which recognize bigger handwritten
elements. Rather than relying on well-segmented glyphs (as in typewriting) or
training to recognize whole words, we focus on overlapping \fragments" of words
composed of zero, one, or two (rarely three) characters, as in Figure 1a. For each
word, we compute possible cut-points yielding di erent ways for segmenting the
word into fragments. Then, we choose the best among di erent segmentation
options using OCR and language models. Cut points are managed similarly to
on-line HTR systems [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], where the text is written on a touch-screen and a
sensor captures all the pen tip movements. Since we do not have access to pen
movements, we need a number of labelled fragments6 for training our model.
6 A fragment can also have empty transcriptions when it does not contain any
character.
      </p>
      <p>Training with fragments has two advantages over training with words (as in
segmentation free HTR):
the number of fragments needed for training is much smaller, because it does
not need to deal with the impressive variety of lexicon;
fragments can be labeled by volunteers in large transcription projects, with
little or no expertise, provided with adequate examples.</p>
      <p>To this end, we set up a simple crowd-sourcing application.</p>
      <p>Proof of Concept. Our hybrid work- ow takes the best of two worlds: we
can handle challenging ligatures as in state-of-the-art HTR, and at the same
time we require limited training e ort like typical OCR systems. Our system
is not mature enough for a thorough experimental evaluation, but for sake of
demonstration, we take into account the \Vatican Registrers" corpus in the
Vatican Secret Archives. The Vatican Registers is a huge collection of volumes
(more than 18.000 pages) produced in the 13-th century, and containing o cial
correspondence of the Roman Curia, such as political letters, opinions on legal
questions, and documents addressed to various religious institutes throughout
Europe. Such records are of unprecedented historical relevance, have not been
transcribed yet, and have a regular writing style (see Figure 1b). For these
reasons, we believe that they can motivate our work.</p>
      <p>Our crowd-sourcing application enrolled 120 high-school students in the city
of Rome, that did the labelling as a part of their work-related learning program.
The program is jointly organized by the engineering and humanities departments
of Roma Tre University, and includes frontal lessons in a variety of topics, such as
paleography, history and machine learning, thus also serving as school guidance.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <sec id="sec-2-1">
        <title>The idea of exploiting large textual corpora to de</title>
        <p>
          tect macroscopic cultural trends has been discussed
for many years [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], promising to empower
historians and other humanities scholars with a tool for
the study of culture and society.
        </p>
        <p>
          Data-driven history. Many studies have been
published over the past few years [
          <xref ref-type="bibr" rid="ref10 ref4">4, 10</xref>
          ] about a
quantitative and data-driven approach to the study
of cultural change and continuity. A seminal study
of 5 million English-language books published over
the arc of 200 years [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] showed the potential of this
approach, for example, measuring the time required
by various technologies to become established or the
duration of celebrity for various categories of people.
        </p>
        <p>Handwriting text recognition. HTR can be
de</p>
        <p>
          ned as the ability to transform handwritten input
represented as graphical marks into symbolic representation as ASCII text.
According to the mode of data acquisition used, HTR can be classi ed into o -line
and on-line. In o -line systems the handwriting is given as an image or scanned
text, without time sequence information. In on-line systems the handwriting is
given as a temporal sequence of coordinates that represents the pen tip
trajectory. For high quality text images, current HTR state-of-the-art prototypes
provide accuracy levels that range from 40 to 80% at the word level [
          <xref ref-type="bibr" rid="ref12 ref2">12, 2</xref>
          ]. On the
other hand, on-line systems are more accurate [
          <xref ref-type="bibr" rid="ref1 ref5 ref6">1, 5, 6</xref>
          ], reaching 90% word-level
accuracy in some cases.
        </p>
        <p>Crowd-sourcing. Expected users of HTR technology belong mainly to two
groups:
individual researchers with experience in handwritten documents;
volunteers which collaborate in large transcription projects.</p>
        <p>
          Recent HTR project [
          <xref ref-type="bibr" rid="ref14 ref8">14, 8</xref>
          ] expose HTR tools through specialised crowd-sourcing
web portals, supporting collaborative work.
        </p>
        <p>
          Other works. Language modeling for on-line handwriting recognition bears
many similarities with OCR and speech recognition, which often employ
statistical n-gram models on the character or word level [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>System</title>
    </sec>
    <sec id="sec-4">
      <title>Work-Flow</title>
      <p>
        Our system rst segments every word into small (possibly overlapping)
fragments, then recognizes characters in fragments, and nally the entire word. The
main steps are schematically shown in Figure 2.
1. Pre-processing. In this phase the color image of a page is cropped into
lines and words. The image is also transformed into a bi-chromatic one.
2. Cut-level operations. For each word we guess cut-points for characters
and build the so-called segmentation lattice data structure [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], such that
each path in the lattice represent a way of segmenting the word.
3. Fragment-level operations. For each pair of cut-points we crop the
corresponding text fragment and classify it, choosing among known characters.
4. Word-level operations. We nally return the best path in the labelled
lattice, which represent a way for transcribing the word.
3.1
      </p>
      <sec id="sec-4-1">
        <title>Transcription Algorithms</title>
        <p>In this section we describe the inner phases of our system. Inner phases are
designed for the Carolingian minuscule script, which is used in the manuscript
for our proof of concept (see Figure 1b of Section 1).</p>
        <p>Cut-level operations. The input for on-line HTR consists of pen up/pen down
switching movements, therefore cut-points for character boundaries can be
selected at certain positions of each stroke. Since we do not have access to stroke
sequences, we design a simple heuristic based on black pixel distribution local
minima, as shown in Figure 3a for a sample occurrence of the word \culpam".
Then, we consider all the possible segmentation options induced by the
cutpoints and let the further phases select the best one. We build a segmentation
lattice where each start-end path represents a segmentation option.</p>
        <p>Each cut-point corresponds to a node. The leftmost cut-point is the start
node, and the rightmost the end node.</p>
        <p>Each edge corresponds to the fragment of word bounded by its endpoints.
We observed experimentally that edges corresponding to fragments smaller than
8 pixels or larger than 34 pixels can be safely dropped our further consideration.
The segmentation lattice for our \culpam" example is shown in Figure 3b. In
the gure, edges have labels, as will be clari ed in the next section.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Fragment-level operations. We call</title>
        <p>fragment any squared portion of a
word that is bounded by two
cutpoints, including incomplete
characters, combinations of incomplete
characters, or multiple characters together.</p>
        <p>Each edge of the segmentation
lattice corresponds to a di erent
fragment, and each path corresponds to
a di erent segmentation option of the
word into fragments. Our next step is a
OCR step, where each fragment/edge
of the segmentation lattice is labeled
with the result of a character
classi</p>
        <p>
          er. There are many principled ways
for the classi cation task at hand. We
decide to use a convolutional Neural
Network (NN), that is one of the most
common and popular approaches [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Since the NN return a score
distribution, rather than a unique answer, we
add multiple edges between the same
two nodes. Some of the fragments are
submitted to a crowd-sourcing
application, for training the NN. The
application guides the workers, providing
sample images of characters and
highlighting variations in shape and style.</p>
        <p>Therefore, workers without experience
in transcription can give answers based
on their perceived similarity of
fragments to sample images. We refer to
the resulting data structure as labeled
lattice.
(a)
(b)
Fig. 3: Cut-points for the word
\culpam" and corresponding lattice. We use
green for actual character boundaries,
and red otherwise.</p>
        <p>Word-level operations. Labeled paths represent candidate transcriptions for
the current word. In order nd the best transcription recognition result, we
use language models (LM). A statistical LM is a representation of a certain
language as a probability distribution over sequences of words or characters.
In other words, a LM expresses the likelihood that certain sequences of words
or characters appear in texts written in the language under analysis. To this
end, we downloaded a large medieval Latin corpus ( 1:5M words) and
computed 3-grams frequencies. Then, we select the best path maximizing the
corresponding word probability. For instance, for the path \culham" in Figure 3b we
have7p($culham^) = p(cj$)p(uj$c)p(ljcu)p(hjul)p(ajlh)p(mjha)p(^jam), where $
and ^ are special symbols denoting the beginning and the end of a word. Every
term in the product can be computed directly from language models.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <sec id="sec-5-1">
        <title>We describe in this section our preliminary experi</title>
        <p>mental results, which serves us as a proof of concept
for In Codice Ratio. Ideas and methods in the proof
have been realized in collaboration with the Vatican
Secret Archives, with the aim of demonstrating in
principle the practical potential of our system.</p>
        <p>Dataset. The \Vatican Registrers" corpus consists
of 43 parchment registers, for a total of 18650 pages
(i.e., writing facades). All the registers are written
with the same script: the so-called Cancelleresca.</p>
        <p>Our dataset consists of 30 pages ( 15K characters)
of register 12 by Pope Honorii III, that is the only
Pope with un-transcribed registers and therefore is
of most interest for the VSR.
7 Using 2-nd order Markov assumption.
8 \uiuscemod" and \i" are processed as separate words.
future, we plan to repeat the training phase by removing such wrong items.
Finally, \culpam" is overthrown by \cullum" because of more serious errors in the
OCR process (\p" and \a" are labelled as \l" and \u', respectively). While word
probability does not help (3-gram \llu" is more frequent than \lpa" in our LM),
we are con dent to correct this kind of error in future works. Furture works
include the following.</p>
        <p>More sophisticated path selection criteria, for instance, taking into account
NN output score (second \l" has lower score than \p").</p>
        <p>More advanced language model tools, such as Hidden Markov Models. This
is currently being developed with promising result.</p>
        <p>Excluding labels that do not t in current line margins, such as excluding
an \l" where the lower margin contains black pixels, such as for \p".
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In Codice Ratio is an automatic transcription
work</p>
      <p>ow with low training e ort. Our proof of concept is
done on a high-resolution digitized copy of the reg- path prob
isters by pope Honorii III. Manuscript pages rst culpam
culnedaenrgvoersaiosneroiefsthoef ttreaxntsifmoramgea.tiTonhse,nt,hparteperxotcreascsteda cccuuulllllpuuanmmi 318 111000 747
pages are decomposed into fragments containing ba- criminis
sic text elements, such as characters and symbols. criminis 8 10 7
Some fragments are labelled by unskilled crowd- ccrriinnuuinuiss 46 1100 88
sourcing workers, which are asked simply to select uiuscemod(i)
matching images to template symbols selected by uiufemod 2 10 2
paleographers. Labelled symbols are used to train a uuifemod 2 10 3
Neural Networks, that is in charge of automatically uiiifemod 5 10 4
compute labels for all the un-labelled fragments. Table 1: Path probabilities
Automatically computed labels are aggregated at for the words in Figure 4.
the word-level into a segmentation lattice, in which
all the traversing paths represent candidate
transcriptions for the word. Selection of the best path is
done based on language models. Our proof of concept suggests that In Codice
Ratio can be applied to the large collection of Vatican Registers in the Vatican
Secret Archives, subject to speci cs improving, that are matter of ongoing work.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank Gaetano Bono glio, Veronica Iovinella e Andrea Salvoni for their
work on the OCR neural network. We thank Matteo Mariani for helping with
all the pre-processing steps. We also thank Gianlorenzo Didonato for developing
the front-end of the In Codice Ratio crowd-sourcing application. Finally, we are
indebted to all the teacher and students of Liceo Keplero and Liceo Montale who
joined the work-related learning program, and did all the labelling e ort.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>Character recognition experiments using unipen data</article-title>
          .
          <source>In Proceedings of the Sixth International Conference on Document Analysis and Recognition, ICDAR '01</source>
          , pages
          <fpage>481</fpage>
          {, Washington, DC, USA,
          <year>2001</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>R.</given-names>
            <surname>Bertolami</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Bunke</surname>
          </string-name>
          .
          <article-title>Hidden markov model-based ensemble methods for o ine handwritten text line recognition</article-title>
          .
          <source>Pattern Recognition</source>
          ,
          <volume>41</volume>
          (
          <issue>11</issue>
          ):
          <volume>3452</volume>
          {
          <fpage>3460</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ciregan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Meier</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          <article-title>. Multi-column deep neural networks for image classi cation</article-title>
          .
          <source>In Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <source>2012 IEEE Conference on</source>
          , pages
          <volume>3642</volume>
          {
          <fpage>3649</fpage>
          . IEEE,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>I.</given-names>
            <surname>Flaounas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
            Lansdall-Welfare, T. De Bie,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Mosdell</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
            , and
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Cristianini</surname>
          </string-name>
          .
          <article-title>Research methods in the age of digital journalism: Massive-scale automated analysis of news-contenttopics, style and gender</article-title>
          .
          <source>Digital Journalism</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <volume>102</volume>
          {
          <fpage>116</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Jaeger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Manke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Reichert</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Waibel</surname>
          </string-name>
          .
          <article-title>Online handwriting recognition: the npen++ recognizer</article-title>
          .
          <source>International Journal on Document Analysis and Recognition</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <volume>169</volume>
          {
          <fpage>180</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>D.</given-names>
            <surname>Keysers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Deselaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Rowley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-L.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Carbune</surname>
          </string-name>
          .
          <article-title>Multilanguage online handwriting recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>T.</given-names>
            <surname>Lansdall-Welfare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sudhahar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Thompson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. N.</given-names>
            <surname>Team</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Cristianini</surname>
          </string-name>
          .
          <article-title>Content analysis of 150 years of british periodicals</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>114</volume>
          (
          <issue>4</issue>
          ):E457{
          <fpage>E465</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Marcus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Parameswaran</surname>
          </string-name>
          , et al.
          <article-title>Crowdsourced data management: Industry and academic perspectives</article-title>
          . Foundations and Trends R in Databases,
          <volume>6</volume>
          (
          <issue>1</issue>
          -2):1{
          <fpage>161</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>J.-B. Michel</surname>
            ,
            <given-names>Y. K.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Aiden</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Veres</surname>
            ,
            <given-names>M. K.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            , ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>Pickett</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Hoiberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Clancy</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Norvig</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Orwant</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Pinker</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Nowak</surname>
            , and
            <given-names>E. L.</given-names>
          </string-name>
          <string-name>
            <surname>Aiden</surname>
          </string-name>
          .
          <article-title>Quantitative analysis of culture using millions of digitized books</article-title>
          .
          <source>Science</source>
          ,
          <volume>331</volume>
          (
          <issue>6014</issue>
          ):
          <volume>176</volume>
          {
          <fpage>182</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          . Distant reading.
          <source>Verso Books</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>R.</given-names>
            <surname>Reddy</surname>
          </string-name>
          and
          <string-name>
            <surname>G. StClair.</surname>
          </string-name>
          <article-title>The million book digital library project</article-title>
          . http://www. rr.cs.cmu.edu/mbdl.htm,
          <year>2016</year>
          . Accessed December 19,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>V.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Alabau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Bened</surname>
          </string-name>
          .
          <article-title>Combination of n-grams and stochastic context-free grammars in an o ine handwritten recognition system</article-title>
          .
          <source>In Proceedings of the 3rd Iberian Conference on Pattern Recognition and Image Analysis</source>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , IbPRIA '
          <volume>07</volume>
          , pages
          <fpage>467</fpage>
          {
          <fpage>474</fpage>
          , Berlin, Heidelberg,
          <year>2007</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>V.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Snchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Depuydt</surname>
          </string-name>
          , and J. de Does.
          <article-title>In uence of text line segmentation in handwritten text recognition</article-title>
          .
          <source>In 2015 13th International Conference on Document Analysis and Recognition (ICDAR)</source>
          , pages
          <fpage>536</fpage>
          {
          <fpage>540</fpage>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>J. A. Sanchez</surname>
            , G. Muhlberger,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Gatos</surname>
            , P. Scho eld, K. Depuydt,
            <given-names>R. M.</given-names>
          </string-name>
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Vidal</surname>
          </string-name>
          , and J. De Does.
          <article-title>transcriptorium: a european project on handwritten text recognition</article-title>
          .
          <source>In Proceedings of the 2013 ACM symposium on Document engineering</source>
          , pages
          <volume>227</volume>
          {
          <fpage>228</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. A.
          <string-name>
            <surname>Stolcke</surname>
          </string-name>
          et al.
          <article-title>Srilm-an extensible language modeling toolkit</article-title>
          .
          <source>In Interspeech</source>
          , volume
          <year>2002</year>
          , page
          <year>2002</year>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>