<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applying Computer Vision Systems to Historical Book Illustrations: Challenges and First Results</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yongho Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chanjong Im</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Schmideler</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wiebke Helm</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Hildesheim</institution>
          ,
          <addr-line>Information Science</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Leipzig, Faculty of Education</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>255</fpage>
      <lpage>260</lpage>
      <abstract>
        <p>Digital humanities still need to unlock the potential of images anlysis algorithms to a large extent. Modern deep learning images processing can contribute much to quantify knowledge about visual components in books. In this study, we report on experiments carried out for historical print. The illustrations in books offer much for humanities research. Object recognition systems can identify the portfolio of objects in book illustrations. In a study with several hundreds of books, we applied systems to find illustrations and classify them. Results show that persons are shown in illustrations within fiction books with a higher frequency than in non-fiction books. We also show the classification results for an analysis of the printing technology. This expert task can still not be perfectly modeled by a CNN. A class activation map analysis can be used to analyze the performance qualitatively.</p>
      </abstract>
      <kwd-group>
        <kwd>Digital Humanities</kwd>
        <kwd>Children Books</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>CNN</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Digital Humanities integrates automatic processing and analysis into research practices
in the Humanities. Image analysis is a growing area within Digital Humanities. The
analysis of books is of great interest to many disciplines. Digital historical corpora
allow the automatic access to illustrations in books and their analysis in large quantities.
This can lead to innovative research questions and quantitative results [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Historical children’s and youth books have yet not often been the subject of research.
Children books typically contain more images than adult books typically. As a
consequence, they are of special interest for an analysis of images. In addition, they form a
closed category on the one hand which contains sufficient variety on the other hand [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Illustrated books have played a significant role in knowledge dissemination. The
declining production costs for printed images have led to a growing exposure of more
and more people to rich visual resources. Research in this area could identify trends in
the objects depicted.</p>
      <p>Copyright © 2021 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
    </sec>
    <sec id="sec-2">
      <title>State of the Art</title>
      <p>
        The progress in computer vision and images analysis in the last decade was substantial.
Deep learning as a new direction in machine learning now can be considered as state of
the art and delivers excellent results in many domains. Deep learning refers to a
departure from feature engineering. Algorithms instead find a representation space for the
problem at hand by themselves. Neural networks have proven to be very successful for
this task. Already early approaches intended to input compresses feature spaces into
backpropagation neural networks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] but recently, new architectures have grown more
complex and successful.
      </p>
      <p>
        Convolution Neural Networks (CNN), the recent state of the art technology is known
to be very effective in automated feature detection and subsequent classification in
many domains. CNNs are composed of recurring sets of two layers: a convolution layer
and a pooling layer. The CNN combines pixels locally and by working through many
layers, more complex features can be extracted [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The applications of computer vision to the digital humanities is a growing area.
However, there are still not very many publications in this domain. One influential
experiment in the art domain by Salah &amp; Elgammel is dedicated to classify the painter of
artistic work. Such work is highly dependent on the type of paintings in the collection
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. An approach to identify objects within art work has also been presented. Similar to
our approach, it needs to deal with the domain shift and apply current technology to
historic print [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        A recent project is focusing on research on graphic novels. Current state of the art
CNNs are applied to tasks like author identification with very good success [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In
addition, the processing is aimed at measuring the drawing style of a graphic novel in
order to find similar book titles. A study of modern children books based on information
available in catalogues has analyzed market structures and book formats [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        A recent approach shows that the visual analysis of a page structure can be carried
out successfully with CNNs. A system can detect elements on a page and analyze tables
from heterogeneous layouts [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        One goal for the research of images in historical children books lies within the
production technologies. As a classification problem with few classes, it seems like a
challenge which could be solved with current technology. The classification of production
technology in the 19th century is still a hard task [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Comparable work is the
application of deep networks and transfer learning for material classification. Cimpoi et al.
conducted material classification with deeper structure and transfer learning [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Collections of Digitized Books</title>
      <p>
        This research is exploiting two collections of books that are digitized. The first
collection is the Wegehaupt corpus maintained by the Staatsbibliothek in Berlin [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The
second data collection is based on the Hobrecker collection [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This collection of
books is maintained by the library of the Technical University of Braunschweig.
All collections are of great interest for cultural research. They contain a rich variety of
different genres of children books mainly from Germany and mainly from the 19th
century: e.g. alphabetization books, picture books, biographies, natural history descriptions
as well as adventure stories. Another resource used is the database pictura paedagogica
online [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. It contains only images, many of which are extracted from books.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Results and Analysis</title>
      <p>We present some results for using a system with pre-defined classes and one with a
self-trained model. All models use deep learning systems and in particular some variant
of a CNN architecture.</p>
      <p>
        After extracting images from the book pages, we use them as input for a object
classification system pre-trained on modern photographs. In previous work, we could see
that the classification results are not perfect and that the greatly differ between books.
Typical performance metrics can lie between 30% and 90% for different books [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
which shows the heterogeneity of the material and the sensitivity of the systems for
that.
      </p>
      <p>
        Yolo to the images and record the recognized object types. The results of Yolo [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
have been recorded for a subset of 321 books from the Hobrecker collection. The
classification is not very reliable; a detailed evaluation is ongoing. However, for a statistical
analysis, it seems sufficient. We manually classified the fiction books and non-fiction
books. The analysis shows that there is no difference in the overall number of
illustrations for the two classes. However, the fiction books contain more images of humans
and horses and thus a more limited scope of object classes than is the case in the
nonfiction books. Non-fiction displaying many different objects and animals seem to cause
that difference (see table 1).
      </p>
      <sec id="sec-4-1">
        <title>Person per illustr.</title>
      </sec>
      <sec id="sec-4-2">
        <title>Person per illustr.</title>
        <p>1.143
0.737
1.446
1.151
It needs to be stressed that the distribution of the classes is highly skewed. The most
frequent classes are humans and a few animals. These results currently do not allow a
quantitative tracing of many different motifs through the century.</p>
        <p>For the study of historical print, questions of materiality are of great importance.
Issues of aesthetic design needs to be considered in relation to the techniques and
printing technologies available. Printing technology like wood cut, wood engraving and
lithography allowed different levels of elaboration. Finding the technique is a tedious
task. It is often not stated in the meta data and it requires experts to identify it from
digitized books.</p>
        <p>
          Therefore, the automatic identification is a important requirement. We trained a
model for distinguishing between three classes and managed to achieve only a
reasonable performance. However, for a statistical analysis, this can be sufficient.
To further analyze the errors and to observe how the algorithms differs from human
experts, we applied a qualitative analysis with class activation mapping (CAM)[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
They show the areas which were relevant for the system to classify the image into a
class. We can see that the algorithm does often not consider content parts of the image
but rather parts of the frame. Also, experts look at long lines or large areas.
Future research needs to also address stylistic and artistic aspects of illustrations. A
deeper analysis of content on a page and in particular of frequent classes (primarily
pictures of humans) offer great potential for advanced analysis tools for digital
humanists. We intend to develop a scene detection system which allows the study of typical
scene types like family, play, schooling and nature.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We thank the Fritz Thyssen Foundation for their funding for the research project
Distant Viewing. We thank the library of the Technische Universität Braunschweig, the
BBF | Bibliothek für Bildungsgeschichtliche Forschung Bibliothek
Bildungsgeschichtliche Forschung, Abteilung des DIPF | Leibniz-Institut für Bildungsforschung und
Bildungsinformation and the Staatsbibliothek Berlin (Preußischer Kulturbesitz) for
providing and facilitating access to their digitized collections.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bentkowska-Kafel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Debating Digital Art History</article-title>
          .
          <source>In: International Journal for Digital Art History</source>
          . vol.
          <volume>1</volume>
          https://doi.org/10.11588/dah.
          <year>2015</year>
          .
          <volume>1</volume>
          .
          <fpage>21634</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Schmideler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Lutherbilder. Ein Streifzug durch die Illustrationsgeschichte der Kinder-und Jugendliteratur des 18</article-title>
          . und 19.
          <string-name>
            <surname>Jahrhunderts</surname>
          </string-name>
          .
          <article-title>Die Reformation in der Kinder-und Jugendliteratur</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>Tolerant information retrieval with backpropagation networks</article-title>
          .
          <source>Neural Computing &amp; Applications</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <fpage>280</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Skansi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2018</year>
          ). Introduction to Deep Learning. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Saleh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Elgammal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Large-scale Classification of Fine-Art Paintings: Learning the Right Metric on The Right Feature</article-title>
          .
          <source>Intl. Journal for Digital Art History</source>
          , (
          <volume>2</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Crowley</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>The State of the Art: Object Retrieval in Paintings using Discriminative Regions</article-title>
          .
          <source>Proc. British Machine Vision Conference</source>
          . BMVA Press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dunst</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Hartel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Auf dem Weg zur Visuellen Stilometrie: Automatische Genre- und Autorunterscheidung in graphischen Narrativen</article-title>
          .
          <source>Kritik der digitalen Vernunft</source>
          . 5. Tagung „Digital Humanities im deutschsprachigen Raum“ http://dhd2018.unikoeln.de/wp-content/uploads/boa-DHd2018
          <string-name>
            <surname>-</surname>
          </string-name>
          web-ISBN.pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Steiner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Conservatism in an Innovative Field: Children's Digital Books in Sweden</article-title>
          .
          <source>DHN 2019 Digital Humanities in the Nordic Countries 4th Conference. ceurws.org/</source>
          Vol-2364
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lehenmeier</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burghardt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mischka</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2020</year>
          ,
          <article-title>August)</article-title>
          .
          <article-title>Layout Detection and Table Recognition-Recent Challenges in Digitizing Historical Documents and Handwritten Tabular Data</article-title>
          .
          <source>In International Conference on Theory and Practice of Digital Libraries</source>
          (pp.
          <fpage>229</fpage>
          -
          <lpage>242</lpage>
          ). Springer, Cham.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Im</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ghauri</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Rothman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Deep Learning Approaches to Classification of Production Technology for 19th Century Books</article-title>
          . LWDA, pp.
          <fpage>150</fpage>
          -
          <lpage>158</lpage>
          . http://ceurws.org/Vol-2191
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Cimpoi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maji</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kokkinos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Deep filter banks for texture recognition, description, and segmentation</article-title>
          .
          <source>International Journal of Computer Vision</source>
          , vol.
          <volume>118</volume>
          (
          <issue>1</issue>
          ) pp.
          <fpage>65</fpage>
          -
          <lpage>94</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Staatsbibliothek</surname>
          </string-name>
          zu Berlin, Preußischer Kulturbesitz. Wegehaupt Digital: https://digitalbeta.staatsbibliothek-berlin.de/suche?category[0]=Kinder- und
          <source>Jugendbücher&amp;queryString =project%3A"wegehauptdigital".</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>UB TU</surname>
          </string-name>
          <article-title>Braunschweig, Hobrecker Kollektion Online</article-title>
          . https://publikationsserver.tu-braunschweig.de/content/collections/childrens_books.xml
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Jornitz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kollmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Ins Bild hinein und aus dem Bild heraus. Anmerkungen zu Erfahrungen im Umgang mit einer pädagogischen Bild-Datenbank</article-title>
          .
          <source>MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung</source>
          ,
          <volume>9</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mitera</surname>
            , H.; Im,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Womser-Hacker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Objekterkennung in historischen Bilderbüchern: Eine Evaluierung des Potenzials von Computer Vision Algorithmen</article-title>
          . In:- KinderBuch. Reihe Studien zu Kinder- und
          <string-name>
            <surname>Jugendliteratur</surname>
          </string-name>
          und
          <article-title>-medien</article-title>
          . J.B.
          <string-name>
            <surname>Metzler</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Redmon</surname>
            , J.; Divvala,
            <given-names>S.</given-names>
          </string-name>
          ; Girshick,
          <string-name>
            <given-names>R.</given-names>
            and
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>You only look once: Unified, real-time object detection</article-title>
          .
          <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>779</fpage>
          -
          <lpage>788</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Selvaraju</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cogswell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vedantam</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Batra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Gradcam: Visual explanations from deep networks via gradient-based localization</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          pp.
          <fpage>618</fpage>
          -
          <lpage>626</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>