<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Genre Classi cation in the Library Catalog</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luschow</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Calvo Tello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Niedersachsische Staatsund Universitatsbibliothek Gottingen</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Library catalogs usually do not contain explicit information about the genre of literary works. However, both for readers and for researchers this information can be useful, e.g, when it comes to creating research corpora based on genre. In this proposal, we o er a rst analysis about how genre information has been integrated in catalogs to date, train several algorithms to classify annotated instances and contrast the results to the elds from the library catalogs used as features. Our study is based on data sets from the union catalog (GVK) of the German Common Library Network (GBV).</p>
      </abstract>
      <kwd-group>
        <kwd>library catalog</kwd>
        <kwd>machine learning</kwd>
        <kwd>genre analysis</kwd>
        <kwd>digital humanities</kwd>
        <kwd>bibliographic data science</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Genres are frequently used as categories to structure literary collections of texts.
This is observable in the way bookstores are organized or how publishers
market their products. But these categories do not only mediate in the economical
process, they also assist in the writing and reading process. Writers often take
actively a decision whether they are producing a novel, and if so, if it is, for
example, a historical novel [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. This is also taken into consideration by readers
who expect a series of characteristics from the genres [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ].
      </p>
      <p>
        Genres have played a central role in academic circles for many centuries,
frequently in a basic three-part schema, most often populated with the categories
drama, epic and lyric [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. For more speci c categories, sceptic opinions are
frequently hold [
        <xref ref-type="bibr" rid="ref10 ref9">9,10</xref>
        ]. However, numerous projects use them to de ne their research
object. Let us take three large projects of the past years in di erent European
countries: The Poetry Standardization and Linked Open Data1 (at the UNED
University in Madrid), The Riddle of Literary Quality2 (Huygens Institute for
the History of the Netherlands), and the Distant Reading for European Literary
History3 (coordinated from the University of Trier). None of them had the aim of
analyzing genres, but all three use genres to de ne their research object: poetry
in the rst case, novels in the latter two.
      </p>
      <p>
        The ubiquity of genres is also observable in literary corpora. Many corpora
use genres as a constant that constitutes a basic criterion for their composition:
dramatic texts in DraCor [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], novels in the ELTeC corpus and the Litbank
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In the case of the Textbox [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], each subcorpus also contain a single main
genre. In more generic corpora such as the German TextGrid Repository [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] or
the Spanish CORDE [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] genres are variable and therefore this information is
integrated as metadata for each text.
      </p>
      <p>
        The interest in genres has also been considered by many computational
approaches from Computer Science, Computational Linguistics, and Digital
Humanities [
        <xref ref-type="bibr" rid="ref16 ref17 ref18 ref24 ref25 ref28 ref3 ref34 ref7 ref8">3,18,28,25,24,17,16,8,34,7</xref>
        ]. These works have been applying mainly
supervised methods to mono-lingual corpora using linguistic cues from the full
text, such as the frequency of the tokens or linguistic annotation. When
multilabel classi cation is applied, the results tend to show relatively high scores, even
when the inter-annotator agreement tends to be low [
        <xref ref-type="bibr" rid="ref2 ref8">2,8</xref>
        ].
      </p>
      <p>Despite this relevance for authors, readers, publishers and researchers,
literary genres are seldom found in library catalogs, as we will explain in the following
section.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Genre in Bibliographic Data</title>
      <p>
        Today, library catalogs do not only serve as a means of nding and identifying
books held in a single library. Union catalogs such as the GVK (union catalog of
the German Common Library Network (GBV)) cover stocks of many di erent
libraries, including all types of media. Over the years an immense amount of
mostly intellectually generated (i.e., human-annotated) metadata was collected
that can be used far beyond its original purpose of registering and describing
books for local access [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Methods from computer science can be applied to
large data sets of bibliographic data to gain valuable insights into the inner
construction of the data, allowing for building tools to automatically structure,
group, enhance and enrich data. In the age of the rise of Digital Humanities,
libraries more and more see the urge and the value of opening their data silos
[
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] and thinking about making their data more useful for researchers.
      </p>
      <p>
        Quantitative approaches to using collections of bibliographic records4 as
research material have been studied to a lesser extent than, for example, the
analysis of full text resources. All too often, they are considered a \mere retrieval
tool" [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], useful in nding, identifying, and gaining access to resources, but not
so much as a resource by itself. Although library catalogs allow for|as
Bubenhofer and Rothenhausler put it|new text arrangements and operations that are
not possible by solely looking at the texts \themselves" [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Bibliographic records can be seen as big data [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Because sophisticated
cataloging standards and conventions exist in librarianship, one may come to the
assumption that these data are consistently structured and largely homogeneous.
However, this is not the case; bibliographic data are the results of multi-layered
4 Hereafter, we refer to catalog data sets as bibliographic data or bibliographic records,
respectively.
historical processes [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and shaped mainly due to di erent cataloging standards
applied over time [
        <xref ref-type="bibr" rid="ref23 ref36">23,36</xref>
        ].
      </p>
      <p>
        Lahti et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] refer to the kind of research that is based on catalog data
as bibliographic data science (BDS). BDS can easily be seen as part of a larger
process called Bibliomining [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] in which libraries use their large data corpora
to generate ndings and new knowledge about their central eld of competence.
Recently, Wallbank et al. [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] illustrated the complexity of bibliographic data
science by pointing out the need for extensive data cleaning|even if the data
extraction from the catalog follows very clear and distinct rules. Or, as Suarez
[
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] phrased it: \it is salutary to calibrate the instrument we are using".
      </p>
      <p>
        Studies focusing on genre information in bibliographic data are sparse and
allude to di erent aspects of the representation of genre, ranging from the
possibilities to use genre facets in user interfaces [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] to the signi cance of genre/form
aspects in American cataloging standards over time [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. In the Functional
Requirements for Bibliographic Records (FRBR), on which modern cataloging
standards such as Resource Description and Access (RDA) are based, the concept
of genre is one of the primary characteristics in distinguishing one work from
another. However, genre was not clearly de ned in the library sector to date [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
and its recording is not obligatory in most cases.
      </p>
      <p>
        Most of the studies related to BDS emphasize the importance of genre,
especially for making access to collections easier and (re)use of bibliographic data
more common. This point of view is already re ected in relevant documents, for
example in version 5.0 of the Metadata Application Pro le (MAP) of the Digital
Public Library of America (DPLA), where a list of values for the edm:hasType
property was compiled based on \areas of interest for researchers", amongst
others [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The values from this list are used in cataloging practice, thus adding to
a more uniform and widespread usage of genre attributes in bibliographic data.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Set, Genre Labels and Features</title>
      <p>In this section we describe our approach to enriching library catalogs with genre
information based on machine learning using already labeled data sets for
classi cation.5</p>
      <p>
        To compile our data set for classi cation we use keywords inside bibliographic
records and their relations to authority records. We extract all subject headings
from the German national authority le Gemeinsame Normdatei (GND)6 that
are classi ed in GND's own taxonomy [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] as \Literaturgattung" (literary genre).
This yields a total of 1319 di erent genre labels. For each genre, we save its label,
GND identi er and union catalog identi er. The union catalog identi er is then
used to retrieve all catalog records from the GVK associated with this keyword
using the Search/Retrieve via URL (SRU) protocol7 via the SRU-API o ered by
5 We also conducted experiments with unsupervised learning (clustering) but for this
short paper we focus only on the classi cation results.
6 https://www.dnb.de/DE/Professionell/Standardisierung/GND/gnd.html
7 http://www.loc.gov/standards/sru/
the GBV8. In total, we are able to download 746,786 records. All these records
are saved into a single data table, where each catalog eld (i.e., author, title,
keywords etc.) is represented by a single column.
      </p>
      <p>For this rst attempt|and to limit both the di culties of the heterogeneity
of the data and the computational costs|, instead of using the precise values in
each catalog eld as features we decide to use only binary features. This leads
to a table where for each record and each eld either a \1" or a \0" denotes
the presence of information in this eld for this record. Our hypothesis is that
the structure of a record can lead to indirect conclusions about the media type,
the publication format, the genre, or at least the cataloging standard used. To a
certain degree, the algorithm should be able to \re-create" the cataloging rules
based on the data elds used for a record, but also identify implicit rules for the
genre assignment.</p>
      <p>Out of 1319 literary genre labels from the authority le, only 1150 have at
least one associated bibliographic record in the union catalog; most genres are
related to very few records (mean: 612 records; median: 10; standard deviation:
6114.8; IQR: 49), and only 31 genres (2.6 %) are used in more than 2500 records
(cf. Fig. 1a).</p>
      <p>(a) original data
(b) main genres</p>
      <p>
        Genre information can be found in the catalog in a very scattered manner.
In fact, over 15 elds in the catalog records contain information related to genre
so it is not a trivial task to determine the genre a record belongs to. After
identifying these 15 elds we nd that 169,610 out of 746,786 records have multiple
genres assigned, i.e., 77.29 % of the records have a single genre assigned. Even if
other more exible theoretical models have been applied to explain the relations
between instances and genres [
        <xref ref-type="bibr" rid="ref16 ref25 ref8">25,16,8</xref>
        ], in this work we remain in the more
sim8 https://verbundwiki.gbv.de/display/VZG/SRU
ple model that each record is associated with only one genre. Thus, we decide to
keep only the rst occurring genre in the data if multiple genres are available for
a record. Additionally, many genres from the authority le are rather too
speci c which leads to very few records assigned to them, e.g., \Universitatsroman"
(college novel). We thus add a column to the data table in which each record's
original genre is mapped to a manually compiled list of 396 main genres (i.e.,
\college novel" will just become \novel"). However, the general distibution of
records across these main genres remains stable, cf. Fig. 1b, showing now 30
genres with more than 2500 records.
      </p>
      <p>Finally, to keep our training data manageable, we further observe only those
records that belong to the top 25 genres.9 We also ignore all catalog elds (i.e.,
columns in our data table) that are present in less than 0.1 % of the records.
Out of 1944 elds available in total, 870 remained (44.75 %). Based on these 870
features, we try to predict the one single genre label for each.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and Evaluation</title>
      <p>
        Using the Python programming language and its machine learning library
scikitlearn [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] we apply three classi cation algorithms to our data: Logistic
Regression, Random Forest [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Decision Tree [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. For all runs the same 10 % of
the data are used as a test set to evaluate the models.10 We evaluate the model
performance on both the original genres (\all") and our manually mapped main
genres (\main").
9 This is an arti cial limitation that allows for reasonable computational costs while
keeping enough analyzable data. At this point, we also have to place emphasis on
the fundamental problem that these top 25 genres are not on the same
conceptual level, containing for example \Roman" (novel) as well as \Prosa" (prose) and
\Fiktionale Darstellung" ( ction). This makes especially these broad categories very
heterogeneous and their constitution nondistinctive.
10 All data and source code are available at https://github.com/alueschow/qurator21
towards genre.
      </p>
      <p>As Table 1 shows, the less detailed main genres yield slightly worse
predicitions compared to when no genre mapping is used, maybe due to the more
heterogeneous and larger genre classes created by the mapping. In both cases,
Random Forest models perform notably better.</p>
      <p>(a) top 25 original genres
(b) top 25 main genres</p>
      <p>Figure 2 illustrates the model performance for the 25 genres observed, clearly
showing that the correct prediction rate highly depends on the individual genres.
Generally, speci c genres such as \Libretto" have a considerably better outcome
than generic genres, e.g., \Frauenliteratur" (women's literature). However, the
genre \Fiktionale Darstellung" ( ction) scores surprisingly well, considering its
very generic scope.</p>
      <p>The confusion matrix for the winning models visualized in Fig. 3 is based on
accuracy values for genre predictions. In both cases|original genres and main
genres|the false positives generated by the model seem to be placed
predominantly in one to two rather broad genres: \Erzahlung" (narrative) and \Epos"
(epic) for original genres, \Epos" and \Sprichwort" (proverb) for the main
genres. A more detailed analysis on the single genre level may be helpful in
determining the reasons behind these assignments, but is beyond the scope of this
short paper.</p>
      <p>(a) Random Forest (all)</p>
      <p>(b) Random Forest (main)
In this paper we have identi ed advancements in library cataloging|understood
as enhancements in data quality and access, but primarily their coverage|as
being fundamental for the widespread use of bibliographic records as a research
source by itself. The use of library catalog data sets is hitherto uncommon, albeit
rst examples of bibliographic data science (BDS) start to take soundings on how
this data can be most e ectively obtained, processed, and analyzed.</p>
      <p>Studying the example of classifying genre information based on bibliographic
data we were able to show that the application of well-established algorithms
(Logistic Regression, Random Forest, Decision Tree) already leads to a robust
performance. We used the existence of single data elds as binary features for
746,786 data sets and observed the most frequent 25 genre labels for both the
original genre assignments and a manually generated mapping to more broadly
de ned genres. Our results indicate that keeping the detailed genre information
may be bene cial for classi cation performance.</p>
      <p>In the near future, we would like to take into consideration other aspects
related to the genre labels (e.g., di erent conceptual levels of genres, deeper
analysis of individual genres, multi-label classi cation, arti cially created distinctive
genres), the features (feature engineering, feature analysis, multilinguality of the
features), resources used for training and the methods used. In any case, in this
rst step we have systematically analyzed genre labels in bibliographic data. Our
hope is that we can awaken the interest in libraries to make their catalogs more
useful for (literary) research through the categories of literary genre.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bamman</surname>
          </string-name>
          , D.:
          <article-title>LitBank: Born-Literary Natural Language Processing</article-title>
          . In: Johnson,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Mimno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Tilton</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (eds.) Computational Humanites, Debates in Digital Humanities (
          <year>2020</year>
          , preprint), https://people.ischool.berkeley.edu/ dbamman/ pubs/pdf/Bamman DH Debates CompHum.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berninger</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ross</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Building a Document Genre Corpus: a Pro le of the KRYS I corpus</article-title>
          .
          <source>In: BCS-IRSG Workshop on Corpus Pro ling. London</source>
          (
          <year>2008</year>
          ), http://eprints.gla.ac.uk/47699/1/ewic2 ir08 s1paper2.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Biber</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The multidimensional approach to linguistic analyses of genre variation: An overview of methodology and nding</article-title>
          .
          <source>Computers in the Humanities</source>
          <volume>26</volume>
          (
          <issue>5-6</issue>
          ),
          <volume>331</volume>
          {
          <fpage>347</fpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olshen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stone</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>Classi cation and Regression Trees</article-title>
          . Wadsworth, Belmont, CA (
          <year>1983</year>
          ). https://doi.org/10.2307/2530946
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.: Random</given-names>
          </string-name>
          <string-name>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Machine Learning 45(1)</source>
          ,
          <volume>5</volume>
          {
          <fpage>32</fpage>
          (
          <year>2001</year>
          ). https://doi.org/10.1023/A:1010933404324
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bubenhofer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , Rothenhausler, K.:
          <article-title>\Korporatheken": Die digitale und verdatete Bibliothek</article-title>
          .
          <volume>027</volume>
          .7 Zeitschrift fur Bibliothekskultur /
          <source>Journal for Library Culture</source>
          <volume>4</volume>
          (
          <issue>2</issue>
          ),
          <volume>60</volume>
          {
          <fpage>71</fpage>
          (
          <year>2016</year>
          ). https://doi.org/10.12685/027.7-
          <issue>4</issue>
          -2-154
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Calvo</given-names>
            <surname>Tello</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>The Novel in the Spanish Silver Age: A Digital Analysis of Genre through Machine Learning</article-title>
          . University of Wurzburg,
          <source>PhD Thesis</source>
          (unpublished)
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Calvo</given-names>
            <surname>Tello</surname>
          </string-name>
          , J.:
          <article-title>Genre Classi cation in Spanish Novels: A Hard Task for Humans and Machines? In: Data in Digital Humanities</article-title>
          . EADH,
          <string-name>
            <surname>Galway</surname>
          </string-name>
          (
          <year>2018</year>
          ), https:// eadh2018.exordo.com/programme/presentation/82
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Estetica come scienza dell'espressione e linguistica generale</article-title>
          . Sandron, Milano, Italy (
          <year>1902</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Derrida</surname>
          </string-name>
          , J.:
          <source>The Law of Genre. Critical Inquiry</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <volume>55</volume>
          {
          <fpage>81</fpage>
          (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Digital</surname>
          </string-name>
          <article-title>Public Library of America: Metadata Application Pro le</article-title>
          ,
          <source>version 5</source>
          .0 (
          <issue>2017</issue>
          ), https://drive.google.com/ le/d/1fJEWhnYy5Ch7 ef -V48-FAViA72OieG/view
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Dragon</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          :
          <article-title>Form and Genre Access to Academic Library Digital Collections</article-title>
          .
          <source>Journal of Library Metadata</source>
          <volume>20</volume>
          (
          <issue>1</issue>
          ),
          <volume>29</volume>
          {
          <fpage>49</fpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1080/19386389.
          <year>2020</year>
          .1723203
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trilcke</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jennifer</surname>
            <given-names>Beine</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Orekhov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            :
            <surname>Dracor</surname>
          </string-name>
          (
          <year>2018</year>
          ), https:// dracor.org/
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Genette</surname>
          </string-name>
          , G.:
          <article-title>Genres, types, modes</article-title>
          .
          <source>Poetique</source>
          <volume>32</volume>
          ,
          <issue>389</issue>
          {
          <fpage>421</fpage>
          (
          <year>1977</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Gorman</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A Systems Librarian's Cataloging Daydream</article-title>
          . In: Sanchez, E. (ed.)
          <source>Conversations With Catalogers in the 21st Century</source>
          , pp.
          <volume>73</volume>
          {
          <fpage>94</fpage>
          .
          <string-name>
            <surname>Libraries</surname>
            <given-names>Unlimited</given-names>
          </string-name>
          , Santa Barbara, California (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Henny-Krahmer</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Betz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , Schlor, D.,
          <string-name>
            <surname>Hotho</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Alternative Gattungstheorien: Das Prototypenmodell am Beispiel hispanoamerikanischer Romane</article-title>
          . In: DHd 2018:
          <article-title>Kritik der digitalen Vernunft</article-title>
          .
          <source>Konferenzabstracts</source>
          . pp.
          <volume>105</volume>
          {
          <fpage>112</fpage>
          . Koln (
          <year>2018</year>
          ), http: //dhd2018.uni-koeln.de/wp-content/uploads/boa-DHd2018
          <string-name>
            <surname>-</surname>
          </string-name>
          web-ISBN.pdf
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Hettinger</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reger</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jannidis</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hotho</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Classi cation of Literary Subgenres</article-title>
          . In: DHd 2016:
          <article-title>Modellierung{Vernetzung{Visualisierung. Die Digital Humanities als facherubergreifendes Forschungsparadigma</article-title>
          .
          <source>Konferenzabstracts</source>
          . pp.
          <volume>154</volume>
          {
          <fpage>158</fpage>
          .
          <string-name>
            <surname>Leipzig</surname>
          </string-name>
          (
          <year>2016</year>
          ), http://dhd2016.de/boa.pdf
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kessler</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Numberg</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Schutze, H.:
          <article-title>Automatic detection of text genre</article-title>
          . In:
          <article-title>Proceedings of the 35th Annual Meeting of the Association for Computational Linguistics and Eighth Conference of the European Chapter of the Association for Computational Linguistics</article-title>
          . pp.
          <volume>32</volume>
          {
          <fpage>38</fpage>
          . ACL '
          <volume>98</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>1997</year>
          ). https://doi.org/10.3115/976909.979622
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marjanen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roivainen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolonen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Bibliographic Data Science and the History of the Book (c</article-title>
          . 1500{
          <year>1800</year>
          ).
          <source>Cataloging &amp; Classi cation Quarterly</source>
          <volume>57</volume>
          (
          <issue>1</issue>
          ),
          <volume>5</volume>
          {
          <fpage>23</fpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1080/01639374.
          <year>2018</year>
          .1543747
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Lukacs</surname>
          </string-name>
          , G.:
          <article-title>Der Historische Roman</article-title>
          . Aufbau-Verlag, Berlin (
          <year>1955</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Nicholson</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The basis for bibliomining: Frameworks for bringing together usage-based data mining and bibliometrics through data warehousing in digital library services</article-title>
          .
          <source>Information Processing Management</source>
          <volume>42</volume>
          (
          <issue>3</issue>
          ),
          <volume>785</volume>
          {
          <fpage>804</fpage>
          (
          <year>2006</year>
          ). https://doi.org/10.1016/j.ipm.
          <year>2005</year>
          .
          <volume>05</volume>
          .008
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , et al.:
          <article-title>Scikit-learn: Machine learning in</article-title>
          <source>Python 12</source>
          ,
          <issue>2825</issue>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Prescott</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Bibliographic Records as Humanities Big Data</article-title>
          . In: Hu,
          <string-name>
            <surname>X.</surname>
          </string-name>
          , et al. (eds.)
          <source>2013 IEEE International Conference on Big Data</source>
          . pp.
          <volume>55</volume>
          {
          <issue>58</issue>
          (
          <year>2013</year>
          ). https://doi.org/10.1109/BigData.
          <year>2013</year>
          .6691670
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Riddell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Schoch, C.:
          <article-title>Progress through Regression</article-title>
          .
          <source>In: Digital Humanities</source>
          <year>2014</year>
          :
          <article-title>Conference Abstracts</article-title>
          . UNIL/EPFL, Lausanne (
          <year>2014</year>
          ), http://dharchive.org/ paper/DH2014/Paper-60.xml
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Santini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Automatic Identi cation of Genre in Web Pages: A new perspective</article-title>
          . LAP Lambert Academic Publishing, Saarbrucken (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Schmunk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Funk</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Das DARIAH-DE- und das</surname>
          </string-name>
          TextGrid-Repositorium:
          <article-title>Geistes- und kulturwissenschaftliche Forschungsdaten persistent und referenzierbar langzeitspeichern</article-title>
          .
          <source>Bibliothek Forschung und Praxis</source>
          <volume>40</volume>
          (
          <issue>2</issue>
          ) (
          <year>2016</year>
          ). https://doi.org/10.1515/bfp-2016
          <source>-0020</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27. Schoch,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Calvo</surname>
          </string-name>
          <string-name>
            <surname>Tello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Henny-Krahmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            ,
            <surname>Popp</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>The CLiGS textbox: Building and Using Collections of Literary Texts in Romance Languages Encoded in XML-TEI</article-title>
          .
          <article-title>Journal of the Text Encoding Initiative (Rolling Issue) (</article-title>
          <year>2019</year>
          ). https://doi.org/10.4000/jtei.2085
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fakotakis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kokkinakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Automatic Text Categorization in Terms of Genre and Author</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>26</volume>
          (
          <issue>4</issue>
          ),
          <volume>471</volume>
          {
          <fpage>497</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Stella</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scheven</surname>
          </string-name>
          , E. (eds.):
          <article-title>GND-Systematik: Leitfaden zu ihrer Vergabe</article-title>
          .
          <source>Deutsche Nationalbibliothek</source>
          , Leipzig (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Suarez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Towards a bibliometric analysis of the surviving record, 1701{1800</article-title>
          . In: Suarez, Michael F.,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Turner</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.L</surname>
          </string-name>
          . (eds.)
          <source>The Cambridge History of the Book in Britain: Volume</source>
          <volume>5</volume>
          :
          <fpage>1695</fpage>
          {
          <year>1830</year>
          , pp.
          <volume>37</volume>
          {
          <fpage>65</fpage>
          . Cambridge University Press (
          <year>2009</year>
          ). https://doi.org/10.1017/CHOL9780521810173.003
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , Hyvonen, N.:
          <article-title>From MARC silos to Linked Data silos? o-bib 4</article-title>
          <issue>, 1</issue>
          {
          <fpage>13</fpage>
          (
          <year>2017</year>
          ). https://doi.org/10.5282/O-BIB/2017H2S1-
          <fpage>13</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Sanchez Sanchez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Dom nguez Cintas,
          <string-name>
            <surname>C.</surname>
          </string-name>
          : El banco de datos de la RAE:
          <article-title>CREA y CORDE</article-title>
          .
          <article-title>Per Abbat: bolet n lologico de actualizacion academica y didactica (2</article-title>
          ),
          <volume>137</volume>
          {
          <fpage>148</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Todorov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <source>The Origin of Genres. New Literary History</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <volume>159</volume>
          {
          <fpage>170</fpage>
          (
          <year>1976</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Underwood</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Distant Horizons: Digital Evidence and Literary Change</article-title>
          . University of Chicago Press, Chicago (
          <year>2019</year>
          ). https://doi.org/10.7208/9780226612973
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Wallbank</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kane</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dickerson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hutchinson</surname>
          </string-name>
          , J.:
          <source>Exploring Bibliographic Records as Research Data. Catalogue and Index</source>
          <volume>197</volume>
          ,
          <issue>3</issue>
          {
          <issue>9</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          :
          <article-title>The Role of Genre in the Bibliographic Universe</article-title>
          .
          <source>Advances in Classi cation Research Online</source>
          <volume>23</volume>
          (
          <issue>1</issue>
          ),
          <volume>38</volume>
          {
          <fpage>45</fpage>
          (
          <year>2012</year>
          ). https://doi.org/10.7152/acro.v23i1.
          <fpage>14236</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>